EXAMPLE 01
Classify a support ticket
A high overall score can hide failures in one language or business rule.
An evaluation workshop
Small datasets. Transparent graders. Changes that sometimes make things worse.
Define success, choose cases, inspect a grader, compare a change, and check the holdout. Each walkthrough explains a failure you can reproduce and what its score cannot establish.
EXAMPLE 01
A high overall score can hide failures in one language or business rule.
EXAMPLE 02
Valid JSON is only the first check. The values must also match the request.
EXAMPLE 03
Check the tool, its arguments, and what actually happens when the simulator executes it.
EXAMPLE 04
Retrieval, answer correctness, and supporting citations can fail independently. Exact evidence checks do not measure general semantic truth.
EXAMPLE 05
Human agreement is not factual correctness. Inspect disagreements, noisy labels, and changes when A/B positions are swapped.
EXAMPLE 06
Replay the trace to verify state transitions. A confident final answer cannot replace successful tools, authorization, or a call budget.