An evaluation workshop

Six ways to ask: does it work?

Small datasets. Transparent graders. Changes that sometimes make things worse.

Follow the whole evaluation loop.

Define success, choose cases, inspect a grader, compare a change, and check the holdout. Each walkthrough explains a failure you can reproduce and what its score cannot establish.

EXAMPLE 04

Retrieve and answer with evidence

Retrieval, answer correctness, and supporting citations can fail independently. Exact evidence checks do not measure general semantic truth.

EXAMPLE 05

Calibrate a pairwise judge

Human agreement is not factual correctness. Inspect disagreements, noisy labels, and changes when A/B positions are swapped.

EXAMPLE 06

Evaluate a multi-step agent

Replay the trace to verify state transitions. A confident final answer cannot replace successful tools, authorization, or a call budget.