An evaluation workshop

Make the test.
Understand the result.

Explore six kinds of evals through recorded experiments. Inspect the grader, compare candidates, and see where a change improves—or fails.

Learn from public incidents →

Explore evidence. Then try your own eval.

This viewer opens actual recorded runs of the bundled local rules. Five suites use synthetic examples; response quality uses a small attributed human-preference sample. These are teaching results, not model benchmarks.

Read the six walkthroughs → · Run and tune the lab locally →

01 · EXPERIMENT

Choose your eval

Comparison, repeated trials, and CI gates

Repeated trials show variability on the same cases; they do not add independent test cases. Registered Python/HTTP candidates invoke your application when run.

Tune this eval ↓

Loading recorded experiments…

02 · TUNE THE EVAL

Define what good means.

Set the expected answer before running. Save a local profile or export it for the command line.

Bundled defaults

Saved profiles and experiments stay on this computer. Earlier profiles remain in the history folder.

03 · SAVED EXPERIMENTS

Keep the evidence.

Reopen an experiment, import a JSON report, or compare compatible saved runs. For saved comparisons, the candidate after the change is used.

Latest 200 experiments shown. Imported reports are user-supplied evidence and are not authenticated. Review case content before sharing.