Pass rate
—Explore evidence. Then try your own eval.
This viewer opens actual recorded runs of the bundled local rules. Five suites use synthetic examples; response quality uses a small attributed human-preference sample. These are teaching results, not model benchmarks.
Read the six walkthroughs → · Run and tune the lab locally →
01 · EXPERIMENT
Choose your eval
Loading recorded experiments…
Change
—Passing gate
—| Case | Expected | Baseline | Result |
|---|
No cases match this filter.
LOOK BENEATH THE AVERAGE
Scores by tag
Measurement details · checks, latency, usage, and repeatability
02 · TUNE THE EVAL
Define what good means.
Set the expected answer before running. Save a local profile or export it for the command line.
Holdout uses the original cases, grader settings, and 80% threshold. Return to Development to edit. These are small teaching examples, not a model benchmark.
Saved profiles and experiments stay on this computer. Earlier profiles remain in the history folder.
03 · SAVED EXPERIMENTS
Keep the evidence.
Reopen an experiment, import a JSON report, or compare compatible saved runs. For saved comparisons, the candidate after the change is used.
Latest 200 experiments shown. Imported reports are user-supplied evidence and are not authenticated. Review case content before sharing.