An evaluation workshop

Recorded offline experiments

Twelve actual offline comparisons, including holdout failures and regressions.

Generated by python -m scripts.reproduce_portfolio. These are actual runs of local rules, not model performance claims. No API calls were made. Timing depends on the machine. Download an HTML file and open it locally, or import JSON through the webpage. Reproduce into an ignored directory with uv run --locked python -m scripts.reproduce_portfolio. The source and grader fingerprints change when implementation changes; recorded reports are comparable with each other, but may require rerunning to compare with a later checkout.

Suite Split Baseline passes Improved passes Regressions Evidence
classification dev 14/20 20/20 0 JSON · HTML
classification holdout 4/10 7/10 0 JSON · HTML
extraction dev 10/20 20/20 0 JSON · HTML
extraction holdout 3/10 8/10 0 JSON · HTML
tool_calling dev 12/20 20/20 0 JSON · HTML
tool_calling holdout 5/10 8/10 1 JSON · HTML
rag dev 2/8 8/8 0 JSON · HTML
rag holdout 1/4 1/4 0 JSON · HTML
response_quality dev 3/6 5/6 0 JSON · HTML
response_quality holdout 2/6 1/6 1 JSON · HTML
agent dev 4/8 8/8 0 JSON · HTML
agent holdout 1/4 4/4 0 JSON · HTML