An evaluation workshop

Practical Eval Lab

Explore the recorded runs here. Clone the toolkit to run candidates and tune your own evals.

Six runnable evaluation examples with a local tuning webpage. Learn how to choose success criteria, inspect failures, compare application versions, and keep the results reproducible. Every example works without an API key.

Status: a public, open-source educational toolkit. Original code and synthetic data use MIT-0. The human-preference sample retains its upstream MIT notice.

Explore recorded results · Read the six walkthroughs · Learn from public incidents

Agent guide · Download documentation context · Read code with GitIngest

The GitHub Pages site lets you inspect actual offline runs and download evidence. Run the local lab below to edit cases, execute candidates, and save experiments.

Start in a source checkout

Use Python 3.10+ and uv:

git clone https://github.com/hk-775/practical-eval-lab.git
cd practical-eval-lab
uv sync --locked
uv run --locked eval-lab serve

Open http://127.0.0.1:8000. Choose an example and click Compare candidates. Edit development cases, references, grading settings, or the passing threshold; then rerun and save the profile. Stop the server with Ctrl+C.

For an occupied port: uv run --locked eval-lab serve --port 8766. The dependency-free source workflow also works with python -m eval_lab serve. The original server.py and eval.py entry points remain available.

See the tuning webpage.

Choose a lesson

Example Skill demonstrated Dev / holdout Walkthrough
Classification Accuracy, confusion matrix, macro-F1, slice analysis 20 / 10 Classify tickets
Structured extraction Schema validity versus field correctness 20 / 10 Extract orders
Tool calling Tool selection, arguments, outcomes, regressions 20 / 10 Evaluate a tool decision
RAG Retrieval recall, reference answers, citation evidence, abstention 8 / 4 Retrieve and answer
Response quality Rubrics, blinded pairwise judging, agreement with human preferences 6 / 6 Calibrate a judge
Multi-step agent Trace replay, authorization, retries, budgets, task completion 8 / 4 Evaluate a workflow

The 126 cases are teaching material: 114 synthetic cases and 12 attributed human preference pairs. Built-in candidates are local rules, not trained models. The candidate named improved describes an intended change, not a promise of better results. Its response-quality holdout score actually regresses.

Recorded JSON/HTML experiments · Learning guide · Public incidents and proposed evals

Evaluate an executed support workflow

The support-workflow evaluation asks whether pinned Strands and Laya checkpoints can select tools and arguments, request missing information, and complete an executable synthetic workflow. It compares them with rules and fitted baselines, calibrates on separate wording families, and records actual tool outcomes, fallback demand, and simulator-episode latency. Model dependencies have a separate uv lockfile; no API key is needed.

The earlier decision-model diagnostic and its design review are retained as a separate experiment. Jev has not been evaluated.

Evaluate your application

Compare two named Python candidates without modifying the runner:

uv run --locked eval-lab compare --suite classification \
  --project examples/python-project.json --before app-v1 --after app-v2 \
  --report reports/application.json --html reports/application.html

Expose those registrations in the webpage:

uv run --locked eval-lab serve --project examples/python-project.json

A complete local HTTP application and endpoint configurations are included too. The integration guide covers Python, HTTP, optional OpenAI calls, environment-based credentials, and replacing the demonstration application. A project file is trusted executable configuration; the browser cannot register arbitrary code or endpoints.

Save and compare evidence

The webpage saves every run and comparison automatically. Reopen saved experiments, import JSON reports, compare compatible saved runs, and export standalone HTML. Reports include case-level outputs, grader checks, slices, gate decisions, candidate identity, fingerprints, candidate/grader timing, and token usage when supplied.

uv run --locked eval-lab compare --suite rag \
  --report reports/rag.json --html reports/rag.html
uv run --locked eval-lab export reports/rag.json --html reports/rag.html
uv run --locked eval-lab run --suite response_quality --candidate improved --trials 3

Repeated trials expose variability on the same cases; they are not independent new test examples. No estimated costs or statistical significance are inferred.

Source-checkout state lives in ignored local/: profiles/, profile history/, and saved runs/. Installed packages use a writable per-user directory. Override with --state-dir PATH or EVAL_LAB_HOME. The CLI uses bundled cases unless given --config or --cases; it never silently loads webpage edits.

Block a regression in CI

uv run --locked eval-lab compare --suite tool_calling --split holdout \
  --fail-on-regression --critical-tag vocabulary --min-slice lookup=1

This intentionally exits 1 despite the average improving from 50% to 80%: one lookup regresses, and critical/slice checks fail. Quality gates apply to the candidate after the change; execution failures in either candidate remain errors.

Exit Meaning
0 Quality gates pass and candidates executed successfully
1 Quality or regression gate failed
2 Invalid configuration, incompatible comparison, or file error
3 At least one candidate execution failed

Install a built artifact

uv build
uv tool install ./dist/practical_eval_lab-0.2.0-py3-none-any.whl
eval-lab serve

This installs an artifact you build locally. No registry package or GitHub Release is claimed. Wheels contain the webpage, datasets, provenance, and license notices. Wheel and source-distribution installs are exercised outside the checkout in CI.

Verify and extend

uv sync --locked
uv run --locked pytest -q
uv run --locked playwright install chromium
uv run --locked python -m scripts.browser_check
uv run --locked python -m scripts.package_check
uv run --locked python -m scripts.reproduce_portfolio

CI covers Python 3.10 and 3.14 on Ubuntu and Python 3.13 on Ubuntu, Windows, and macOS, with Chromium checks on Ubuntu. Tests exercise grader failure paths, real loopback HTTP integration, report persistence/import, gates, packaging, and the six webpage flows. The optional OpenAI adapter is tested with a mock response; no paid live call has been used to substantiate the bundled results.

Dataset, candidate, grader, report, and local tuning profile

Contracts and extension points · Data provenance · Website build and hosting · Architecture page · Publication inventory · Contributing · Support · Security · Community conduct

This is a small local toolkit. It does not provide a hosted service, distributed execution, general semantic grounding, a validated safety judge, or production certification. The walkthroughs state what each grader can and cannot establish.