Practical Eval Lab — public documentation context Generated from allowlisted source documents. Source fingerprints describe the documentation; recorded experiments retain their own dates and implementation fingerprints. ===== FILE: index.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/index.md # Practical Eval Lab > Six runnable AI evaluation examples, transparent graders, and a local tuning webpage. This website displays recorded offline results. Candidate execution and tuning run locally. Five suites use synthetic cases; response quality uses a small attributed Anthropic human-preference sample. Bundled candidates are local rules. Results demonstrate evaluation methods, not production model quality or safety certification. Created by [Harleen Kaur](https://hk-775.github.io/hk-775/). [Source repository](https://github.com/hk-775/practical-eval-lab) · [Local setup](https://hk-775.github.io/practical-eval-lab/getting-started.md). ## Walkthroughs - [Classify a support ticket](https://hk-775.github.io/practical-eval-lab/guides/classification.md): A high overall score can hide failures in one language or business rule. - [Extract an order](https://hk-775.github.io/practical-eval-lab/guides/extraction.md): Valid JSON is only the first check. The values must also match the request. - [Choose and call a tool](https://hk-775.github.io/practical-eval-lab/guides/tool-calling.md): Check the tool, its arguments, and what actually happens when the simulator executes it. - [Retrieve and answer with evidence](https://hk-775.github.io/practical-eval-lab/guides/rag.md): Retrieval, answer correctness, and supporting citations can fail independently. Exact evidence checks do not measure general semantic truth. - [Calibrate a pairwise judge](https://hk-775.github.io/practical-eval-lab/guides/response-quality.md): Human agreement is not factual correctness. Inspect disagreements, noisy labels, and changes when A/B positions are swapped. - [Evaluate a multi-step agent](https://hk-775.github.io/practical-eval-lab/guides/agent.md): Replay the trace to verify state transitions. A confident final answer cannot replace successful tools, authorization, or a call budget. [Recorded JSON and HTML evidence](https://hk-775.github.io/practical-eval-lab/results.md) · [Data provenance and attribution](https://hk-775.github.io/practical-eval-lab/data-provenance.md). --- Source: [scripts/build_pages.py](https://github.com/hk-775/practical-eval-lab/blob/main/scripts/build_pages.py) Source SHA-256: `e6c887df7116f9b8a2bc4997a3cc26d15654cc301e835ddf1d1d2b6e0ec53ce2` ===== FILE: getting-started.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/getting-started.md # Practical Eval Lab Six runnable evaluation examples with a local tuning webpage. Learn how to choose success criteria, inspect failures, compare application versions, and keep the results reproducible. Every example works without an API key. **Status:** a public, open-source educational toolkit. Original code and synthetic data use [MIT-0](https://hk-775.github.io/practical-eval-lab/downloads/LICENSE.txt). The human-preference sample retains its [upstream MIT notice](https://hk-775.github.io/practical-eval-lab/notices.md). [Explore recorded results](https://hk-775.github.io/practical-eval-lab/) · [Read the six walkthroughs](https://hk-775.github.io/practical-eval-lab/guides.html) · [Learn from public incidents](https://hk-775.github.io/practical-eval-lab/incidents.html) [Agent guide](https://hk-775.github.io/practical-eval-lab/llms.txt) · [Download documentation context](https://hk-775.github.io/practical-eval-lab/agent-context.txt) · [Read code with GitIngest](https://gitingest.com/hk-775/practical-eval-lab) The GitHub Pages site lets you inspect actual offline runs and download evidence. Run the local lab below to edit cases, execute candidates, and save experiments. ## Start in a source checkout Use Python 3.10+ and [uv](https://docs.astral.sh/uv/): ```bash git clone https://github.com/hk-775/practical-eval-lab.git cd practical-eval-lab uv sync --locked uv run --locked eval-lab serve ``` Open . Choose an example and click **Compare candidates**. Edit development cases, references, grading settings, or the passing threshold; then rerun and save the profile. Stop the server with Ctrl+C. For an occupied port: `uv run --locked eval-lab serve --port 8766`. The dependency-free source workflow also works with `python -m eval_lab serve`. The original `server.py` and `eval.py` entry points remain available. [See the tuning webpage](https://hk-775.github.io/practical-eval-lab/assets/tuning-lab.png). ## Choose a lesson | Example | Skill demonstrated | Dev / holdout | Walkthrough | |---|---|---:|---| | Classification | Accuracy, confusion matrix, macro-F1, slice analysis | 20 / 10 | [Classify tickets](https://hk-775.github.io/practical-eval-lab/guides/classification.md) | | Structured extraction | Schema validity versus field correctness | 20 / 10 | [Extract orders](https://hk-775.github.io/practical-eval-lab/guides/extraction.md) | | Tool calling | Tool selection, arguments, outcomes, regressions | 20 / 10 | [Evaluate a tool decision](https://hk-775.github.io/practical-eval-lab/guides/tool-calling.md) | | RAG | Retrieval recall, reference answers, citation evidence, abstention | 8 / 4 | [Retrieve and answer](https://hk-775.github.io/practical-eval-lab/guides/rag.md) | | Response quality | Rubrics, blinded pairwise judging, agreement with human preferences | 6 / 6 | [Calibrate a judge](https://hk-775.github.io/practical-eval-lab/guides/response-quality.md) | | Multi-step agent | Trace replay, authorization, retries, budgets, task completion | 8 / 4 | [Evaluate a workflow](https://hk-775.github.io/practical-eval-lab/guides/agent.md) | The 126 cases are teaching material: 114 synthetic cases and 12 attributed human preference pairs. Built-in candidates are local rules, not trained models. The candidate named `improved` describes an intended change, not a promise of better results. Its response-quality holdout score actually regresses. [Recorded JSON/HTML experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Learning guide](https://hk-775.github.io/practical-eval-lab/learning-guide.md) · [Public incidents and proposed evals](https://hk-775.github.io/practical-eval-lab/incidents.md) ## Evaluate an executed support workflow The [support-workflow evaluation](https://hk-775.github.io/practical-eval-lab/tool-workflow.md) asks whether pinned Strands and Laya checkpoints can select tools and arguments, request missing information, and complete an executable synthetic workflow. It compares them with rules and fitted baselines, calibrates on separate wording families, and records actual tool outcomes, fallback demand, and simulator-episode latency. Model dependencies have a separate uv lockfile; no API key is needed. The [earlier decision-model diagnostic](https://hk-775.github.io/practical-eval-lab/decision-models.md) and its [design review](https://hk-775.github.io/practical-eval-lab/decision-model-design-review.md) are retained as a separate experiment. Jev has not been evaluated. ## Evaluate your application Compare two named Python candidates without modifying the runner: ```bash uv run --locked eval-lab compare --suite classification \ --project examples/python-project.json --before app-v1 --after app-v2 \ --report reports/application.json --html reports/application.html ``` Expose those registrations in the webpage: ```bash uv run --locked eval-lab serve --project examples/python-project.json ``` A complete local HTTP application and endpoint configurations are included too. The [integration guide](https://hk-775.github.io/practical-eval-lab/integrations.md) covers Python, HTTP, optional OpenAI calls, environment-based credentials, and replacing the demonstration application. A project file is trusted executable configuration; the browser cannot register arbitrary code or endpoints. ## Save and compare evidence The webpage saves every run and comparison automatically. Reopen saved experiments, import JSON reports, compare compatible saved runs, and export standalone HTML. Reports include case-level outputs, grader checks, slices, gate decisions, candidate identity, fingerprints, candidate/grader timing, and token usage when supplied. ```bash uv run --locked eval-lab compare --suite rag \ --report reports/rag.json --html reports/rag.html uv run --locked eval-lab export reports/rag.json --html reports/rag.html uv run --locked eval-lab run --suite response_quality --candidate improved --trials 3 ``` Repeated trials expose variability on the same cases; they are not independent new test examples. No estimated costs or statistical significance are inferred. Source-checkout state lives in ignored `local/`: `profiles/`, profile `history/`, and saved `runs/`. Installed packages use a writable per-user directory. Override with `--state-dir PATH` or `EVAL_LAB_HOME`. The CLI uses bundled cases unless given `--config` or `--cases`; it never silently loads webpage edits. ## Block a regression in CI ```bash uv run --locked eval-lab compare --suite tool_calling --split holdout \ --fail-on-regression --critical-tag vocabulary --min-slice lookup=1 ``` This intentionally exits **1** despite the average improving from 50% to 80%: one lookup regresses, and critical/slice checks fail. Quality gates apply to the candidate after the change; execution failures in either candidate remain errors. | Exit | Meaning | |---:|---| | 0 | Quality gates pass and candidates executed successfully | | 1 | Quality or regression gate failed | | 2 | Invalid configuration, incompatible comparison, or file error | | 3 | At least one candidate execution failed | ## Install a built artifact ```bash uv build uv tool install ./dist/practical_eval_lab-0.2.0-py3-none-any.whl eval-lab serve ``` This installs an artifact you build locally. No registry package or GitHub Release is claimed. Wheels contain the webpage, datasets, provenance, and license notices. Wheel and source-distribution installs are exercised outside the checkout in CI. ## Verify and extend ```bash uv sync --locked uv run --locked pytest -q uv run --locked playwright install chromium uv run --locked python -m scripts.browser_check uv run --locked python -m scripts.package_check uv run --locked python -m scripts.reproduce_portfolio ``` CI covers Python 3.10 and 3.14 on Ubuntu and Python 3.13 on Ubuntu, Windows, and macOS, with Chromium checks on Ubuntu. Tests exercise grader failure paths, real loopback HTTP integration, report persistence/import, gates, packaging, and the six webpage flows. The optional OpenAI adapter is tested with a mock response; no paid live call has been used to substantiate the bundled results. ![Dataset, candidate, grader, report, and local tuning profile](https://hk-775.github.io/practical-eval-lab/architecture/pipeline.svg) [Contracts and extension points](https://hk-775.github.io/practical-eval-lab/contracts.md) · [Data provenance](https://hk-775.github.io/practical-eval-lab/data-provenance.md) · [Website build and hosting](https://hk-775.github.io/practical-eval-lab/hosting.md) · [Architecture page](https://hk-775.github.io/practical-eval-lab/architecture.html) · [Publication inventory](https://hk-775.github.io/practical-eval-lab/publication.md) · [Contributing](https://hk-775.github.io/practical-eval-lab/contributing.md) · [Support](https://hk-775.github.io/practical-eval-lab/support.md) · [Security](https://hk-775.github.io/practical-eval-lab/security.md) · [Community conduct](https://hk-775.github.io/practical-eval-lab/conduct.md) This is a small local toolkit. It does not provide a hosted service, distributed execution, general semantic grounding, a validated safety judge, or production certification. The walkthroughs state what each grader can and cannot establish. --- Source: [README.md](https://github.com/hk-775/practical-eval-lab/blob/main/README.md) Source SHA-256: `668ed2efef563d943281f22eceda167834bca2663a78ef466526a1e333febb3c` ===== FILE: learning-guide.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/learning-guide.md # Run → inspect → change → compare An eval is a repeatable experiment: a dataset, a candidate, a grader, and a decision rule. This lab makes each part visible. ## 1. Classification: read failures before improving rules Start the webpage with `uv run --locked python server.py`, choose Classification, and click **Compare candidates**. The baseline passes 14 of 20 development cases. The improved rules pass 20. Filter to **Changed results** and inspect the Spanish ticket and procurement requests. The improvement comes from a small vocabulary extension and checking business requests before hardware keywords. Now choose Holdout. The improved rules pass only 7 of 10. They miss unfamiliar vocabulary and another language. Development success does not prove broad language understanding. The original 15-case starter is preserved in root `cases.jsonl`: ```bash uv run --locked python eval.py --cases cases.jsonl ``` That original baseline still passes 13/15. ## 2. Extraction: separate formatting from meaning Choose Structured extraction. A response must include: ```json {"order_id":"A-104","quantity":3,"item":"keyboard","priority":"normal"} ``` The grader checks that the output is an object, fields have the right types, and each field matches the reference. `true` is not a valid quantity, even though Python treats booleans as a numeric subtype. JSON with the wrong quantity still fails. Inspect the **not urgent** case. The baseline sees `urgent` and assigns high priority. The improved rules recognize the negation and assign low priority. This suite defines “not urgent” and “no rush” as low priority. That convention is part of this example's task contract, not a universal business rule. Toggle **Allow extra fields** to change the schema rule. Required fields and their values remain checked. The current local candidates do not add extra fields, so this setting may leave their scores unchanged. The improved rules still miss `six` and `printers` in holdout. Adding cases is useful only when their reference answers reflect the actual task. ## 3. Tool calling: improvement can include regressions Choose Tool calling, then Holdout, then compare. The score improves from 50% to 80%, but one case regresses: “Could you locate A-104?” The baseline looks up any mentioned order number. This answers that case correctly but also acts on cancellation requests and illustrative IDs. The improved candidate requires a recognized status-query phrase. It rejects more inappropriate actions, but does not recognize “locate.” Filter to **Regressions**. The useful question is not only “did the average go up?” but also “which behavior did we lose?” The tools never contact real services: - `lookup_order(order_id)` reads three synthetic order statuses. - `create_ticket(category, summary)` returns a simulated outcome. - `no_action()` records that nothing should happen. There is no autonomous agent loop. This suite tests a single tool decision, its arguments, and the simulator's result. ## 4. Tune an eval in the webpage 1. Choose Development and select a case. 2. Edit its input, expected answer, or tags. For extraction and tools, the expected answer must be a JSON object matching the suite contract. 3. Apply the case to the draft. Running, saving, and exporting also apply pending case edits and validate them. 4. Adjust grading settings or the passing threshold. 5. Rerun both candidates. The comparison uses the same edited data for both. 6. Save tuning to keep it across reloads. Export a profile to share or use in CLI. Add case duplicates the selected example into a new `custom-N` case so you can adapt a valid reference. Remove case affects only the development draft. Saved overrides live in `local/profiles/`; earlier saved versions are in `local/history/`. Restore defaults archives the saved version and returns to the bundled cases. Unsaved drafts are not persisted across closing the page. Concurrent saves detect stale versions instead of overwriting another tab's work. Exported profiles can be imported back into the same suite. A profile contains all case text; use synthetic or otherwise shareable inputs. ```bash uv run --locked python -m eval_lab compare \ --suite extraction --config extraction-profile.json \ --report reports/my-comparison.json ``` The CLI does not silently load local webpage overrides. Pass `--config` explicitly. Holdout uses the bundled grader settings and 80% gate in the webpage. Your development draft remains available when you switch back. ## 5. Keep the experiment honest Changing the passing threshold changes only the gate. Changing labels, cases, or grading rules changes the test. Reports record hashes of the dataset and grader; comparisons reject mismatched configurations. Pass rate is the fraction of cases passing **every** enabled check. Slice scores show the numerator and denominator because these samples are tiny. Tags overlap; their totals should not be added together. The original three suites contain 90 synthetic teaching examples. Both rule candidates were developed for this lab. Holdout is a reserved teaching split in a public repository, not a secret benchmark or a guarantee against contamination. The guide discusses its failures openly. For a real product, collect a larger, representative dataset and reserve fresh examples before tuning. Continue with the [RAG](https://hk-775.github.io/practical-eval-lab/guides/rag.md), [response-quality](https://hk-775.github.io/practical-eval-lab/guides/response-quality.md), and [agent](https://hk-775.github.io/practical-eval-lab/guides/agent.md) walkthroughs. They separate evidence, human agreement, and workflow completion. Inspect saved runs, measurement details, and standalone HTML exports in the webpage. Do not use exact text equality as a general judge of open-ended quality. The RAG grader is explicitly extractive; the response-quality example uses real human preferences with documented limitations. --- Source: [docs/learning-guide.md](https://github.com/hk-775/practical-eval-lab/blob/main/docs/learning-guide.md) Source SHA-256: `ff452891eac567b3c9301347fc890b131ab162840b39d0975a1ef3e08cf523fc` ===== FILE: contracts.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/contracts.md # Data, candidates, grades, and reports ## Case A JSONL dataset contains 1–500 objects with exactly `id`, `input`, `expected`, and `tags`. IDs are unique nonempty strings up to 100 characters. Inputs fit within 20,000 serialized characters. Tags are up to 20 distinct nonempty strings of at most 60 characters. JSON numbers must be finite. The first three examples use string inputs. RAG, pairwise judging, and agents use structured objects. Each suite validates both its input and reference contract before a candidate runs. `eval_lab/data//{dev,holdout}.jsonl` contains concrete examples. RAG references must quote an existing current, trusted relevant document. ## Candidate and grade Built-ins are `baseline` and `improved`; `openai` is an optional CLI adapter. Trusted project registrations add named Python, HTTP, or OpenAI candidates. A candidate receives a deep copy of the input, never the reference. Return a JSON value or `CandidateOutput(output, usage, metadata)` as described in the [integration guide](https://hk-775.github.io/practical-eval-lab/integrations.md). A grader returns `passed`, nonempty `checks` with `name`, boolean `passed`, and human-readable `detail`, plus optional numeric `metrics`. A case passes only if every enabled check passes. Exceptions fail execution and are counted separately. Classification, extraction, and one-call tools are in `suites.py`; the new tasks are in `advanced.py`. To add a suite, register its metadata and settings, implement input/reference validation, a grader, and local candidates, and add separate dev and holdout files. Cover meaningful negative examples in tests. This is an explicit teaching catalog, not a dynamic untrusted-grader plugin system. ## Tuning profile, version 2 ```json { "schema_version": 2, "suite": "classification", "cases": [{"id":"example", "input":"My keyboard is broken", "expected":"Hardware", "tags":["critical"]}], "settings": {"normalize_labels": true}, "threshold": 0.8 } ``` Legacy three-field profiles (`cases`, `settings`, `threshold`) are accepted and normalized to version 2. Current profiles include suite identity to reject an accidental cross-suite import. Save revisions detect stale browser edits, and prior saved versions are archived before replacement or reset. ## Run report, version 2 Reports record suite, split, candidate identity, lab version, timestamp, dataset and grader hashes, threshold, settings, gate policy, results, slices, errors, trials, case count, and derived measurements. Repeated execution IDs append a trial suffix; `case_id` retains the original identity. Both candidate and grader latency are recorded, while token usage is present only when actually supplied. Candidate source/configuration fingerprints identify the local implementation or registration without copying secret environment values. Grader fingerprints cover the shared runner and grader source plus settings. They do not fingerprint your external service or prove which remote deployment answered; version that service and its candidate configuration explicitly. Derived check and measurement rates report their denominators. An execution error may have only an execution check, so a per-field or per-metric rate can describe fewer outputs than the entire dataset. Overall pass rate always includes errors. The report includes total errors and usage coverage to make this visible. Matched comparison requires identical schema, suite, split, dataset hash, grader hash, threshold, trial count, and gate policy, plus the same execution IDs. It reports per-execution improvements/regressions, delta, and gate results. A total score increase does not override a regression or critical-case gate. Reports import up to 20 MB. Validation checks structures, counts, fingerprints against embedded cases, trial completeness, slices, and gate consistency. Derived measurements are recomputed. This is integrity checking, not a digital signature or proof that a run happened. Imported evidence remains labeled, including when it is used in a new comparison. Version 1 reports need rerunning; profile migration is supported, but old report migration is intentionally not implied. ## Persistence and export The local state directory contains profiles, profile history, and UUID-named run JSON files. The UI lists the latest 200 runs; older files are retained. Export JSON to preserve machine-readable results and HTML for a standalone readable artifact. HTML escapes every output, contains no scripts, and loads no external assets. Both formats include data-license notices. Neither anonymizes case contents. Source checkouts use `local/`; installed applications use the operating system's per-user application-data directory. `--state-dir` or `EVAL_LAB_HOME` overrides it. Deleting an exported file does not delete saved server state. --- Source: [docs/contracts.md](https://github.com/hk-775/practical-eval-lab/blob/main/docs/contracts.md) Source SHA-256: `70bcd0e56d075616df816b7028c7a357411d95c0c602bba4e253504ac13b096e` ===== FILE: integrations.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/integrations.md # Bring your application A candidate receives only a case's `input`. References, tags, IDs, and human labels remain with the runner. Return a JSON-serializable output matching the selected suite. The runnable demos intentionally call a tiny rules application so the integration can be verified without a provider account. ## Python callable Start with the included complete example: ```bash uv run --locked eval-lab compare --suite classification \ --project examples/python-project.json --before app-v1 --after app-v2 uv run --locked eval-lab serve --project examples/python-project.json ``` Replace `eval_lab.demo_app:classify_ticket` in a trusted project file with your own installed `module:function`. Optional `options` become keyword arguments; `suites` limits which tasks the registration may run. Project schema version is 1. ```json { "schema_version": 1, "candidates": { "my-application": { "kind": "python", "target": "my_package:answer", "options": {"version": "v2"}, "suites": ["classification"] } } } ``` Install your application into the same environment. Source-only modules can be made importable through your normal Python packaging or PYTHONPATH workflow. No arbitrary filesystem import path is accepted from the webpage. For direct Python integration, no project file is needed: ```python from eval_lab.core import run_eval from eval_lab.integrations import CandidateOutput def evaluate_ticket(text): # Replace with your application call. Token usage is optional. return CandidateOutput("Other", usage={}) report = run_eval("classification", runner=evaluate_ticket, trials=2) ``` Use `CandidateOutput` only when supplying token counts or model metadata. A raw dict remains an ordinary candidate output; extraction objects are never confused with an adapter envelope. Supported usage keys are `input_tokens`, `output_tokens`, and `total_tokens`, with nonnegative integer values. Optional metadata preserves `model` and `response_id`. Reports never infer missing usage or billing. ## HTTP application In terminal one: ```bash uv run --locked python -m eval_lab.demo_app --port 8765 ``` In terminal two: ```bash uv run --locked eval-lab compare --suite classification \ --project examples/http-project.json --before endpoint-v1 --after endpoint-v2 uv run --locked eval-lab serve --project examples/http-project.json ``` The adapter sends a POST with `{"input": ...}` and expects `{"output": ...}`, optionally with `usage` and `metadata`. It has a configurable 1–120 second timeout (fractional positive seconds also work), a 2 MB response limit, and no retries or redirects. Remote endpoints require HTTPS; HTTP is allowed only for loopback hosts. URL credentials, query strings, and fragments are rejected. To authenticate, set `token_env` to an environment-variable name; its value is sent as a Bearer token and is not copied into the report. Candidate exceptions become failed executions with the exception type only. Inspect your application's own protected logs for detailed errors. Reports still contain inputs and outputs, so review them before sharing. ## Optional OpenAI judge Choose a model explicitly and set `OPENAI_API_KEY` in your environment or secret manager. A live run sends inputs to the provider and may incur costs. ```bash uv sync --locked --extra openai uv run --locked --extra openai eval-lab run --suite response_quality \ --candidate openai --model YOUR_MODEL --trials 3 \ --report reports/live-judge.json --html reports/live-judge.html ``` `YOUR_MODEL` is a placeholder, not a claimed available model. Use `--prompt FILE` to replace the suite prompt. The adapter uses the Responses API, records the requested model, prompt, returned model identifier, response ID, and token usage. Each call has a 30-second timeout and zero automatic retries. Output parsing is strict: Markdown fences around JSON fail the relevant JSON grader. To compare prompts or models, register two `kind: "openai"` candidates with explicit `model`, optional `prompt`, and `suites`, then use `compare --project` with `--before` and `--after`. The webpage refuses direct OpenAI registrations; run live-model experiments through the CLI. Registered HTTP/Python applications can themselves use paid services, so choose those registrations deliberately. The offline pairwise example measures heuristic agreement with real human labels. No paid model judge run or model performance result is bundled. ## Execution boundaries A project file is trusted code/network configuration, not an untrusted tuning profile. Loading a Python candidate executes its module. Functions run sequentially in-process; Python candidates have no forced timeout or cancellation. Keep runs small and use your application's own request limits. The browser disables controls while a run is active. Distributed workers, resumability, and live job cancellation are outside this portfolio's implemented scope. The server binds only to loopback and has no multi-user authentication. Do not expose it as a public application. Credentials never belong in project options, cases, profiles, output text, or source control. --- Source: [docs/integrations.md](https://github.com/hk-775/practical-eval-lab/blob/main/docs/integrations.md) Source SHA-256: `9dace76c2dafd2db18d89a4ca55f6d3da189250eb0d40820d83e63d7667ca3ee` ===== FILE: data-provenance.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/data-provenance.md # Dataset provenance The first three suites contain 90 synthetic cases retained from the original lab. RAG and multi-step agents add 24 synthetic cases. These 114 cases use invented support tickets, orders, policies, and tool results. Their references were authored as part of the teaching task, not collected from model users or production systems. Original project code and these synthetic cases use MIT-0. ## Human-preference calibration sample Response quality uses **12 existing human-preference pairs** from [Anthropic HH-RLHF](https://github.com/anthropics/hh-rlhf), revision `c72f5cee8eb7b4d2ea5617657f4430d5e333af07`, file `helpful-base/test.jsonl.gz`. The source describes `chosen`/`rejected` responses ranked through human preference collection. It does not supply this project's rubric scores; we do not invent those scores or call heuristic labels human. [Dataset description](https://github.com/anthropics/hh-rlhf/blob/c72f5cee8eb7b4d2ea5617657f4430d5e333af07/README.md) · [Upstream license](https://github.com/anthropics/hh-rlhf/blob/c72f5cee8eb7b4d2ea5617657f4430d5e333af07/LICENSE) · [Included MIT notice](https://hk-775.github.io/practical-eval-lab/downloads/anthropic-MIT.txt) · [Per-record manifest](https://hk-775.github.io/practical-eval-lab/downloads/response-quality-provenance.json) Selection: short, everyday-topic conversations were manually inspected for this small demonstration. Selection was purposeful, not random or representative. The local dev lines are 18, 68, 104, 137, 148, and 354. Local holdout lines are 535, 595, 603, 53, 384, and 397. Both local splits come from the upstream test file; “holdout” here means reserved from local tuning, not Anthropic's original train/test partition or a hidden benchmark. Transformation: split each `chosen` and `rejected` string at its last `\n\nAssistant:` marker; require identical preceding conversation; retain exact conversation and response substrings, including whitespace. Alternate the original human-preferred response between A and B within each local split. Only the reference contains the winner; the candidate receives conversation, A, and B. The manifest includes each original record's SHA-256 and 1-based source line. The original preference is not a certificate of correctness. For example, the circle-area pair contains a questionable preferred response; its original label is preserved and tagged `noisy-reference`. Review disagreements rather than assuming the human reference or automated judge is always right. The original larger dataset contains sensitive material; only these inspected pairs are bundled. These texts and reproductions in recorded reports retain Anthropic's MIT license, including attribution. JSON/HTML response-quality exports carry that notice too. When adapting or editing cases, retain applicable notices and document provenance for any new data. Changing a reference after examining results invalidates claims that the new score measures the same test. ## How to use these datasets responsibly as experiments The candidates were authored for this lab, and the walkthroughs openly discuss holdout failures. The data is educational and small. It does not establish general model quality, real-world safety, fairness, or resistance to contamination. For a real application, define the target population and risk categories, obtain properly licensed or consented examples, adjudicate ambiguous labels, and reserve fresh data before tuning. Repeated runs of the same case measure variability, not new coverage. The optional `examples/incidents/rag-profile.json` adds four explicitly fictional MIT-0 exercises outside the six-suite case counts. Incident documentation paraphrases public primary sources and links the original accounts; it does not redistribute full articles, customer transcripts, or the affected systems. --- Source: [docs/data-provenance.md](https://github.com/hk-775/practical-eval-lab/blob/main/docs/data-provenance.md) Source SHA-256: `a2bb696deca36d885148b270bf39adc43f13f272e9b00dbb3319d74628115202` ===== FILE: notices.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/notices.md # Third-party material The project's original code and synthetic teaching cases use [MIT-0](https://hk-775.github.io/practical-eval-lab/downloads/LICENSE.txt). The following data retains its own license; MIT-0 does not replace it. | Material | Source | License | Included notice | |---|---|---|---| | 12 human-preference pairs, their conversation context, and reproductions in example reports | Anthropic HH-RLHF, `helpful-base/test.jsonl.gz`, revision `c72f5cee8eb7b4d2ea5617657f4430d5e333af07` | MIT, copyright (c) 2022 Anthropic | [Full MIT notice](https://hk-775.github.io/practical-eval-lab/downloads/anthropic-MIT.txt) | The [per-record manifest](https://hk-775.github.io/practical-eval-lab/downloads/response-quality-provenance.json) records source line numbers, SHA-256 hashes, local splits, and A/B mappings. Text and human preference labels were retained; see the [data provenance explanation](https://hk-775.github.io/practical-eval-lab/data-provenance.md). The source repository's [license](https://github.com/anthropics/hh-rlhf/blob/c72f5cee8eb7b4d2ea5617657f4430d5e333af07/LICENSE) and [dataset description](https://github.com/anthropics/hh-rlhf/blob/c72f5cee8eb7b4d2ea5617657f4430d5e333af07/README.md) were checked when selecting this sample. Dependencies have their own licenses and are recorded in `uv.lock`. No external fonts, analytics scripts, image libraries, customer traces, or production data are bundled. --- Source: [THIRD_PARTY_NOTICES.md](https://github.com/hk-775/practical-eval-lab/blob/main/THIRD_PARTY_NOTICES.md) Source SHA-256: `ffed9461d73d03c1b9973c52d5f4f29cbf55ac8f24c3042b66a127ffce49d141` ===== FILE: architecture.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/architecture.md # The local evaluation loop Dataset inputs go to a named candidate. Reference answers go directly to the grader. Outputs, checks, gates, and metrics form a saved report. The CLI and local tuning webpage share this runner. The public Pages site serves static documents and recorded JSON/HTML reports. It has no execution backend, model credentials, private API, WebSockets, or cloud integration. [Rendered diagram](https://hk-775.github.io/practical-eval-lab/architecture/pipeline.svg) · [Editable draw.io source](https://hk-775.github.io/practical-eval-lab/architecture/pipeline.drawio) · [Contracts](https://hk-775.github.io/practical-eval-lab/contracts.md) · [Hosting](https://hk-775.github.io/practical-eval-lab/hosting.md). --- Source: [scripts/build_pages.py](https://github.com/hk-775/practical-eval-lab/blob/main/scripts/build_pages.py) Source SHA-256: `e6c887df7116f9b8a2bc4997a3cc26d15654cc301e835ddf1d1d2b6e0ec53ce2` ===== FILE: guides/classification.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/guides/classification.md # Classification: a score can hide the wrong failures **Task.** Route a support ticket to Hardware, Software, or Other. Procurement is Other even when the request names a device. Success is the correct label, with optional case/whitespace normalization. Extra explanations do not count as labels. ```bash uv run --locked eval-lab compare --suite classification --html reports/classification.html uv run --locked eval-lab compare --suite classification --split holdout ``` The development comparison is 14/20 → 20/20. The change gives procurement/general requests precedence over hardware keywords and adds a few device words in another language. Holdout improves from 4/10 to 7/10; unfamiliar vocabulary still fails. The holdout command exits 1 at the default 80% gate. **Why these measurements?** Accuracy answers how often the routing decision is right. The confusion matrix identifies which classes get mixed up. Macro-F1 gives each of the three labels equal weight; the implementation uses normalized labels for this diagnostic even when the exact-label pass check is enabled. Missing classes receive F1=0, so interpret small custom datasets carefully. Tag slices show specific language and business-rule failures with their denominators. **Limits.** Keyword rules do not understand arbitrary tickets. These small, authored examples cannot establish deployment accuracy or demographic fairness. A language slice is not a substitute for a representative multilingual dataset. **Try it.** Export the development profile, add a ticket where hardware is mentioned but the problem is a software driver, and decide the reference before running. Compare the per-class errors, not just the overall percentage. Then run your own application using [named candidates](https://hk-775.github.io/practical-eval-lab/integrations.md). [Recorded experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Dataset](https://github.com/hk-775/practical-eval-lab/tree/main/eval_lab/data/classification) --- Source: [examples/classification.md](https://github.com/hk-775/practical-eval-lab/blob/main/examples/classification.md) Source SHA-256: `4c61d766a6ec9c3fe2ded4e10a112d97ea489932127a7801319988223c91c43e` ===== FILE: guides/extraction.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/guides/extraction.md # Structured extraction: valid JSON can still be wrong **Task.** Extract `order_id`, positive integer `quantity`, singular `item`, and `priority` from an order request. Missing values are null; unspecified priority is normal. This lesson defines “not urgent” and “no rush” as low priority. ```bash uv run --locked eval-lab compare --suite extraction --html reports/extraction.html uv run --locked eval-lab compare --suite extraction --split holdout ``` Development improves 10/20 → 20/20; holdout improves 3/10 → 8/10. The change adds case-insensitive matching, selected number words, and negation handling. It still misses vocabulary such as `six` and `printers` in the holdout data. **Why these measurements?** JSON-object validity and schema checks distinguish formatting failure from incorrect values. Per-field pass rates identify whether quantity, item, priority, or order ID is responsible. `true` is rejected as a quantity even though Python booleans are integer subclasses. An incorrect value still fails when extra fields are allowed. **Limits.** Exact field equality assumes one agreed normalization. It does not measure whether a different wording means the same thing. The schema is intentionally small, and execution errors are reported separately from per-field measurements. **Try it.** Add a missing-quantity case and an ambiguous priority case. Agree the null/ambiguity convention first. Compare two prompts through the optional model adapter, preserving the same profile, and use multiple trials to inspect variability. [Recorded experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Dataset](https://github.com/hk-775/practical-eval-lab/tree/main/eval_lab/data/extraction) --- Source: [examples/extraction.md](https://github.com/hk-775/practical-eval-lab/blob/main/examples/extraction.md) Source SHA-256: `97851441f1d05bc543e60f7dc7c48873e559ee19a18e8914ddc4d1cfba36ba3b` ===== FILE: guides/tool-calling.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/guides/tool-calling.md # Tool calling: an average improvement can hide a regression **Task.** Select one simulated `lookup_order`, `create_ticket`, or `no_action` call. Grade the tool, arguments, and the side-effect-free simulator outcome. This suite is one tool decision; the [agent example](https://hk-775.github.io/practical-eval-lab/guides/agent.md) covers a workflow. ```bash uv run --locked eval-lab compare --suite tool_calling --split holdout uv run --locked eval-lab compare --suite tool_calling --split holdout \ --fail-on-regression --critical-tag vocabulary --min-slice lookup=1 ``` Development is 12/20 → 20/20. Holdout is 5/10 → 8/10, with four improvements and one regression. The second command intentionally exits 1. The baseline acts on any order number, including cancellation requests and sample IDs. The change requires a recognized lookup intent, but misses “Could you locate A-104?” **Why these measurements?** Tool selection alone misses wrong order IDs or ticket arguments. Simulated outcome checks make effects inspectable. A regression gate protects previously passing behavior, a critical tag requires every tagged execution to pass, and a slice minimum protects a subset. Gate tags must exist; a typo cannot silently disable a check. **Limits.** The simulator has three tools and no external effects. It cannot validate a real service's authorization, delivery, or idempotency behavior. Those need integration tests against an appropriate isolated service. **Try it.** Add “locate” support to a candidate and check whether it accidentally acts on a negated lookup. Keep the regression visible until the application change passes both positive and negative cases; changing the label is a different test. [Recorded experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Dataset](https://github.com/hk-775/practical-eval-lab/tree/main/eval_lab/data/tool_calling) --- Source: [examples/tool-calling.md](https://github.com/hk-775/practical-eval-lab/blob/main/examples/tool-calling.md) Source SHA-256: `d7370866b2c22f982102a7155e64e489ca183eadc0576ef1c0d91a54930149c0` ===== FILE: guides/rag.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/guides/rag.md # RAG: retrieval, answers, and evidence are different tests **Task.** Given a question and a small policy library, retrieve document IDs, quote one answer sentence, and cite it. Return null and empty lists when the library cannot answer. Each document has explicit `trusted` and `current` flags. ```bash uv run --locked eval-lab compare --suite rag --html reports/rag.html uv run --locked eval-lab compare --suite rag --split holdout ``` Development is 2/8 → 8/8. The baseline uses the first overlapping document and its first sentence. The change ranks lexical overlap using inverse document frequency, filters stale/untrusted documents, and chooses a matching sentence. Holdout stays at 1/4: paraphrases, multiple required sources, and conflicting authority remain unresolved. The holdout command is expected to fail the default gate. **Why these measurements?** Retrieval recall checks the labeled relevant IDs. Answer correctness uses exact reference equality. Citation checks require cited IDs to have been retrieved and be relevant. Evidence checks require the full answer to be an exact span in each current, trusted cited document. Unanswerable cases test abstention; a missing answer is not automatically a safe success. **Limits.** This is extractive RAG. Exact spans are a deliberately narrow evidence proxy, not semantic entailment or factual truth. A false source can still contain an exact matching sentence. Real systems need source-quality decisions, factuality review, semantic grounding, and suitable domain references. The trust flags are fixture inputs, not a learned source-trust classifier. Retrieval recall for an unanswerable case is defined here as 1 only when nothing is retrieved; it is not a universal retrieval convention. **Try it.** Replace a question with a paraphrase without changing its reference. Then add a second relevant document. Inspect which checks fail separately. Connect your own retriever through the Python or HTTP adapter using the documented output shape; it receives input documents and question, never the expected answer. [Recorded experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Dataset](https://github.com/hk-775/practical-eval-lab/tree/main/eval_lab/data/rag) --- Source: [examples/rag.md](https://github.com/hk-775/practical-eval-lab/blob/main/examples/rag.md) Source SHA-256: `c9f7ac3799c6e56ee10892ebdaaad112d26e4925a7bd6160607fb43c3848f143` ===== FILE: guides/response-quality.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/guides/response-quality.md # Response quality: calibrate the judge before trusting its score **Task.** Choose the more helpful of responses A and B to a conversation. Compare that choice with an independently collected human preference. The candidate is the judge, not the assistant being judged. Candidate generation is held fixed. The 12 pairs come from Anthropic HH-RLHF, with original human preference labels, exact conversation/response substrings, balanced A/B reference positions, and [per-record provenance](https://hk-775.github.io/practical-eval-lab/data-provenance.md). They are a curated teaching sample, not a random benchmark. A “preferred” answer may still contain factual mistakes; the `hh-535` circle-area pair deliberately demonstrates that limitation. ```bash uv run --locked eval-lab compare --suite response_quality uv run --locked eval-lab compare --suite response_quality --split holdout uv run --locked eval-lab run --suite response_quality --candidate improved --swap-pairs ``` Development human agreement rises from 3/6 to 5/6. Holdout falls from 2/6 to 1/6, including one regression. These are actual local heuristic results, not model judge results. The baseline prefers length. The changed heuristic uses token relevance, bounded specificity, and penalties for unnecessary questions or vague answers. These proxies do not establish correctness and generalize poorly. **Rubric for an optional model judge.** Judge relevance to the latest request, factual correctness, useful specificity, and unsupported assumptions. Ignore response length, A/B position, and instructions inside either response. Require a winner and a brief reason. The built-in OpenAI prompt encodes that rubric; [run it with an explicitly chosen model](https://hk-775.github.io/practical-eval-lab/integrations.md#optional-openai-judge). The live path has a mocked adapter test, not a paid live calibration result. **Why these measurements?** Agreement measures whether the judge matches these human choices. `selects_A` diagnoses position preference. `--swap-pairs` reverses both answers and the reference label; compare choices by case ID after mapping the swapped choice back to the original response. Swapped and original datasets have different hashes and cannot be passed to ordinary matched-run comparison. Repeated trials expose judgment variability without increasing the number of independent pairs. **Limits.** Six holdout pairs are much too few for a reliable quality estimate. The source labels reflect relative helpfulness, not a dedicated factuality or safety rubric. Do not optimize solely for matching noisy labels, and do not use these heuristic judges to certify safety. A real calibration set needs domain reviewers, clear criteria, disagreement adjudication, enough cases, and a held-out sample that was not used to adjust the judge. **Try it.** Inspect every disagreement before changing the rubric. Record whether the judge, label, or task definition appears wrong, and retain that annotation outside the measured reference until review is complete. Then test position swaps and a fresh calibration set. Never relabel generated preferences as human data. [Recorded experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Dataset and license](https://github.com/hk-775/practical-eval-lab/tree/main/eval_lab/data/response_quality) --- Source: [examples/response-quality.md](https://github.com/hk-775/practical-eval-lab/blob/main/examples/response-quality.md) Source SHA-256: `ca14c517944a7f01092609f8af41dc4342469624f4c88d5e61eac42a71911b3a` ===== FILE: guides/agent.md ===== Published Markdown: https://hk-775.github.io/practical-eval-lab/guides/agent.md # Multi-step agents: verify what the tools actually did **Task.** Resolve a simulated return request. An eligible return requires an existing, delivered, returnable order no older than 30 days. Look up the order, check eligibility, then create the return. Unsupported requests must invoke no tools. Retry transient failures within the call budget. Report missing orders or exhausted budgets as unavailable. ```bash uv run --locked eval-lab compare --suite agent --html reports/agent.html uv run --locked eval-lab compare --suite agent --split holdout --critical-tag critical ``` Development improves 4/8 → 8/8; holdout improves 1/4 → 4/4. The baseline follows the three-step workflow but stops at its first transient failure and mishandles unsupported requests. The changed candidate checks intent first and retries within the budget. The recorded comparisons show which failures change; every tool remains an in-memory simulation with no actual orders or money. **Why these measurements?** A trace records `{tool, arguments, result}` for each call. The grader independently replays calls through the simulator. It rejects unknown tools, wrong order IDs, skipped eligibility checks, duplicate returns, fabricated results, and excessive calls. Goal completion depends on replayed state, not the candidate's final claim. “Unavailable” requires evidence: an actual missing-order result or an exhausted retry budget. An empty trace cannot prove it. **Limits.** The candidate and grader share the simulator's tool semantics; a bug in that simulator is a shared assumption. Negative tests and explicit reference goals help, but a real integration needs independent service contracts and tests. The fixture's failure schedule is visible in input for reproducibility. It is not a hidden environment or a measurement of planning capability. General agents need richer traces, permissions, idempotency, and real environment outcome checks. **Try it.** Add a second transient failure and reduce the budget. Check both goal completion and termination. Register a Python candidate that runs a real agent against isolated test tools and adapts its trace to this contract. Do not connect write-capable production tools merely to run a teaching eval. [Recorded experiments](https://hk-775.github.io/practical-eval-lab/results.md) · [Dataset](https://github.com/hk-775/practical-eval-lab/tree/main/eval_lab/data/agent) --- Source: [examples/agent.md](https://github.com/hk-775/practical-eval-lab/blob/main/examples/agent.md) Source SHA-256: `334631e3d3c8039e9000662942273b2d41484ff0de932708cd4494a60b3f77c3`