# Practical Eval Lab

> Six runnable AI evaluation examples, transparent graders, and a local tuning webpage.

This website displays recorded offline results. Candidate execution and tuning run locally. Five suites use synthetic cases; response quality uses a small attributed Anthropic human-preference sample. Bundled candidates are local rules. Results demonstrate evaluation methods, not production model quality or safety certification.

Created by [Harleen Kaur](https://hk-775.github.io/hk-775/). [Source repository](https://github.com/hk-775/practical-eval-lab) · [Local setup](https://hk-775.github.io/practical-eval-lab/getting-started.md).

## Walkthroughs
- [Classify a support ticket](https://hk-775.github.io/practical-eval-lab/guides/classification.md): A high overall score can hide failures in one language or business rule.
- [Extract an order](https://hk-775.github.io/practical-eval-lab/guides/extraction.md): Valid JSON is only the first check. The values must also match the request.
- [Choose and call a tool](https://hk-775.github.io/practical-eval-lab/guides/tool-calling.md): Check the tool, its arguments, and what actually happens when the simulator executes it.
- [Retrieve and answer with evidence](https://hk-775.github.io/practical-eval-lab/guides/rag.md): Retrieval, answer correctness, and supporting citations can fail independently. Exact evidence checks do not measure general semantic truth.
- [Calibrate a pairwise judge](https://hk-775.github.io/practical-eval-lab/guides/response-quality.md): Human agreement is not factual correctness. Inspect disagreements, noisy labels, and changes when A/B positions are swapped.
- [Evaluate a multi-step agent](https://hk-775.github.io/practical-eval-lab/guides/agent.md): Replay the trace to verify state transitions. A confident final answer cannot replace successful tools, authorization, or a call budget.

[Recorded JSON and HTML evidence](https://hk-775.github.io/practical-eval-lab/results.md) · [Data provenance and attribution](https://hk-775.github.io/practical-eval-lab/data-provenance.md).

---
Source: [scripts/build_pages.py](https://github.com/hk-775/practical-eval-lab/blob/main/scripts/build_pages.py)

Source SHA-256: `e6c887df7116f9b8a2bc4997a3cc26d15654cc301e835ddf1d1d2b6e0ec53ce2`
