# Six evaluation walkthroughs

- [Classify a support ticket](https://hk-775.github.io/practical-eval-lab/guides/classification.md): A high overall score can hide failures in one language or business rule.
- [Extract an order](https://hk-775.github.io/practical-eval-lab/guides/extraction.md): Valid JSON is only the first check. The values must also match the request.
- [Choose and call a tool](https://hk-775.github.io/practical-eval-lab/guides/tool-calling.md): Check the tool, its arguments, and what actually happens when the simulator executes it.
- [Retrieve and answer with evidence](https://hk-775.github.io/practical-eval-lab/guides/rag.md): Retrieval, answer correctness, and supporting citations can fail independently. Exact evidence checks do not measure general semantic truth.
- [Calibrate a pairwise judge](https://hk-775.github.io/practical-eval-lab/guides/response-quality.md): Human agreement is not factual correctness. Inspect disagreements, noisy labels, and changes when A/B positions are swapped.
- [Evaluate a multi-step agent](https://hk-775.github.io/practical-eval-lab/guides/agent.md): Replay the trace to verify state transitions. A confident final answer cannot replace successful tools, authorization, or a call budget.

---
Source: [scripts/build_pages.py](https://github.com/hk-775/practical-eval-lab/blob/main/scripts/build_pages.py)

Source SHA-256: `e6c887df7116f9b8a2bc4997a3cc26d15654cc301e835ddf1d1d2b6e0ec53ce2`
