# Practical Eval Lab > Six runnable AI evaluation examples, transparent graders, and a local tuning webpage. This website displays recorded offline results. Candidate execution and tuning run locally. Five suites use synthetic cases; response quality uses a small attributed Anthropic human-preference sample. Bundled candidates are local rules. Results demonstrate evaluation methods, not production model quality or safety certification. Created by [Harleen Kaur](https://hk-775.github.io/hk-775/). [Source repository](https://github.com/hk-775/practical-eval-lab) ยท [Local setup](https://hk-775.github.io/practical-eval-lab/getting-started.md). ## Walkthroughs - [Classify a support ticket](https://hk-775.github.io/practical-eval-lab/guides/classification.md): A high overall score can hide failures in one language or business rule. - [Extract an order](https://hk-775.github.io/practical-eval-lab/guides/extraction.md): Valid JSON is only the first check. The values must also match the request. - [Choose and call a tool](https://hk-775.github.io/practical-eval-lab/guides/tool-calling.md): Check the tool, its arguments, and what actually happens when the simulator executes it. - [Retrieve and answer with evidence](https://hk-775.github.io/practical-eval-lab/guides/rag.md): Retrieval, answer correctness, and supporting citations can fail independently. Exact evidence checks do not measure general semantic truth. - [Calibrate a pairwise judge](https://hk-775.github.io/practical-eval-lab/guides/response-quality.md): Human agreement is not factual correctness. Inspect disagreements, noisy labels, and changes when A/B positions are swapped. - [Evaluate a multi-step agent](https://hk-775.github.io/practical-eval-lab/guides/agent.md): Replay the trace to verify state transitions. A confident final answer cannot replace successful tools, authorization, or a call budget. ## Setup, contracts, and evidence - [Local setup](https://hk-775.github.io/practical-eval-lab/getting-started.md): Install, run, and tune the toolkit. - [Evaluation fundamentals](https://hk-775.github.io/practical-eval-lab/learning-guide.md): How to choose cases, graders, and gates. - [Application adapters](https://hk-775.github.io/practical-eval-lab/integrations.md): Connect Python or HTTP applications. - [Contracts](https://hk-775.github.io/practical-eval-lab/contracts.md): Input/output, grading, reporting, and extension boundaries. - [Recorded evidence](https://hk-775.github.io/practical-eval-lab/results.md): Twelve comparisons with JSON and standalone HTML. - [Decision model benchmark](https://hk-775.github.io/practical-eval-lab/decision-models.md): Pinned Strands and Laya comparison; Jev optional. - [Decision model results](https://hk-775.github.io/practical-eval-lab/decision-model-results.md): Local measurements on a frozen synthetic fixture; raw evidence and limits. - [Decision model design review](https://hk-775.github.io/practical-eval-lab/decision-model-design-review.md): Why the first diagnostic cannot support a general model ranking. - [Executed support workflow](https://hk-775.github.io/practical-eval-lab/tool-workflow.md): Tool and argument selection, clarification, actual simulator outcomes, and baselines. - [Support-workflow protocol](https://hk-775.github.io/practical-eval-lab/tool-workflow-protocol.md): Frozen partitions, calibration, outcome definitions, and limitations. - [Support-workflow results](https://hk-775.github.io/practical-eval-lab/tool-workflow-results.md): Recorded tool outcomes and replayable traces; model stages compared with rules. - [Public incidents](https://hk-775.github.io/practical-eval-lab/incidents.md): Documented incidents, sources, and proposed checks. - [Provenance](https://hk-775.github.io/practical-eval-lab/data-provenance.md): Synthetic cases and attributed human preferences. - [Architecture](https://hk-775.github.io/practical-eval-lab/architecture.md): Local runtime and static hosting boundaries. - [Combined context](https://hk-775.github.io/practical-eval-lab/agent-context.txt): A bounded bundle of the principal guides with source fingerprints. - [Source manifest](https://hk-775.github.io/practical-eval-lab/discovery.json): Document URLs and SHA-256 fingerprints. ## Optional - [Coding-agent instructions](https://hk-775.github.io/practical-eval-lab/coding-agents.md): Setup, tests, repository map, and conventions. - [Licenses](https://hk-775.github.io/practical-eval-lab/notices.md): MIT-0 code and retained third-party MIT notice. - [Security](https://hk-775.github.io/practical-eval-lab/security.md): Reporting and local/public trust boundaries.