An evaluation workshop

Tool calling: an average improvement can hide a regression

Read the example, inspect its evidence, then reproduce it in the local lab.

Task. Select one simulated lookup_order, create_ticket, or no_action call. Grade the tool, arguments, and the side-effect-free simulator outcome. This suite is one tool decision; the agent example covers a workflow.

uv run --locked eval-lab compare --suite tool_calling --split holdout
uv run --locked eval-lab compare --suite tool_calling --split holdout \
  --fail-on-regression --critical-tag vocabulary --min-slice lookup=1

Development is 12/20 → 20/20. Holdout is 5/10 → 8/10, with four improvements and one regression. The second command intentionally exits 1. The baseline acts on any order number, including cancellation requests and sample IDs. The change requires a recognized lookup intent, but misses “Could you locate A-104?”

Why these measurements? Tool selection alone misses wrong order IDs or ticket arguments. Simulated outcome checks make effects inspectable. A regression gate protects previously passing behavior, a critical tag requires every tagged execution to pass, and a slice minimum protects a subset. Gate tags must exist; a typo cannot silently disable a check.

Limits. The simulator has three tools and no external effects. It cannot validate a real service's authorization, delivery, or idempotency behavior. Those need integration tests against an appropriate isolated service.

Try it. Add “locate” support to a candidate and check whether it accidentally acts on a negated lookup. Keep the regression visible until the application change passes both positive and negative cases; changing the label is a different test.

Recorded experiments · Dataset