Practical Eval Lab · saved experiment

agent · baseline → improved

4/4 passed (100.0%) · change +75.0% · 0 regressions

holdout · 2026-10-01T20:51:14.728010+00:00 · 1 trial(s)

overall: PASS

Teaching data; results describe these cases only. Candidate kind: local_rules. Imported reports are user-supplied evidence, not authenticated measurements. Review inputs and outputs before sharing.

Measurements and reproducibility
{
  "checks": {
    "schema": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "budget": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "trace": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "goal": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "final_state": {
      "passed": 4,
      "total": 4,
      "rate": 1.0
    },
    "termination": {
      "passed": 1,
      "total": 1,
      "rate": 1.0
    }
  },
  "measurements": {
    "tool_calls": {
      "mean": 2.5,
      "count": 4
    }
  },
  "candidate_latency_ms": {
    "median": 0.0085,
    "p95": 0.017
  },
  "usage": {},
  "usage_coverage": 0,
  "trial_scores": [
    1
  ],
  "unstable_cases": 0
}

Dataset 7733052626a5c754339e5f185de98f09ca88e018db0954cd8b2073e5cfc2c9c5
Grader 8ed15f453523fc8778082785bbd46add899da329d8e3f6ab69a7edb5c6e5b3f8

{
  "name": "improved",
  "kind": "local_rules",
  "model": null,
  "source_hash": "7665f5dde4f1c1a9ea008dbca97d9c101a4af0510720b6f00d5da49ad6c12e20"
}

Case results

agent-boundary · PASS

retry, boundary

Input and reference
{
  "request": "return",
  "order": {
    "id": "R-boundary",
    "delivered": true,
    "age_days": 30,
    "returnable": true,
    "exists": true
  },
  "failures": {
    "create_return": 1
  },
  "budget": 5
}

Reference

{
  "final": "returned"
}

Before

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "found": true
      }
    },
    {
      "tool": "check_eligibility",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "eligible": true
      }
    },
    {
      "tool": "create_return",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "error": "temporary"
      }
    }
  ],
  "final": "unavailable"
}

Output

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "found": true
      }
    },
    {
      "tool": "check_eligibility",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "eligible": true
      }
    },
    {
      "tool": "create_return",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "error": "temporary"
      }
    },
    {
      "tool": "create_return",
      "arguments": {
        "order_id": "R-boundary"
      },
      "result": {
        "return_created": true
      }
    }
  ],
  "final": "returned"
}
Grader checks
PASS schema: Return a bounded trace and final status.
PASS budget: At most 5 tool calls.
PASS trace: Every recorded result matches independent tool replay.
PASS goal: The replayed outcome must satisfy the reference goal.
PASS final_state: Reported final state must agree with replay.

agent-nonreturnable · PASS

eligibility, critical

Input and reference
{
  "request": "return",
  "order": {
    "id": "R-nonreturnable",
    "delivered": true,
    "age_days": 5,
    "returnable": false,
    "exists": true
  },
  "failures": {},
  "budget": 5
}

Reference

{
  "final": "declined"
}

Before

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-nonreturnable"
      },
      "result": {
        "found": true
      }
    },
    {
      "tool": "check_eligibility",
      "arguments": {
        "order_id": "R-nonreturnable"
      },
      "result": {
        "eligible": false
      }
    }
  ],
  "final": "declined"
}

Output

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-nonreturnable"
      },
      "result": {
        "found": true
      }
    },
    {
      "tool": "check_eligibility",
      "arguments": {
        "order_id": "R-nonreturnable"
      },
      "result": {
        "eligible": false
      }
    }
  ],
  "final": "declined"
}
Grader checks
PASS schema: Return a bounded trace and final status.
PASS budget: At most 5 tool calls.
PASS trace: Every recorded result matches independent tool replay.
PASS goal: The replayed outcome must satisfy the reference goal.
PASS final_state: Reported final state must agree with replay.

agent-purchase · PASS

unsupported, critical

Input and reference
{
  "request": "purchase",
  "order": {
    "id": "R-purchase",
    "delivered": true,
    "age_days": 5,
    "returnable": true,
    "exists": true
  },
  "failures": {},
  "budget": 5
}

Reference

{
  "final": "no_action"
}

Before

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-purchase"
      },
      "result": {
        "found": true
      }
    },
    {
      "tool": "check_eligibility",
      "arguments": {
        "order_id": "R-purchase"
      },
      "result": {
        "eligible": true
      }
    },
    {
      "tool": "create_return",
      "arguments": {
        "order_id": "R-purchase"
      },
      "result": {
        "error": "rejected"
      }
    }
  ],
  "final": "unavailable"
}

Output

{
  "trace": [],
  "final": "no_action"
}
Grader checks
PASS schema: Return a bounded trace and final status.
PASS budget: At most 5 tool calls.
PASS trace: Every recorded result matches independent tool replay.
PASS goal: The replayed outcome must satisfy the reference goal.
PASS final_state: Reported final state must agree with replay.

agent-exhausted · PASS

budget, retry

Input and reference
{
  "request": "return",
  "order": {
    "id": "R-exhausted",
    "delivered": true,
    "age_days": 5,
    "returnable": true,
    "exists": true
  },
  "failures": {
    "lookup_order": 2,
    "check_eligibility": 2
  },
  "budget": 4
}

Reference

{
  "final": "unavailable"
}

Before

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-exhausted"
      },
      "result": {
        "error": "temporary"
      }
    }
  ],
  "final": "unavailable"
}

Output

{
  "trace": [
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-exhausted"
      },
      "result": {
        "error": "temporary"
      }
    },
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-exhausted"
      },
      "result": {
        "error": "temporary"
      }
    },
    {
      "tool": "lookup_order",
      "arguments": {
        "order_id": "R-exhausted"
      },
      "result": {
        "found": true
      }
    },
    {
      "tool": "check_eligibility",
      "arguments": {
        "order_id": "R-exhausted"
      },
      "result": {
        "error": "temporary"
      }
    }
  ],
  "final": "unavailable"
}
Grader checks
PASS schema: Return a bounded trace and final status.
PASS budget: At most 4 tool calls.
PASS trace: Every recorded result matches independent tool replay.
PASS goal: The replayed outcome must satisfy the reference goal.
PASS termination: Unavailable requires a missing-order result or an exhausted retry budget.
PASS final_state: Reported final state must agree with replay.
Data attribution and license
Bundled cases are synthetic teaching data under MIT-0. Custom cases may have different provenance.