Held-out comparison
Hybrid wins the routing tradeoff
Local rules remained sub-millisecond and free. GPT-4o Mini improved accuracy, while the hybrid reached the highest score by escalating only the 28.3% of prompts the heuristic marked uncertain.
| Strategy | Accuracy | Macro F1 | p50 latency | p95 latency | LLM calls | Cost / 1K | Errors |
|---|---|---|---|---|---|---|---|
| Improved heuristic | 86.7% | 0.875 | 0.307 ms | 0.896 ms | 0 / 60 | $0 | 0 |
| LLM router | 91.7% | 0.926 | 657.7 ms | 844.0 ms | 60 / 60 | $0.045082 | 1 |
| Hybrid · 0.3 | 95.0% | 0.950 | 1.457 ms | 851.3 ms | 17 / 60 | $0.011590 | 0 |
Regex improvement
38.3% → 86.7% on the frozen test split
Only development-set misses informed the generalized intent families. After the rules reached 100% on development, they were locked before final per-case inspection.
Frozen before tuning
The 120 generated prompts were split deterministically into 60 development and 60 final cases. Corpus and split hashes are published for verification.
Generalized rules, not prompt IDs
The classifier added reusable families for requested artifacts, evidence weighing, compact outputs, quantitative operations, and multilingual intent.
Regression score kept separate
The earlier 98.6% result on the known 72-prompt corpus remains published, but it is labeled as a tuned regression result rather than held-out evidence.
Scenario accuracy
Where the strategies differ
Small scenario counts are diagnostic, not statistically conclusive. The complete held-out case output is published with the repository.
| Scenario | Cases | Heuristic | LLM | Hybrid |
|---|---|---|---|---|
| Clear | 12 | 100.0% | 91.7% | 100.0% |
| Implicit | 24 | 83.3% | 91.7% | 95.8% |
| Keyword collision | 12 | 83.3% | 83.3% | 83.3% |
| Multilingual | 12 | 83.3% | 100.0% | 100.0% |
Interpretation
What this experiment supports
Keep the deterministic first pass
The local classifier delivered 86.7% held-out accuracy with sub-millisecond p95 latency and no direct router spend.
Escalate uncertainty, not every prompt
The hybrid reached 95.0% accuracy while invoking the LLM on 17 of 60 prompts, about one quarter of the full-router cost.
Count operational failures
The LLM-only run produced one malformed classification after consuming its completion budget. Errors are scored as wrong routes, not silently removed.
Methodology
Reproducible, with the caveats attached
Every prompt is evaluated against the same six-label taxonomy. Router cache reuse is disabled with a per-request nonce, and direct router token cost is calculated from recorded usage. The corpus, split, hashes, and case-level decisions are committed.
# Reproduce the frozen held-out comparison
uv run --locked axon-benchmark-routing \
--corpus docs/benchmarks/autorouting-final-test-2026-09-16.jsonl \
--strategy heuristic \
--strategy llm \
--strategy hybrid \
--hybrid-threshold 0.3 \
--llm-base-url http://127.0.0.1:8001/v1 \
--llm-model gpt-4o-mini \
--llm-api-key-env AXON_BENCHMARK_API_KEY \
--llm-input-cost-per-million 0.15 \
--llm-output-cost-per-million 0.60
- The 60 final cases were excluded from rule tuning. No classifier changes were made after final per-case inspection.
- The corpus is synthetic and human-reviewed, not a substitute for labeled customer traffic from the target deployment.
- One repetition captures a reproducible snapshot, not a confidence interval. Provider latency and stochastic classification can vary.
- The supplied token prices were benchmark inputs, not a promise of current provider pricing. Recheck prices before financial decisions.
- The benchmark measures routing classification only. It does not claim that a correct label always selects the best downstream model.