Auto-routing benchmark · Frozen held-out snapshot

How much should routing intelligence cost?

AxonLLM compares a local keyword-and-regex classifier, an LLM classifier, and a confidence-gated hybrid on a frozen held-out corpus. The benchmark isolates the routing decision so generation latency and downstream model spend cannot hide the router's own cost.

Read this result correctly: the heuristic was improved using only the 60-prompt development split, where it moved from 35.0% to 100.0%. The rules were locked before final per-case inspection and scored 86.7% on the separate 60-prompt test split.

Hybrid wins the routing tradeoff

Local rules remained sub-millisecond and free. GPT-4o Mini improved accuracy, while the hybrid reached the highest score by escalating only the 28.3% of prompts the heuristic marked uncertain.

Read methodology
Improved heuristic Local only
86.7%
Accuracy · 52 / 60
Macro F10.875
p50 / p950.307 / 0.896 ms
LLM call rate0%
Router cost / 1K$0
LLM router Full escalation
91.7%
Accuracy · 55 / 60 · 1 parse error
Macro F10.926
p50 / p95657.7 / 844.0 ms
LLM call rate100%
Router cost / 1K$0.045082
Hybrid · 0.3 Best held out
95.0%
Accuracy · 57 / 60 · 0 errors
Macro F10.950
p50 / p951.457 / 851.3 ms
LLM call rate28.3%
Router cost / 1K$0.011590
Strategy Accuracy Macro F1 p50 latency p95 latency LLM calls Cost / 1K Errors
Improved heuristic86.7%0.8750.307 ms0.896 ms0 / 60$00
LLM router91.7%0.926657.7 ms844.0 ms60 / 60$0.0450821
Hybrid · 0.395.0%0.9501.457 ms851.3 ms17 / 60$0.0115900

38.3% → 86.7% on the frozen test split

Only development-set misses informed the generalized intent families. After the rules reached 100% on development, they were locked before final per-case inspection.

+48.3 pp
Improvement on the same frozen 60-prompt test split
Original · test
38.3%
Improved · dev
100.0%
Improved · test
86.7%

Frozen before tuning

The 120 generated prompts were split deterministically into 60 development and 60 final cases. Corpus and split hashes are published for verification.

Generalized rules, not prompt IDs

The classifier added reusable families for requested artifacts, evidence weighing, compact outputs, quantitative operations, and multilingual intent.

Regression score kept separate

The earlier 98.6% result on the known 72-prompt corpus remains published, but it is labeled as a tuned regression result rather than held-out evidence.

Where the strategies differ

Small scenario counts are diagnostic, not statistically conclusive. The complete held-out case output is published with the repository.

ScenarioCasesHeuristicLLMHybrid
Clear12100.0%91.7%100.0%
Implicit2483.3%91.7%95.8%
Keyword collision1283.3%83.3%83.3%
Multilingual1283.3%100.0%100.0%

What this experiment supports

01

Keep the deterministic first pass

The local classifier delivered 86.7% held-out accuracy with sub-millisecond p95 latency and no direct router spend.

02

Escalate uncertainty, not every prompt

The hybrid reached 95.0% accuracy while invoking the LLM on 17 of 60 prompts, about one quarter of the full-router cost.

03

Count operational failures

The LLM-only run produced one malformed classification after consuming its completion budget. Errors are scored as wrong routes, not silently removed.

Reproducible, with the caveats attached

Every prompt is evaluated against the same six-label taxonomy. Router cache reuse is disabled with a per-request nonce, and direct router token cost is calculated from recorded usage. The corpus, split, hashes, and case-level decisions are committed.

Corpus
60 dev + 60 test
Test labels
6 × 10 balanced
Repetitions
1 snapshot
Concurrency
1 sequential
Corpus generator
Claude Sonnet
Label review
Manual + GPT-4.1
Router LLM
GPT-4o Mini
Generation
Excluded
# Reproduce the frozen held-out comparison
uv run --locked axon-benchmark-routing \
  --corpus docs/benchmarks/autorouting-final-test-2026-09-16.jsonl \
  --strategy heuristic \
  --strategy llm \
  --strategy hybrid \
  --hybrid-threshold 0.3 \
  --llm-base-url http://127.0.0.1:8001/v1 \
  --llm-model gpt-4o-mini \
  --llm-api-key-env AXON_BENCHMARK_API_KEY \
  --llm-input-cost-per-million 0.15 \
  --llm-output-cost-per-million 0.60
  • The 60 final cases were excluded from rule tuning. No classifier changes were made after final per-case inspection.
  • The corpus is synthetic and human-reviewed, not a substitute for labeled customer traffic from the target deployment.
  • One repetition captures a reproducible snapshot, not a confidence interval. Provider latency and stochastic classification can vary.
  • The supplied token prices were benchmark inputs, not a promise of current provider pricing. Recheck prices before financial decisions.
  • The benchmark measures routing classification only. It does not claim that a correct label always selects the best downstream model.