Model-reviewed synthetic benchmark · October 2, 2026

Four routers.
One frozen test set.

Regex, Laya, Strands Decider, and GPT-4o Mini face the same requests. Compare routing accuracy, uncertainty, latency, and API cost—and inspect every decision.

Explore the evidence Download aggregate JSON Reproduce the run ↗
Test prompts
960
Scenario families
240
Development prompts
240
Task categories
6

What this measures: selecting a task category, including recognizing requests that need reasoning. A downstream LLM would still generate the answer. These are synthetic English prompts with AI-reviewed labels, not production traffic or human annotations.

Measured results

Accuracy comes with uncertainty.

One primary pass per router. Intervals resample whole scenario families, keeping correlated prompt variants together.

Regex / heuristic

49.79%

46.9%–52.7% · 95% family interval

Median latency: 0.34 ms

API cost / 1,000: $0

Laya

66.25%

63.3%–69.2% · 95% family interval

Median latency: 37.63 ms

API cost / 1,000: $0

Strands Decider

91.35%

89.2%–93.3% · 95% family interval

Median latency: 164.18 ms

API cost / 1,000: $0

GPT-4o Mini

91.25%

88.9%–93.5% · 95% family interval

Median latency: 736.22 ms

API cost / 1,000: $0.051203

RouterAccuracy / 95% intervalMacro F1 Median / p95 latencyAPI cost / 1,000Hardware / electricityErrors
Regex / heuristic49.79% / 46.9%–52.7%0.5070.34 / 1.40 ms$0Unmeasured0
Laya66.25% / 63.3%–69.2%0.66237.63 / 46.53 ms$0Unmeasured0
Strands Decider91.35% / 89.2%–93.3%0.914164.18 / 207.52 ms$0Unmeasured0
GPT-4o Mini91.25% / 88.9%–93.5%0.911736.22 / 966.04 ms$0.051203Unmeasured0

Measured host: Apple M4 Pro, 48 GiB RAM; MPS for local models. Laya: float32; Strands: bfloat16 with a reference causal-convolution kernel. GPT-4o Mini latency includes its network round trip. Hardware/electricity is unmeasured; provider infrastructure is included in API prices.

Paired comparisons

Is the difference clear on this set?

Positive values favor the candidate. The intervals below adjust for all six comparisons. An interval crossing zero does not distinguish that pair at this confidence level.

Candidate vs baselineAccuracy difference Adjusted intervalSeparated on this set?
Laya vs Regex / heuristic+16.46 pp+11.35 to +21.67 ppYes
Strands Decider vs Regex / heuristic+41.56 pp+37.08 to +46.15 ppYes
GPT-4o Mini vs Regex / heuristic+41.46 pp+36.56 to +46.35 ppYes
Strands Decider vs Laya+25.10 pp+20.83 to +29.27 ppYes
GPT-4o Mini vs Laya+25.00 pp+20.52 to +29.48 ppYes
GPT-4o Mini vs Strands Decider-0.10 pp-4.06 to +3.65 ppNo

10,000 class-stratified, paired family bootstrap samples. Bonferroni family-wise 95% intervals. These describe the synthetic sampling scheme, not real-world representativeness.

Task-level evidence

Look past the aggregate.

Each category has 160 test prompts in 40 families. “Reasoning” means recognizing a reasoning request; it does not score the quality of an eventual reasoning answer.

TaskRegex / heuristicLayaStrandsGPT-4o Mini
coding63.7%81.9%98.8%91.9%
reasoning21.9%48.1%79.4%96.9%
creative writing36.2%90.0%95.0%98.1%
summarization63.1%53.1%95.6%98.1%
math17.5%46.2%95.6%93.8%
general96.2%78.1%83.8%68.8%
Repeatability: 120 test prompts, three repetitions each

The panel reuses 30 deterministically selected test families. It is reported separately and adds nothing to the primary accuracy sample size.

  • Regex / heuristic: 0/120 prompts changed predicted label.
  • Laya: 0/120 prompts changed predicted label.
  • Strands Decider: 0/120 prompts changed predicted label.
  • GPT-4o Mini: 0/120 prompts changed predicted label.

Inspect the evidence

Read the prompts. Check the routes.

Browse the frozen test set with expected categories and all four predictions. Everything below loads from a public static file; this page makes no inference calls.

Loading the frozen test cases…

Download test prompts and predictions · Full dataset, development split, audits and manifest ↗

Methodology and boundaries

Built to be inspected and repeated.

Author → audit → freeze

GPT-4.1 authored 1,200 prompts across ten domains. Claude Sonnet 4.6 reviewed labels without seeing expected categories. Disputes were resolved before scoring, and revision logs remain public. Agreement is a check, not proof of label truth.

Keep families together

Each family has explicit, implicit, keyword-collision, and quoted-distraction variants. All four stay in one split. Routers and their instructions were held fixed; neither labels nor audit notes enter model requests.

Separate quality and cost

Failed requests count as incorrect. Model loading and three warmups are excluded from primary latency. API costs use reported usage; local hardware and electricity remain unmeasured.

Keep the scope visible

This tests short English prompts. It does not establish production performance, multilingual quality, long-context behavior, calibrated confidence, or downstream answer quality. Synthetic-author and reviewer biases remain.

History: September heuristic / LLM / hybrid benchmark · Original 60-prompt four-router pilot. These use different datasets and should be read as separate experiments.