Regex / heuristic
46.9%–52.7% · 95% family interval
Median latency: 0.34 ms
API cost / 1,000: $0
Model-reviewed synthetic benchmark · October 2, 2026
Regex, Laya, Strands Decider, and GPT-4o Mini face the same requests. Compare routing accuracy, uncertainty, latency, and API cost—and inspect every decision.
What this measures: selecting a task category, including recognizing requests that need reasoning. A downstream LLM would still generate the answer. These are synthetic English prompts with AI-reviewed labels, not production traffic or human annotations.
Measured results
One primary pass per router. Intervals resample whole scenario families, keeping correlated prompt variants together.
46.9%–52.7% · 95% family interval
Median latency: 0.34 ms
API cost / 1,000: $0
63.3%–69.2% · 95% family interval
Median latency: 37.63 ms
API cost / 1,000: $0
89.2%–93.3% · 95% family interval
Median latency: 164.18 ms
API cost / 1,000: $0
88.9%–93.5% · 95% family interval
Median latency: 736.22 ms
API cost / 1,000: $0.051203
| Router | Accuracy / 95% interval | Macro F1 | Median / p95 latency | API cost / 1,000 | Hardware / electricity | Errors |
|---|---|---|---|---|---|---|
| Regex / heuristic | 49.79% / 46.9%–52.7% | 0.507 | 0.34 / 1.40 ms | $0 | Unmeasured | 0 |
| Laya | 66.25% / 63.3%–69.2% | 0.662 | 37.63 / 46.53 ms | $0 | Unmeasured | 0 |
| Strands Decider | 91.35% / 89.2%–93.3% | 0.914 | 164.18 / 207.52 ms | $0 | Unmeasured | 0 |
| GPT-4o Mini | 91.25% / 88.9%–93.5% | 0.911 | 736.22 / 966.04 ms | $0.051203 | Unmeasured | 0 |
Measured host: Apple M4 Pro, 48 GiB RAM; MPS for local models. Laya: float32; Strands: bfloat16 with a reference causal-convolution kernel. GPT-4o Mini latency includes its network round trip. Hardware/electricity is unmeasured; provider infrastructure is included in API prices.
Paired comparisons
Positive values favor the candidate. The intervals below adjust for all six comparisons. An interval crossing zero does not distinguish that pair at this confidence level.
| Candidate vs baseline | Accuracy difference | Adjusted interval | Separated on this set? |
|---|---|---|---|
| Laya vs Regex / heuristic | +16.46 pp | +11.35 to +21.67 pp | Yes |
| Strands Decider vs Regex / heuristic | +41.56 pp | +37.08 to +46.15 pp | Yes |
| GPT-4o Mini vs Regex / heuristic | +41.46 pp | +36.56 to +46.35 pp | Yes |
| Strands Decider vs Laya | +25.10 pp | +20.83 to +29.27 pp | Yes |
| GPT-4o Mini vs Laya | +25.00 pp | +20.52 to +29.48 pp | Yes |
| GPT-4o Mini vs Strands Decider | -0.10 pp | -4.06 to +3.65 pp | No |
10,000 class-stratified, paired family bootstrap samples. Bonferroni family-wise 95% intervals. These describe the synthetic sampling scheme, not real-world representativeness.
Task-level evidence
Each category has 160 test prompts in 40 families. “Reasoning” means recognizing a reasoning request; it does not score the quality of an eventual reasoning answer.
| Task | Regex / heuristic | Laya | Strands | GPT-4o Mini |
|---|---|---|---|---|
| coding | 63.7% | 81.9% | 98.8% | 91.9% |
| reasoning | 21.9% | 48.1% | 79.4% | 96.9% |
| creative writing | 36.2% | 90.0% | 95.0% | 98.1% |
| summarization | 63.1% | 53.1% | 95.6% | 98.1% |
| math | 17.5% | 46.2% | 95.6% | 93.8% |
| general | 96.2% | 78.1% | 83.8% | 68.8% |
The panel reuses 30 deterministically selected test families. It is reported separately and adds nothing to the primary accuracy sample size.
Inspect the evidence
Browse the frozen test set with expected categories and all four predictions. Everything below loads from a public static file; this page makes no inference calls.
Loading the frozen test cases…
Download test prompts and predictions · Full dataset, development split, audits and manifest ↗
Methodology and boundaries
GPT-4.1 authored 1,200 prompts across ten domains. Claude Sonnet 4.6 reviewed labels without seeing expected categories. Disputes were resolved before scoring, and revision logs remain public. Agreement is a check, not proof of label truth.
Each family has explicit, implicit, keyword-collision, and quoted-distraction variants. All four stay in one split. Routers and their instructions were held fixed; neither labels nor audit notes enter model requests.
Failed requests count as incorrect. Model loading and three warmups are excluded from primary latency. API costs use reported usage; local hardware and electricity remain unmeasured.
This tests short English prompts. It does not establish production performance, multilingual quality, long-context behavior, calibrated confidence, or downstream answer quality. Synthetic-author and reviewer biases remain.
History: September heuristic / LLM / hybrid benchmark · Original 60-prompt four-router pilot. These use different datasets and should be read as separate experiments.