All evidence records

reported · independent experiment

Model routing: test whether Jev adds value

No-Jev ablation matched the hybrid result.

TokenTrim; external author; not independently reproduced here.

Evidence confidence

moderate. A described comparison supports a bounded conclusion. Workload transfer and independent reproduction remain unresolved.

What was observed

Select a model for a user query.

Author reports offline scoring on 5,835 held-out LLMRouterBench queries across 13 models; answers are precomputed.

Baseline

Best fixed model and identical retrieval router without Jev.

Finding

Retrieval routing improved the accuracy-cost tradeoff, but the no-Jev ablation matched it. This does not establish that the Jev difficulty signal caused savings.

  • Hybrid accuracy 62.4%; best fixed 60.3%.
  • No-Jev ablation accuracy 62.4%.

What the result does not establish

Cached downstream answers; no live end-to-end answer latency. Reported cost uses benchmark assumptions.

What we would test in System One

Before adding a model router, compare a fixed default and a no-Jev retrieval policy. Ship a router only if the added signal earns its overhead.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Try a related workflow

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result