All evidence records

reported · independent experiment

Choosing tools from real MCP inventories

Better tool prediction can still take longer.

BillionsBobby/JevRouter; external project author; not rerun here.

Evidence confidence

moderate. A described comparison supports a bounded conclusion. Workload transfer and independent reproduction remain unresolved.

What was observed

Predict the first five tool calls for Toolathlon tasks.

Ten tasks; inventories from nine live MCP servers. Serial and decomposed Jev routing compared with DeepSeek V4.1 Flash.

Baseline

DeepSeek tool predictions; serial versus decomposed Jev.

Finding

Serial Jev was faster in this report; decomposition increased position-wise matches but also increased latency.

  • Position-wise hits: serial Jev 38%, decomposed 44%, DeepSeek 24%.
  • Per-task latency: 1.58s, 10.6s and 8.65s respectively.

What the result does not establish

Only ten tasks and predicted sequences; this is not demonstrated task completion. Provider, prompts and mode affect the tradeoff.

What we would test in System One

Offer a next-tool recommendation over explicit candidates. Do not market a predicted multi-step sequence as a reliable autonomous plan.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Try a related workflow

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result