All evidence records

reported · vendor benchmark

Typed evaluation versus chat-model wrappers

A reason to test typed decisions, not a savings guarantee.

TypeSafe AI, the Jev model vendor; not independently reproduced here.

Evidence confidence

low. This source suggests a useful experiment but does not establish a reliable benefit in a real agent workflow.

What was observed

Bounded decisions inside four application workflows.

Vendor launch comparison of typed Jev evaluations and chat-model decision wrappers.

Baseline

Vendor-selected chat-model baselines.

Finding

The launch report supports evaluating typed decisions as a separate operation. Its headline speed and cost ratios are vendor-specific workload results.

  • Vendor advertises up to 193.6× speed and 444.6× cost improvements in its evaluation; not System One Engine results.

What the result does not establish

Vendor-authored comparison, selected tasks and wrappers. Reported ratios do not predict a coding agent's total cost or time. Pricing and Gateway promotions change.

What we would test in System One

Batch independent questions over shared evidence. Measure host tool overhead and final outcomes before promising savings.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result