Check a claim against tool output
Ask whether the evidence supports a statement before an agent repeats it as fact.
Read the guide Guide 02 · Coding and research agentsShortlist context for an agent
Evaluate a small set of retrieved passages without losing the source material.
Read the guide Guide 03 · Agents with several connected toolsChoose the next available tool
Give an agent a tool recommendation from a list it can actually use.
Read the guide Guide 04 · Anyone creating a System One patternWrite a question Jev can evaluate
Define the evidence, possible answers and review policy before adjusting a threshold.
Read the guide Guide 05 · Individual coding and chat agent usersMeasure whether offloading helps
Compare completed tasks with and without System One, including review and fallback costs.
Read the guideExamples you can inspect
The JSON examples use current service contracts. They are teaching examples, not new benchmark results. Every guide links to related studies and explains what still needs testing.
These guides also live in the public System One Bench repository. The website uses a reviewed source revision so a changed report cannot silently change the claim.