System One Bench

What can your agent
do with Jev?

Start with a useful task and a human example. Then read the evidence, see how confident we are and inspect the test that would change our view.

We include firsthand Reddit and X posts, builder demos, official examples and measured studies. Potential value and evidence confidence are editorial ratings. They are not model probabilities or promised savings.

Read the rating method

A report can be useful before it becomes a benchmark

A person trying Jev in a long coding session tells us which problem matters. A builder demo shows a possible implementation. A controlled comparison helps decide whether it works better. Bench keeps those kinds of evidence distinct and includes conflicting accounts.

Testing the current recipes

Our first check of 15 new recipes matched the authored labels in 20 of 21 synthetic cases. One recovery question exposed ambiguity. We retained every result and did not raise any broader confidence rating.

Read the recipe check and its limits

Each idea has a test to run

The protocols define cases, baselines, measurements and a product decision. They are planned work, not completed results. We will update the claim and its confidence when a finding supports, narrows or rejects it.

Read the test backlog · Contribute a use case or firsthand report · Try a current recipe