Evidence confidence
moderate. A described comparison supports a bounded conclusion. Workload transfer and independent reproduction remain unresolved.
What was observed
Identify the responsible agent, step and error category in a failed run.
6,257 text traces from Who&When Pro. Jev answers three choice questions. The author uses the official scorer and compares with published paper baselines.
Baseline
GPT-5.4 results from the original paper, not a fresh matched API run.
Finding
Reported error-category macro-F1 was 23.7 for Jev and 15.3 for GPT-5.4 on the 100-point scale. Joint accuracy was 31.3% and 21.3%.
- Error-category macro-F1: Jev 23.7; paper GPT-5.4 baseline 15.3.
- All three labels correct: Jev 31.3%; paper GPT-5.4 baseline 21.3%.
- Reported Jev input-price estimate: $1.28 for 6,257 traces.
What the result does not establish
Failures are injected. Jev selects enumerated agents and steps while paper baselines generate them. Mode-confidence ECE is 0.287. No latency comparison is reported.
What we would test in System One
Try advisory trace triage with a fixed error taxonomy. A low joint success rate does not justify automatic repair or blame assignment.
This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.
Primary sources
Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.
Metric definitions and review method · Submit a correction or new result