All evidence records

reported · synthetic test

Separate questions for separate hazards

Explicit checks helped; calibration and latency limits remain.

AnshChoudhary; external project author; not rerun here.

Evidence confidence

low. This source suggests a useful experiment but does not establish a reliable benefit in a real agent workflow.

What was observed

Flag hazards in proposed agent tool calls before execution.

600 authored records, 320 held out; full hazard battery compared with a broad dangerousness question.

Baseline

Single generic hazard question.

Finding

Decomposed questions sharply reduced false blocks on hard negatives in this corpus. Calibration and latency targets still failed.

  • Hard-negative block rate: full battery 0%; generic question 39.2%.
  • Reported ECE 0.156; p95 added latency 595ms.

What the result does not establish

Author-designed synthetic attacks are not evidence of adversarial security in production. Approval friction must be measured separately.

What we would test in System One

Use independent, explicit checks as advisory signals. Keep authorization, execution safeguards and human approval outside Jev.

This recommendation is our interpretation of the study. Related research does not establish the quality of every Engine recipe.

Try a related workflow

Primary sources

Read this record in System One Bench. Source commits are pinned where available. Review dates describe our inspection, not the original run date.

Metric definitions and review method · Submit a correction or new result