Measure agent reliability by executing tasks 20 times against isolated backends and checking terminal state changes, not parsing trajectories—reveals 67% of clean-looking failures have wrong field values or missing side effects.
Summary
Single-attempt benchmarks hide consistency gaps: Claude Opus 5.5 retains only 71% of its pass@1 score across 20 repeats. You need to know whether your agent will execute stateful workflows correctly every time, not just once.
Why it matters
Single-attempt benchmarks hide consistency gaps: Claude Opus 5.5 retains only 71% of its pass@1 score across 20 repeats. You need to know whether your agent will execute stateful workflows correctly every time, not just once.
Implementation verdict
Replaces pass@1-only leaderboard thinking with pass@20 and observed 20/20 metrics. Requires isolated MCP tool sessions and executable state checks on real backend objects. Available now via Hugging Face—worth running on your own workflows before production deployment to surface silent failures current evals miss.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.