GPT-4.5 solves SQL injection 70% of the time; Claude Sonnet 4.6 hits budget limits before breaching—guardrails work, but inconsistently.
June 5, 2026
Summary
Security teams need empirical data on LLM attack surface before deploying agents with data access. This benchmark reveals which models leak user data and which ones stop themselves.
Why it matters
Security teams need empirical data on LLM attack surface before deploying agents with data access. This benchmark reveals which models leak user data and which ones stop themselves.
Implementation verdict
Replaces hand-waving about LLM safety with actual exploit metrics. Requires building a vulnerable test app matching your threat model. Worth running now if you're shipping agents with sensitive context—the variance between models is stark.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.