Announcing the Artificial Analysis Cyber Index and the Artificial Analysis Cyber Index Alliance, a new standard for evaluating AI models on enterprise cyber defense The Artificial Analysis Cyber Index Alliance brings together industry partners to create a new standard for evaluating how AI models perform on enterprise cyber defense tasks. The Alliance launches alongside the Artificial Analysis Cyber Index, which combines three partner-contributed and open benchmarks to evaluate how well agents find and fix vulnerabilities. As models demonstrate increasingly advanced cyber offense capabilities, it becomes more relevant for AI labs and companies alike to understand how models perform on cyber defense tasks and which perform best. We’re announcing the Cyber Index Alliance today with Collinear AI, IBM, NVIDIA, and Vercel as launch partners. Benchmarks in the Artificial Analysis Cyber Index: ➤ CWE-Bench-AA, from Collinear AI covers auditing and patching: 120 held-out tasks spanning all ten OWASP Top 10 (2025) categories, across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust. ➤ DeepsecBench-AA, from Vercel, isolates discovery: Given a codebase and a budget, the agent needs to find every vulnerability present, and is scored against a golden set of findings from human security reviewers. Real findings are rewarded and benign code flagged as vulnerable is penalized. ➤ CyberGym-E2E-AA, from Berkeley RDI, runs end to end: Find the memory-safety bug, write a proof-of-concept that triggers the crash, then patch it so the crash no longer reproduces. Key results: ➤ Grok 4.7 (xhigh) and MiMo-V2.6-Pro lead the Cyber Index scoring 56, followed by GPT-6 Luna (max, 53), GLM-5.3-Flash (50) and Muse Spark 1.3 (xhigh, 44). ➤ Safety refusals hold back several frontier models: GPT-6 Sol (max), GPT-6 Astra (max), Claude Opus 5.5 (max with fallback), Claude Fable 5.1 (max with fallback) and Gemini 3.8 Flash (high) decline tasks representing 32-38% of the Cyber Index on safety grounds. Despite frontier agentic coding capabilities, they trail the leaders by 19 to 31 points. Most of the gap comes from CyberGym-E2E-AA, where GPT-6 Sol and GPT-6 Astra refuse every task, Claude Opus 5.5 refuses 98% and Claude Fable 5.1 refuses 99%. For more details, read the article: https://lnkd.in/g8t68-mD
Glad to see a defense-side index. Offense benchmarks get the headlines, but teams picking a model for vuln triage need to know whether its fixes hold up and don't break the build. Will the index report false positive rates alongside fix rates?
A useful move towards evidence-led cyber defence. For CIOs, the value will come from seeing how benchmark results translate into prioritised remediation, reliable fixes and measurable reduction in exposure across live engineering workflows.
A thoughtful example of technology moving from promise to practical impact. The CIO priority is to pair innovation with clear accountability, measurable outcomes, and a resilient operating model. That balance will help organisations scale with confidence.
The off-target passes deserve a place beside the headline score. CyberGym counts a fixed, reproducible crash even when it is not the vulnerability the task was built around; your article says 31% of passes land there. That is useful bug fixing, but a security team still has an open ticket. I would report target-vulnerability closure separately.
A valuable development for measuring AI’s ability to support cyber defense. Standardized benchmarks that test vulnerability discovery, remediation, and end-to-end security tasks can provide useful insight into how AI systems perform in real-world security workflows
#Grok 4.7 and #MiMo-V2.6-Pro lead the #Cyber Index at 56. #CyberGym-E2E scores a patch after a proof-of-concept that triggers the crash. GPT-6 #Sol and #GPT-6 #Astra refuse every task on that bench; Claude #Opus 5.5 refuses 98% and #Claude #Fable 5.1 refuses 99. That refuse is a halt the model already obeys. NVIDIA sits on the same alliance as a launch partner. What an Accident Is Worth is the Berkeley-led offense eval that kept going. #ExploitGym ran inside a lab harness; roughly 1,200 agents exchanged more than 70,000 messages, around 700 reached a production platform, and about 17,600 attacker actions followed over four and a half days. On 27 June the on-call staff advised that stopping the evaluation was unnecessary: https://seldondance.substack.com/p/what-an-accident-is-worth Who May Keep Going is the halt a company has to obey. NVIDIAs Eos dropped from roughly 4mw to 3 inside a minute on a utility signal, and the facility has honored more than 200 such signals: https://seldondance.substack.com/p/who-may-keep-going Saurabh wrote that teams should decline to be hostage to guardrails when defending. A 100% refuse on the crash-and-patch bench is already a stop. The stop that was missing sat on the offense run.
We are glad to partner on this important benchmark. Now you know which models to use for Cyber Defense. Don't be hostage to guardrails when defending yourself.
Artificial Analysis is the most beautiful and useful thing in AI this 2026 🎉🍾👏 My congratulations to all the team!