Collinear AI Launches CWE-bench to Test Frontier Coding Agents on Defensive Cybersecurity Capabilities
PR Newswire
SAN FRANCISCO, Sept. 2, 2026
Held-out cybersecurity benchmark spans 54 weakness types; the leading agent passes less than 50% of tasks, and 18 remain unsolved.
SAN FRANCISCO, Sept. 2, 2026 /PRNewswire-PRWeb/ -- Collinear AI has launched CWE-bench, a held-out benchmark that tests whether frontier coding agents can defend real software vulnerabilities.
AI systems are now finding and exploiting security flaws across real infrastructure, largely on their own. If agents can discover and exploit cyber weaknesses, we need to know how well they can also find and repair them. CWE-bench accurately measures whether they can.
To ensure broad vulnerability coverage, CWE-bench is built around MITRE's CWE taxonomy. The benchmark of 100 agentic tasks currently spans 54 weakness types and all 10 OWASP Top 10 2025 categories, with the aim of further expanding coverage across MITRE's catalog.
The tasks are designed so that memorizing published fixes is not enough. In one example, three of four leading agents fixed the publicly documented token-revocation paths but missed a newly introduced path, earning zero credit. They recognized the known version of the vulnerability but failed to reason through how the weakness appeared elsewhere in the code.
"Cybersecurity is one of the toughest remaining hill climbs in coding. We built CWE-bench to make that climb faster with hard but fair environments that expose useful failures," said Nazneen Rajani, CEO of Collinear AI, who previously led post-training at Hugging Face. "CWE-bench applies all known vulnerabilities to known open-source code bases. All frontier models have the knowledge of these vulnerabilities, and these codebases are already in their training data but the leading model still scores less than 50%. The benchmark gives model builders a trusted signal about what to improve next."
Initial results include:
- Fable 5 leads the current leaderboard with a 47% pass@1 score at maximum reasoning.
- 18 of the 100 tasks remain unsolved by every model tested.
- No agent tested successfully repairs a majority of the benchmark.
- Performance varies across weakness types, exposing specific areas where defensive software reasoning still needs to improve.
- Models on the Pareto front of performance vs. cost include Fable 5, Gemini 3.8 Flash Cyber, and GPT-5.6 Sol — all at high reasoning.
Explore the leaderboard and methodology at https://cwe-bench.com/. Read more on how the benchmark was constructed at https://blog.collinear.ai/p/cwe-bench.
About Collinear AI
CWE-bench is produced by Collinear AI. Collinear builds agent evaluations, reinforcement-learning environments and verifier-graded training data for frontier models. The company is headquartered in Sunnyvale, California.
Media Contact
Richard Darnielle, Collinear AI, 1 (650) 772-6449, richard@collinear.ai, collinear.ai
View original content to download multimedia:https://www.prweb.com/releases/collinear-ai-launches-cwe-bench-to-test-frontier-coding-agents-on-defensive-cybersecurity-capabilities-302868204.html
SOURCE Collinear AI