Grok matches a cheaper AI on security tests, at nearly 65 times the cost
Artificial Analysis has launched a new Cyber Index designed to measure how well AI models find and fix software vulnerabilities, while showing how much each evaluation costs to run.
The index combines three cybersecurity evaluations: CWE-Bench-AA, DeepsecBench-AA and CyberGym-E2E-AA. Together, they cover vulnerability discovery, validation and remediation.
The launch also establishes the Cyber Index Alliance, with Collinear AI, IBM, NVIDIA and Vercel as founding partners. Collinear AI and Vercel contributed benchmark work, while IBM and NVIDIA provided expert input, according to Artificial Analysis.
What counts as a successful fix?
Finding a vulnerability is only part of the task. In CWE-Bench-AA, an AI agent receives a real open-source repository and an instruction describing the area of concern, but not the vulnerability’s exact location. It must audit the code, identify the weakness and patch it. This comes as a broader move toward giving agents autonomy to execute tasks.
A fix counts as successful only when a programmatic verifier confirms that the vulnerability can no longer be triggered and legitimate functionality still works.
There is no partial credit or language-model judge deciding whether the patch looks good. Artificial Analysis reports results as pass@1, meaning the percentage of 120 held-out tasks solved on the first attempt. Tasks run in an isolated sandbox without internet access.
This creates an important distinction: removing vulnerable code is not enough if the patch breaks legitimate functionality. Similarly, closing one route to a vulnerability while leaving a related weakness open does not receive credit.
Artificial Analysis says partial fixes were the main failure mode in CWE-Bench-AA. Excluding refusals and timeouts, 55% of failed attempts fixed the primary issue while leaving a related weakness open.

Three tests, not one
The three evaluations measure different defensive tasks.
CWE-Bench-AA tests vulnerability identification and remediation across 120 held-out tasks covering the OWASP Top 10 2025 categories and several programming languages.
DeepsecBench-AA focuses on vulnerability discovery in open-source application code, with results compared against expert-verified findings.
CyberGym-E2E-AA tests whether models can discover, reproduce and patch memory-safety vulnerabilities in C and C++ projects. The launch version contains 131 tasks.
Each evaluation contributes equally to the overall Cyber Index score. The index does not ask models to develop working exploits.
Performance comes with a cost
The first results show why cost is worth considering alongside the score.
Artificial Analysis reports that Grok 4.7 (xhigh) and MiMo-V2.6-Pro both score 56 while the cost per task is $11.67 and $0.18 respectively. GPT-6 Luna (max) scores 53 and the cost per task is $0.12.
These are Artificial Analysis’ evaluation costs, not estimates of production deployment costs. They are calculated from the tokens used during the evaluations and available model pricing.
The comparison is therefore useful for understanding the trade-off between benchmark performance and evaluation cost, but it does not capture infrastructure, human review, monitoring or failure-handling costs.
What the Cyber Index shows
Artificial Analysis also reports safety blocks separately. When a model or provider declines a task on safety grounds, that task receives zero credit, although the refusal rate is reported separately.
The launch index also has clear limits. It does not cover areas such as incident response, writing new code without introducing vulnerabilities or testing targets without source-code access.
For security teams, the useful takeaway is therefore not simply which model has the highest score. The benchmark provides a way to compare defensive capability, task coverage and cost under a defined testing methodology.
Artificial Analysis says it plans to expand the index. For now, its results should be read as evidence of how models performed on these specific source-code-based security tasks, rather than as a complete measure of an AI system’s ability to handle cybersecurity work.
The post Grok matches a cheaper AI on security tests, at nearly 65 times the cost appeared first on Search Engine Watch.