Analysis
Does AI help cyber attackers or defenders more?
Between 2024 and 2026, AI systems went from finding a single new bug to, by Anthropic's account, thousands of high-severity vulnerabilities in every major operating system and browser.[1][2] The same capabilities have been used offensively: Anthropic reported a 2025 espionage campaign in which AI carried out 80-90% of the work, so the open question is whether defenders or attackers gain more.[3]
The question
AI is cutting the cost of finding software vulnerabilities: DARPA’s 2025 challenge put it at about $152 per task.[4] The debate is whether that helps defenders, who must fix everything, more than attackers, who need only one way in.
What we know
- Defensive results. In November 2024 Google’s Big Sleep agent found an exploitable memory-safety bug in SQLite before it reached a release. Google called it the first public example of an AI agent doing so in widely used software.[1][5] At DARPA’s AI Cyber Challenge final in August 2025, the systems found 54 synthetic vulnerabilities, patched 43 and found 18 real ones. They averaged 45 minutes per patch at about $152 per task.[6][4] DARPA said the systems would be released as open source.[7]
- Frontier capability. In April 2026 anthropic said its unreleased Claude Mythos Preview had found thousands of high-severity vulnerabilities, including some in every major operating system and web browser.[2] Anthropic said AI models can now beat all but the most skilled humans at finding and exploiting vulnerabilities, and warned that such capabilities may soon spread.[8] See claude-mythos.
- Defensive deployment. Through Project Glasswing, Anthropic gave Mythos Preview to AWS, Apple, Cisco, CrowdStrike, Google, Microsoft, the Linux Foundation and others. It committed up to $100 million in usage credits and $4 million in donations to open-source security groups.[9][10] On 6 October 2026 it expanded a tiered Cyber Verification Program that gives vetted security professionals reduced blocking.[11]
- Offensive use. Anthropic reported in November 2025 that a group it assessed with high confidence to be Chinese state-sponsored used Claude Code against roughly thirty targets in September 2025, succeeding in a small number of cases.[12] It said AI performed 80-90% of the campaign, and called this the first documented large-scale cyberattack executed without substantial human intervention.[3] The model sometimes hallucinated credentials or overstated its findings.[13]
Three views
Defenders gain more. Finding and fixing bugs at about $152 a task changes the economics for open-source maintainers, who have long been under-resourced.[4] If the strongest models reach defenders first, as Glasswing intends, flaws can be fixed before attackers find them.[9] Patching is also measurable. DARPA’s finalists fixed most of what they found.[6]
Attackers gain more. An attacker needs one working exploit, while a defender must patch every system, and patches take time to deploy. The 2025 campaign shows AI can already do most of an intrusion’s work at machine speed.[3] Anthropic itself expects these capabilities to spread beyond careful developers.[8]
It depends on access. The balance may be set less by the technology than by who can use it. Labs are experimenting with gated access for verified defenders and tighter safeguards for everyone else.[11] That works only as long as comparable models are not freely available. Meanwhile real attacks still suffer from model errors.[13]
Links to post-quantum migration
The two frontiers interact. Post-quantum migration means rewriting cryptographic code across almost every system, and new code brings new bugs.[14] AI tools that audit code could reduce that risk, or help attackers exploit it. Both trends also strain the same security teams. That pressure is one reason governments emphasise inventories and staged timelines; see migration timeline debate.
What to watch
Likely over the next year, with moderate confidence: more reports of AI-assisted intrusions from AI labs and security vendors, and more open-source projects running AI audits. A key unknown is whether models with Mythos-level capability become widely available without safeguards. If they do, the “attackers gain more” view becomes more plausible. Independent measurements of how quickly AI-found bugs get patched would also help settle the debate. Today most figures come from the companies building the tools.[2][6]
Competing views
Defenders gain more
Automated systems can find and patch flaws at scale and low cost, and the most capable tools are being steered to software maintainers first.[6][4][9][10]
Questions readers ask
Can AI really find unknown software vulnerabilities?
Yes. Google's Big Sleep found an exploitable SQLite bug in 2024, DARPA's AIxCC finalists found 18 real vulnerabilities in 2025, and Anthropic said in 2026 that its Mythos Preview model had found thousands of high-severity flaws.[1][6][2]
Have attackers used AI to run cyberattacks?
Anthropic reported in November 2025 that a group it assessed as Chinese state-sponsored used Claude Code to attempt intrusions into about thirty organizations, with AI performing 80-90% of the campaign. This is Anthropic's own assessment.[12][3]
How cheap is AI vulnerability finding?
In DARPA's AIxCC final, teams spent about $152 per competition task and submitted patches in an average of 45 minutes.[4]
Are fully autonomous AI cyberattacks possible yet?
Not reliably, according to Anthropic. In the 2025 campaign the model sometimes hallucinated credentials or overstated its findings, which Anthropic said remains an obstacle to fully autonomous attacks.[13]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
In November 2024 Google's Big Sleep AI agent was reported to have found a previously unknown exploitable memory-safety bug in SQLite, which Google called the first public example of an AI agent doing so in widely used real-world software. confirmedas of 2024-11-01
- From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code · Google Project Zero · 2024-11-01 (retrieved 2026-10-10)
- [2]
Anthropic said Claude Mythos Preview had found thousands of high-severity vulnerabilities, including some in every major operating system and web browser. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [3]
Anthropic said the AI performed 80-90% of that campaign, with humans intervening at perhaps 4-6 critical decision points, and called it the first documented large-scale cyberattack executed without substantial human intervention. confirmedas of 2025-11-13
- Disrupting an AI-orchestrated cyber espionage campaign · Anthropic · 2025-11-13 (retrieved 2026-10-10)
- Disrupting an AI-orchestrated cyber espionage campaign · Anthropic · 2025-11-13 (retrieved 2026-10-10)
- [4]
AIxCC finalists analyzed more than 54 million lines of code, submitted patches in an average of 45 minutes and spent about $152 per competition task. confirmedas of 2025-08-08
- AI Cyber Challenge marks pivotal inflection point for cyber defense · DARPA · 2025-08-08 (retrieved 2026-10-10)
- AI Cyber Challenge marks pivotal inflection point for cyber defense · DARPA · 2025-08-08 (retrieved 2026-10-10)
- [5]
The SQLite bug was reported in early October 2024, fixed the same day and never reached an official release. confirmedas of 2024-11-01
- From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code · Google Project Zero · 2024-11-01 (retrieved 2026-10-10)
- [6]
In the AIxCC final, competing systems found 54 of the synthetic vulnerabilities across 63 challenges, patched 43 of them, and also discovered 18 real, non-synthetic vulnerabilities. confirmedas of 2025-08-08
- AI Cyber Challenge marks pivotal inflection point for cyber defense · DARPA · 2025-08-08 (retrieved 2026-10-10)
- AI Cyber Challenge marks pivotal inflection point for cyber defense · DARPA · 2025-08-08 (retrieved 2026-10-10)
- [7]
DARPA said the finalists' AI systems were being made available as open source for broad adoption. confirmedas of 2025-08-08
- AI Cyber Challenge marks pivotal inflection point for cyber defense · DARPA · 2025-08-08 (retrieved 2026-10-10)
- [8]
Anthropic said AI models had reached a level of coding capability at which they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities, and warned that such capabilities may soon proliferate. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [9]
On 7 April 2026 Anthropic announced Project Glasswing with partners including AWS, Apple, Cisco, CrowdStrike, Google, Microsoft and the Linux Foundation, to use its unreleased Claude Mythos Preview model to find and fix vulnerabilities in critical software. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [10]
Anthropic committed up to $100 million in Mythos Preview usage credits and $4 million in direct donations to open-source security organizations. confirmedas of 2026-04-07
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [11]
On 6 October 2026 Anthropic expanded its Cyber Verification Program into three access tiers that give qualifying security professionals advanced cyber capabilities with reduced blocking classifiers, moving Project Glasswing members into the top tier. confirmedas of 2026-10-06
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- Project Glasswing: Securing critical software for the AI era · Anthropic · 2026-04-07 (retrieved 2026-10-10)
- [12]
Anthropic reported in November 2025 that a group it assessed with high confidence to be Chinese state-sponsored used its Claude Code tool to attempt intrusions into roughly thirty organizations in September 2025, succeeding in a small number of cases. confirmedas of 2025-11-13
- Disrupting an AI-orchestrated cyber espionage campaign · Anthropic · 2025-11-13 (retrieved 2026-10-10)
- [13]
Anthropic reported that the model sometimes hallucinated credentials or overstated what it had found, which it said remains an obstacle to fully autonomous cyberattacks. confirmedas of 2025-11-13
- Disrupting an AI-orchestrated cyber espionage campaign · Anthropic · 2025-11-13 (retrieved 2026-10-10)
- [14]
A large-scale quantum computer would make insecure the public-key systems based on integer factorization, such as RSA, and those based on the discrete logarithm problem, which includes elliptic-curve cryptography. confirmedas of 2026-10-10
- NIST IR 8105: Report on Post-Quantum Cryptography · NIST · 2016-04-01 (retrieved 2026-10-10)
- NIST IR 8105: Report on Post-Quantum Cryptography · NIST · 2016-04-01 (retrieved 2026-10-10)
Revision history (2)
- Page created.
- Linked the Anthropic entity page.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Does AI help cyber attackers or defenders more?." ContentLora, updated Oct 10, 2026. https://contentlora.com/analysis/ai-cybersecurity-debate
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerPost-quantum cryptography and security in 2026: a crash courseA sourced crash course on post-quantum cryptography: the quantum threat, NIST's new standards, deployment, migration deadlines and AI in security.
- DevelopingPost-quantum cryptography and security trackerA dated, sourced timeline of post-quantum cryptography and AI security milestones: NIST standards, deployment, government deadlines, 2024-2026.
- WikiCrypto-agilityCrypto-agility is the ability to replace cryptographic algorithms without rebuilding systems. Why the post-quantum transition made it a priority.
- WikiDARPA AI Cyber Challenge (AIxCC)DARPA's two-year competition for AI systems that find and patch software vulnerabilities, won by Team Atlanta in August 2025. Results and significance.
- WikiNIST Post-Quantum Cryptography projectNIST's open, multi-year competition that produced the ML-KEM, ML-DSA and SLH-DSA standards, plus HQC, FN-DSA and the work still in progress.
- AnalysisHow fast must we move to post-quantum cryptography?Q-Day timing, 2029 corporate targets versus 2030-2035 government deadlines, and whether to rush new algorithms: the post-quantum migration debate.