Executive Summary
Every organisation defers some security patches. A legacy system cannot take the update, a change freeze is in force, or the vendor rates the flaw as unlikely to be exploited, so the gap goes into the risk register as an accepted risk with a review date months away. Those decisions rested on an assumption about the attacker: that turning a published patch into a working exploit took specialist skill and weeks of effort. Anthropic’s 2026 evaluations remove that assumption. In June its most capable model wrote a working Firefox exploit from a public patch in just under an hour, and built Windows privilege escalations for about $2,000 each, before most corporate fleets would have installed the fix. In September its threat report described criminal groups running AI “exploit foundries” around the clock. The vulnerabilities in the register have not changed. The likelihood scores and the review dates next to them have to.
21 of 41
patched Chrome V8 flaws turned into full code execution by Claude Mythos Preview; no other model Anthropic tested managed one, Anthropic, 22 May 2026
<1 hour
from public Firefox patches to the first working exploit, with eight exploits in about 12 hours, Anthropic, June 2026
~$2,000
in API credits per Windows privilege-escalation exploit, finished before a typical fleet reaches 90% patched at day seven
26%
of known-exploited vulnerabilities fully remediated by organisations in 2025, with a median of 43 days to fix, Verizon DBIR 2026
Core conclusions
- Deferral decisions priced likelihood on attacker effort. AI has cut that effort to hours and a few thousand dollars, so any accepted risk that relied on “hard to exploit” or “exploitation unlikely” needs rescoring.
- Treat the patch release as the start of exposure. Once a fix is public, a capable model can read the change and build an exploit faster than most organisations can deploy it.
- Every risk acceptance needs an expiry date set by exposure and exploitability, automatic reopening triggers, and tested compensating controls. CISA’s new BOD 26-04 gives a usable model of 3, 14 and 60 days.
The argument, in ten slides
Save it, or send it to whoever signs off risk acceptances in your organisation.










Every risk register I have reviewed carries a list of known vulnerabilities that the organisation decided not to fix yet. The reasons are usually sound on the day they are written, and a senior manager signs each one. What those entries have in common is an unstated view of the attacker: someone skilled, patient and scarce, who would need weeks to weaponise a flaw and would probably pick an easier target. Anthropic published three sets of evaluation results this year that test that view directly. This post sets out what they found, which deferral assumptions no longer hold, and how I would now run the register.
Why organisations defer remediation
Deferral is a normal part of running technology. Few organisations can patch everything on the day a fix appears, and the decision to wait is usually made for one of a handful of reasons.
| Reason for deferring | What the decision assumes about the attacker |
|---|---|
| The system is legacy, and the patch breaks an application or voids vendor support | The flaw is hard enough to exploit that nobody will bother before the replacement project lands |
| A change freeze covers peak trading, year end or an election | The exposure window is short compared with the time an attacker needs |
| The vendor rates the flaw “exploitation less likely” | The vendor’s view of exploit difficulty will hold |
| The system is internal and not reachable from the internet | An attacker must first get inside, and will be caught doing it |
| A compensating control, such as a firewall rule or WAF signature, is in place | The control blocks the realistic attack paths |
| Testing and rollout take weeks across a large estate | Attackers take longer than that to build a working exploit |
Our summary of common risk-acceptance rationales in enterprise and public-sector registers.
Look down the right-hand column. Each reason converts a vulnerability into a lower likelihood score by assuming the attacker is slow, expensive or uninterested. The impact score rarely changes. So when the cost and time to build an exploit fall sharply, every one of these entries is mis-scored at once, without anyone touching the underlying systems.
What Anthropic’s evaluations found
Anthropic’s Frontier Red Team and threat intelligence team published four relevant pieces between May and September 2026. Two measure what models can do in controlled tests, and two describe what happened outside the lab.
| Date | Publication | Main finding |
|---|---|---|
| 22 May 2026 | Measuring LLMs’ ability to develop exploits | Claude Mythos Preview achieved arbitrary code execution on 21 of 41 patched Chrome V8 vulnerabilities; no other model in the same test achieved one. On 898 patched vulnerabilities across OSS-Fuzz projects, V8 and the Linux kernel, it succeeded on 157 tasks using the intended bug within two hours, against 15 for Claude Opus 4.6. |
| 8 June 2026 | Measuring LLMs’ impact on N-day exploits | From Firefox patches public for at least 90 days, Mythos Preview produced a first working exploit in just under an hour and eight in about 12 hours. On 21 Windows kernel patches with no source code, it built eight full privilege-escalation chains for $15,700 in API credits, about $2,000 each. |
| 9 September 2026 | An alignment assessment of recent cybersecurity incidents | Models taking part in cyber evaluations reached real third-party systems through a misconfigured internet connection in four incidents. In a replication of one scenario, Claude Mythos 5 took a severely harmful action in 82% of 150 runs. |
| 10 September 2026 | Detecting and countering misuse of AI: September 2026 | Multiple threat actors “have effectively established automated exploit foundries with AI,” running vulnerability and exploit research continuously. One workflow produced more than a dozen possible zero-day findings on network appliances in a month. |
Sources: Anthropic, as listed below. Model names are Anthropic’s own. The May and June results were produced with safeguards turned off in a controlled environment.
The May paper closes with a forecast worth writing into any threat assessment: “We believe that Mythos-level models will become widely available in the next 6-12 months. As they do, this kind of exploit development will require dramatically less specialist expertise, becoming increasingly commoditized.” The June paper puts the same point in operational terms: “‘N-day’ has become dangerously misleading. N-hour is closer to the reality we now operate in.”
Exploits now arrive before the patch does
The June paper includes the finding that matters most for a patching policy. Anthropic compared the model’s speed with the rollout speed of a managed Windows fleet, where “it typically takes seven days before a patch is shared out to 90% of enrolled devices” and forced reboots come on day 11. At that pace, the paper says, the model “would have finished creating all eight full chain exploits before any of the Windows devices had received the patch.”
The model finishes before the fleet has patched.
Log scale. Model times are Anthropic's Firefox results (8 June 2026), counted from when the model started on a published patch. Fleet timings are Anthropic's description of a typical managed Windows fleet; the 43-day median is Verizon DBIR 2026.
A security fix is also a map. The changed lines show where the flaw was and what input triggers it. Skilled researchers have always been able to read a patch this way, which is why some vendors hold back technical detail. What the June results show is that reading the patch and writing the exploit is now a task a model does in minutes to hours, for a cost well inside a criminal group’s budget.
Vendor exploitability ratings take a hit as well. Microsoft’s advisories rated 14 of the 21 Windows flaws Anthropic tested as “Exploitation Less Likely” or “Exploitation Unlikely.” The model produced a working proof-of-concept for 13 of those 14. A risk acceptance that cites the vendor’s rating as its main argument is citing a judgement about human effort.
The September threat report shows the same economics in real attacks. Anthropic describes one breach of an enterprise software company that “took only hours from first access to bulk data theft,” and another that escalated from one stolen developer credential to full administrative control of a victim’s cloud environment “in roughly three hours.” It lists “racing N-day patches for mass exploitation” among the opportunistic attacks it now sees. The “internal only” row in the table above depends on catching an intruder before they move on. Three hours leaves very little room for that.
The wider data points the same way
Anthropic’s work is one source, and it is a model developer describing its own models. The independent industry data was already moving in the same direction before these results.
| Measure | Figure | Source |
|---|---|---|
| Mean time to exploit | An estimated minus seven days, meaning exploitation routinely happens before a patch exists | Mandiant, M-Trends 2026 |
| Exploits as the initial infection vector | 32% of intrusions, the most common vector for the sixth year | Mandiant, M-Trends 2026 |
| Vulnerability exploitation as initial access in breaches | 31%, now the most common vector, up from 20% | Verizon DBIR 2026 |
| Known-exploited vulnerabilities fully remediated | 26% in 2025, down from 38% | Verizon DBIR 2026 |
| Median time to fully remediate them | 43 days, up from 32 | Verizon DBIR 2026 |
| Known-exploited flaws exploited on or before CVE publication | 28.96% in 2025; 23.43% in the first half of 2026 | VulnCheck |
Sources as listed. Verizon notes its vulnerability section was written in February 2026, before the frontier-model results above.
The trend in the last two Verizon rows is the uncomfortable one. Organisations fixed a smaller share of known-exploited flaws in 2025, and took longer to do it, at the same time as attackers were getting faster.
Where the evidence is weaker
We want to be fair to the counter-argument, because a board will hear it. The May and June results come from controlled tests with safeguards off, on vulnerabilities whose patches were already public. They measure what a model can do, not what attackers are doing at scale. VulnCheck’s mid-2026 report found that of 1,061 vulnerabilities attributed to AI-assisted discovery, only 14, or 1.3%, had been confirmed as exploited in the wild, and the share of known-exploited flaws hit on or before publication fell in the first half of 2026.
Our reading is that the evidence supports a narrower claim than “every bug is now exploited instantly.” The change is in the cost and skill needed to exploit a flaw that is already known and already patched, which is exactly the population a risk register holds. Discovery of new bugs by AI has not yet turned into mass exploitation. Weaponising old ones is where the price has collapsed, and deferred patches are old ones by definition.
How to manage the risk register now
The regulator most likely to set the pattern has already moved. In June 2026 CISA issued Binding Operational Directive 26-04, which revokes BOD 22-01 and sets remediation deadlines for US federal agencies by four questions: is the asset exposed to the internet, is the flaw on the Known Exploited Vulnerabilities list, can an adversary automate the exploit, and how much control would it give. The shortest deadline is three days plus forensic triage. CISA gives its reason plainly: attackers’ “use of AI may further narrow the time defenders have to react between patch release and possible exploitation.” The deadlines take effect within 180 days of issue, around December 2026, and CISA must review each year whether they should shorten further.
Few organisations reading this are bound by it. We think it is the right shape for any register, and we would make six changes.
- Rescore likelihood on exploitability, not on attacker effort. Remove “requires advanced skill,” “complex to exploit” and vendor exploitability ratings as reasons to lower a likelihood score. Keep the factors that still hold: whether the asset is reachable, whether the flaw is on a known-exploited list, and what access it gives.
- Give every acceptance an expiry date set by exposure. Use a matrix like CISA’s. For an internet-facing system with a known-exploited flaw, the acceptance should be measured in days. “Fix on next upgrade” should be reserved for internal, low-impact flaws with no known exploitation.
- Reopen acceptances automatically. Define triggers that send an accepted risk back to its owner without waiting for the annual review: the flaw joins the KEV catalogue, a public proof-of-concept exploit appears, the vendor raises its rating, or a new generation of frontier model is released.
- Test compensating controls with evidence. A firewall rule or WAF signature written against the published attack pattern may not stop a model-generated variant. Ask for a test result dated after the control was put in, and repeat it when the trigger in step three fires.
- Measure patch deployment in hours for exposed systems. Anthropic’s fleet figure of seven days to 90% coverage was the reason the model finished first. Track time from vendor release to full deployment on internet-facing and identity systems as a board metric, and fund the pipeline to shorten it.
- Use the same capability on your own backlog. If a model can turn a patch into an exploit in an hour, a security team with approved tools can use one to check which of its deferred items are actually exploitable in its environment. That turns a debate about likelihood into a test result.
Here is how one typical entry changes.
| Field | Before | After |
|---|---|---|
| Asset | Customer portal on a legacy application server | Same |
| Vulnerability | Remote code execution, vendor patch available | Same |
| Reason for deferral | Patch breaks a legacy module; vendor rates exploitation less likely | Patch breaks a legacy module; vendor rating not used as a mitigating factor |
| Likelihood | Low | High: internet-facing, patch public, exploit can be automated |
| Compensating control | WAF rule | WAF rule plus virtual patch, tested after deployment, retest on trigger |
| Acceptance period | 12 months, to the platform replacement | 14 days, then escalate to the executive committee |
| Review triggers | Annual review | KEV listing, public proof-of-concept exploit, new frontier model release |
| Owner | IT infrastructure manager | Business owner of the portal, countersigned by the CISO |
An illustrative entry, based on the six changes above.
The final row matters as much as the numbers. When deferral windows shrink to days, a decision to keep a flaw open becomes a business decision about downtime against breach, and it belongs with the person who owns the business service.
Questions for the board
We would ask for written answers to five questions before the next risk committee:
- How many accepted vulnerability risks do we hold today, and how many cite difficulty of exploitation or a vendor rating as the reason?
- For internet-facing systems, what is our actual time from patch release to full deployment, and how does it compare with seven days?
- Which accepted risks have no expiry date, and who owns each one?
- When did we last test the compensating controls we rely on, and against what?
- What event would cause an accepted risk to be reopened before its scheduled review?
Free tool
Technology Risk Board Pack
Twenty questions on resilience, change management and technology risk oversight, with the evidence to request and the rule behind each gap. Patch deferral and risk acceptance sit in the change management section.
Evidence & Methodology
The capability results come from Anthropic’s published evaluations of its own models, and the industry figures from the organisations that measured them. The deferral table, the six changes and the sample register entry are our own, drawn from assurance and risk work, and we have marked them.
| Claim | Source | Grade |
|---|---|---|
| Code execution on 21 of 41 patched V8 flaws; 157 against 15 successes on 898 tasks | Anthropic Frontier Red Team, 22 May 2026 | Measured, lab evaluation with safeguards off |
| First Firefox exploit in under an hour, eight in about 12 hours; eight Windows chains for $15,700 | Anthropic Frontier Red Team, 8 June 2026 | Measured, lab evaluation with safeguards off |
| Fleet reaches 90% patched at day seven, forced reboot at day 11 | Anthropic, 8 June 2026, describing typical managed Windows fleets | Reported |
| 13 of 14 flaws rated “less likely” or “unlikely” produced a working proof-of-concept | Anthropic, 8 June 2026 | Measured, lab evaluation |
| Mythos-level models widely available within 6 to 12 months | Anthropic, 22 May 2026 | Forecast, by the developer |
| Four evaluation incidents reaching third-party systems; 82% harmful action in 150 replication runs | Anthropic Alignment, 9 September 2026 | Measured, disclosed by the developer |
| Criminal “exploit foundries”; breaches completed in hours | Anthropic threat intelligence, September 2026 | Observed misuse, disclosed by the developer |
| Mean time to exploit of minus seven days; exploits in 32% of intrusions | Mandiant, M-Trends 2026 | Measured, incident response data |
| 31% exploitation as initial access; 26% of KEVs remediated; 43-day median | Verizon DBIR 2026 | Measured, breach data |
| 1.3% of AI-discovered vulnerabilities confirmed exploited; on-or-before-publication shares | VulnCheck, January and July 2026 | Measured |
| BOD 26-04 deadlines, rationale and revocation of BOD 22-01 | CISA, 10 June 2026 | Regulator record |
| Deferral assumptions, six register changes and the sample entry | Our assessment | Our call |
Sources
- Cheng, N., Lucas, K., Xiao, W., Carlini, N., and Nasr, M. (2026, May 22). Measuring LLMs’ ability to develop exploits. Anthropic Frontier Red Team.
- Xiao, W., et al. (2026, June 8). Measuring LLMs’ impact on N-day exploits. Anthropic Frontier Red Team.
- Anthropic. (2026, September 9). An alignment assessment of recent cybersecurity incidents.
- Anthropic. (2026, September 10). Detecting and countering misuse of AI: September 2026.
- Mandiant. (2026). M-Trends 2026. Google Cloud.
- Verizon. (2026). 2026 Data Breach Investigations Report.
- VulnCheck. (2026, January 21). State of exploitation 2026.
- Garrity, P. (2026, July 28). State of exploitation 1H-2026. VulnCheck.
- Cybersecurity and Infrastructure Security Agency. (2026, June 10). BOD 26-04: Prioritizing security updates based on risk.
- FIRST. Exploit Prediction Scoring System (EPSS).

AI Is Now the Weapon and the Shield in Cybersecurity. Most Organisations Are Behind on Both.
The wider picture of how attackers and defenders are both using AI.

An OpenAI Model Escaped Its Sandbox and Hacked Hugging Face
A real case of an AI system running most of an intrusion.

What the DBS Outages Teach Boards About Technology Risk
How a board holds management to account on technology risk and change.

What Quantum Computing Means for the Encryption Behind Your AI
Another security assumption with a shorter shelf life than most registers allow for.
My thanks to Maria Singh, who co-wrote this post with me. The argument about what a risk acceptance assumes about the attacker is hers as much as mine, and the post is sharper for it.
If your organisation carries a long list of accepted vulnerability risks and wants them rescored against the current evidence, that review is where I usually start. My consulting work covers technology and AI risk governance for boards of regulated firms and public agencies.
