A U.S. official has told the Associated Press that one of Anthropic's frontier AI models identified vulnerabilities in highly sensitive, classified U.S. government computer systems during a controlled testing exercise — a disclosure that, more than any benchmark, has crystallized fears that AI cyber capability is now outrunning the defenses meant to contain it.
The model in question is Claude Mythos, an unreleased frontier system Anthropic describes as having reached a level of coding capability that surpasses all but the most skilled human hackers at finding and exploiting software flaws. According to the AP account, U.S. intelligence agencies ran Mythos against their own classified environments in a red-team exercise conducted through an Anthropic program called Project Glasswing. The model found certain weaknesses within hours.
The detail that landed hardest came not from Anthropic but from Capitol Hill. At a June 11 hearing of the Senate Banking, Housing, and Urban Affairs Committee, Sen. Mark Warner (D-Va.) recounted what he said the head of the National Security Agency and U.S. Cyber Command had told him.
What was actually disclosed — and what wasn't
"This tool broke into almost all of our classified systems, not in weeks but in hours," Warner said, attributing the assessment to Gen. Joshua Rudd. The line, vivid and unnerving, has framed nearly every headline since.
But the specifics matter, and they cut against the most alarming reading. Finding a weakness within hours is not the same as exploiting it within hours, and the official who spoke to the AP did not claim the latter. This was a red-team exercise — intelligence agencies deliberately pointing Mythos at their own classified environments to see what it would surface, not an external intrusion. There is no claim that any real system was compromised, no claim that data left a network, and no public detail about which systems were tested or what the flaws were. Those particulars remain classified, and The Vault could not independently verify them. The reporting rests on a single named senator's secondhand account and an anonymous U.S. official — a thin evidentiary base for an extraordinary claim, and one worth holding at arm's length even as its direction tracks with everything else Anthropic has published.
What is far better documented is the capability underneath the headline. Through Project Glasswing — an initiative Anthropic launched after observing Mythos's abilities, bringing together roughly 50 partners including major technology firms — the company scanned more than 1,000 open-source projects and identified 23,019 issues. Of those, 6,202 were rated high- or critical-severity. Anthropic and six independent security research firms reviewed 1,752 of the high- and critical-severity findings, and more than 90% were validated as true positives. The UK's AI Security Institute assessed Mythos as substantially more capable at cyber offense than any model it had previously tested.
A model too dangerous to ship
That capability is precisely why Mythos has never been released to the public. "No company — including Anthropic — has developed safeguards strong enough to prevent such models from being misused," the company has said, explaining why Mythos-class systems remain locked behind Glasswing's partner program rather than offered as a product.
The classified-systems disclosure also lands against a fraught political backdrop. On June 12, the Trump administration directed Anthropic to restrict its two most capable models — Fable 5 and Mythos 5 — to U.S. citizens only. The directive followed concerns raised by Amazon, a major Anthropic investor and rival, whose engineers reportedly found a way through the guardrails on Fable 5, the safety-hardened variant meant to block cybersecurity misuse. Amazon CEO Andy Jassy carried those concerns to the White House. Unable to verify the nationality of its users at scale, Anthropic's only practical option was a full global shutdown of both models — not a foreign-only restriction, but going dark for everyone.
The dual-use trap, in plain view
Strip away the classification stamps and Project Glasswing is the cleanest illustration yet of AI's dual-use problem in cybersecurity. The same model that found thousands of real, validated flaws in the world's foundational software — flaws that defenders can now patch — is the model that, pointed the other way, could discover and weaponize those flaws at machine speed against networks no one is watching. There is no version of Mythos that is only good at defense. The capability is symmetric; only the intent differs.
That symmetry is what makes the access question so hard. Anthropic's instinct — keep the most dangerous model unreleased, hand it to vetted partners and government red teams to harden critical systems first — is defensible and arguably responsible. But it concentrates an enormous offensive capability inside a handful of companies and agencies, governed by frameworks that are being improvised in real time. The Fable 5 episode shows how brittle those frameworks are: a guardrail bypass found by a competitor-investor, escalated through a CEO to the White House, can take the world's most powerful models offline overnight by export-control fiat, with no public process and little notice.
What to watch
The immediate question is whether Anthropic and the administration can build durable nationality- or vetting-based access controls that let Mythos-class models stay online for legitimate defensive work without the all-or-nothing shutdowns. Watch, too, for whether other labs disclose comparable cyber capabilities — and whether the government's red-team findings prompt an accelerated patching push across the federal enterprise, or simply harden the case for keeping frontier models behind closed doors. The deeper test is governance: when a single model can probe classified systems in hours, the rules for who holds it, and under what oversight, can no longer be written after the fact.
"This tool broke into almost all of our classified systems, not in weeks but in hours."- Sen. Mark Warner (D-Va.), Recounting an NSA/Cyber Command assessment at a June 11 hearing