When the US government yanked Claude Fable 5 offline in June, the jailbreak that triggered it had not come through any formal research channel. It came from internal testing by Amazon. Now, as part of the deal that got the model restored, Anthropic has built the channel it lacked: a dedicated HackerOne program for security researchers to report cyber jailbreaks in Fable 5.
From export-control order to bug bounty
The program is one thread in a larger settlement. Anthropic released Fable 5 and its more capable sibling Mythos 5 on June 9, 2026. Three days later, on June 12, the Commerce Department applied emergency export controls citing national-security authorities, ordering Anthropic to cut off access for any foreign national inside or outside the United States. Because the order took effect immediately and the company had no reliable way to verify nationality in real time, it suspended both models for everyone. Access was restored beginning July 1 after Commerce lifted the controls on June 30.
The trigger, by Anthropic's own account, was a report in which Amazon researchers found a way to prompt Fable 5 into identifying software vulnerabilities and, in one case, producing code demonstrating how a vulnerability could be exploited. Anthropic played the finding down, saying its testing confirmed that weaker models — including Claude Opus 4.8, OpenAI's GPT-5.5, and China's Kimi K2.7 — could identify the same vulnerabilities, and that every model it tested could reproduce the single exploit demonstration. The government and the reporting partner nonetheless treated it as serious enough to justify emergency action.
To resolve the standoff, Anthropic trained a new safety classifier that it says blocks the reported technique in more than 99% of cases, routing flagged requests to the weaker Opus 4.8 and notifying users. It also committed to deeper government collaboration and helped draft an industry framework for scoring jailbreak severity. Announcing the HackerOne channel in its June 30 post, the company wrote that it is "launching a new HackerOne program where security researchers can submit potential cyber jailbreaks they've discovered in Fable 5 (once available) for our review."
A narrow, deliberate scope
The program, hosted at hackerone.com/anthropic-cyber-jailbreak, is scoped tightly to cyber jailbreaks — techniques that bypass Fable 5's cybersecurity classifiers — rather than general model safety or non-cyber misuse, which continue to route through Anthropic's standard responsible-disclosure process. HackerOne, the long-running vulnerability-coordination platform, manages submission, triage, and researcher communications. Vetted researchers attempt bypasses under controlled conditions, and successful submissions are handled as responsible disclosures.
Anthropic has run HackerOne bug bounties before, including invite-only programs in 2025 that challenged red-teamers to find universal jailbreaks in unreleased safety classifiers, with rewards up to $25,000 for verified findings on CBRN topics. Several aggregator accounts of the new Fable 5 effort describe it primarily as a disclosure channel rather than a fixed-reward bounty; Anthropic's own post frames it as a submission-and-review program and does not publish a reward schedule for it, so the exact incentive structure is best described cautiously pending program details.
Why the channel matters
The episode exposed a structural gap in how frontier AI gets stress-tested. Software security has long relied on coordinated disclosure and severity scoring — the Common Vulnerability Scoring System is the standard reference — but AI jailbreaks have had no equivalent. Anthropic, with Amazon, Microsoft, Google, and other Project Glasswing partners, is now proposing a four-factor severity framework scoring capability gain, breadth, ease of weaponization, and discoverability. The HackerOne channel is the intake pipe that feeds such a framework: a formal front door so that the next significant bypass surfaces through a managed disclosure process rather than an ad hoc internal report that escalates to a Cabinet secretary.
That points to the harder problem the redeployment left unresolved: government-gated frontier models. Fable 5 never went through the voluntary pre-release review path created by the June 2 executive order; the government reached for export controls instead. As The Hacker News observed, when Washington wants to move fast on a frontier model, it still has "no binding process, only improvised ones." A bug bounty cannot fix that, but it does something narrower and real — it gives outside researchers a sanctioned way to probe a model whose capabilities regulators now treat as national-security relevant, and it gives Anthropic a defensible claim that it, not an adversary, will find the next major jailbreak first.
What to watch
Watch whether Anthropic publishes reward tiers and researcher-eligibility rules, which will signal how seriously it wants to attract top red-teamers versus tightly control who can probe the model. Watch whether the four-factor severity framework gets adopted beyond the Glasswing group and whether any government body blesses it. And watch the first high-severity submission that arrives through HackerOne rather than a partner lab — the real test of whether this channel changes how frontier-model risk gets surfaced, triaged, and disclosed.
"We're also launching a new HackerOne program where security researchers can submit potential cyber jailbreaks they've discovered in Fable 5 (once available) for our review."— Anthropic, from its June 30 redeployment post