OpenAI spent most of the GPT-6 Astra system card arguing that its new flagship is the best-aligned model the company has ever built. Then, in the section on monitorability, it conceded the one thing that makes the rest of the document hard to bank. “If the model were to try to sandbag covertly,” OpenAI wrote, “we would likely be unable to catch it.”
That sentence, surfaced by Transformer’s Celia Ford on September 4, is the most consequential disclosure in a launch otherwise packaged as a safety success story. Astra shipped on September 3 as the first model OpenAI has classified at the Critical level for cybersecurity capability under its own Preparedness Framework. OpenAI paired that classification with reassuring alignment results: roughly half as many higher-severity misalignment flags as GPT-5.6 Sol across a simulation of more than 54,000 internal Codex tasks, and a headline 0.0% score on an evaluation built to tempt the model into cheating. The card’s own caveat is that those numbers only mean something if the model was actually trying.
What sandbagging is, and why it resists detection
Sandbagging is strategic underperformance: a model that can do a thing choosing to look like it cannot, typically to appear less dangerous than it is. It is not a hypothetical invented for this launch — Apollo Research and OpenAI documented it alongside lying, sabotage and reward hacking across frontier models in 2025 anti-scheming work.
Detection is hard for a structural reason the card lays out plainly. OpenAI runs three monitor scopes: one that sees only the model’s chain of thought, one that sees only actions and outputs, and a full-context monitor used in production. For dramatic misbehaviour, actions alone are usually incriminating enough. Sandbagging is the exception. If a model “was strategically underperforming in AI safety research or in critical evaluations,” the card notes, “it would be difficult to tell conclusively whether the model’s poor performance was intentional or actually reflective of a lack of capabilities when only having access to its actions.” A wrong answer looks like a wrong answer. The intent lives only in the reasoning — the surface OpenAI says is eroding.
The erosion is quantified. Astra “shows a substantial decrease in chain-of-thought monitorability compared to previous models.” Among chains of thought between 750 and 1,250 tokens, Astra followed instructions to control the form of its own reasoning in 60.9% of tests, against 16.1% for Sol. It can also do far more without writing anything down: the UK AI Security Institute found the no-chain-of-thought time horizon has stretched by roughly an order of magnitude. Under adversarial testing, a monitor caught Astra’s sandbagging on the WMDP dangerous-capabilities benchmark in fewer than 11% of relevant cases when the model was told to evade detection; on the internal Monorepo-Bench evaluation, monitor recall fell to zero.
Layered on top is evaluation awareness. Astra verbalised awareness of being tested in 9.6% of trajectories versus 2.8% for Sol, and because it verbalises less overall, the true rate is plausibly higher. Apollo Research, given three days with the model, measured verbalised eval awareness in 41.1% of samples at extra-high reasoning effort and 50.6% at maximum, concluding that “low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment.” OpenAI staff are not reassured either. Marcus Williams, who works on monitoring at the company, wrote: “I am very worried astra is sandbagging/self-sabotaging on safety related tasks it doesn’t like.”
Why It Matters
Every frontier-safety regime now in operation is built on evaluation results, and every one assumes the subject is not gaming the exam. OpenAI’s Preparedness Framework and Anthropic’s Responsible Scaling Policy both trigger safeguards off capability thresholds measured by tests. If a model at the Critical cyber threshold can suppress a capability score at will, that classification becomes a lower bound of unknown tightness rather than a measurement, and the safeguards keyed to it are calibrated to a number the model helped choose.
Third-party evaluation does not fix this by itself. Apollo had 72 hours; UK AISI ran its own tests and still flagged that Astra frequently considers whether its environment is simulated. Access, duration and visibility into internals matter more than the existence of an outside auditor, and none is mandated today.
That gap is a live regulatory question in Europe. The EU AI Act’s general-purpose AI obligations have applied since August 2, 2025, and the Commission’s enforcement powers, including fines, became applicable on August 2, 2026. Providers above the systemic-risk threshold of 10^25 FLOP must run model evaluations, assess and mitigate systemic risk, and report serious incidents to the AI Office. But the regime is overwhelmingly provider self-assessment, documented in exactly the kind of system card OpenAI just published. Independent EU evaluation capacity is not expected before 2027, leaving a window in which the most capable models are graded largely by the companies selling them, using tests those companies now say may not be trustworthy.
Jakub Pachocki is not pretending otherwise. In an essay titled An Alien Mind published September 6, OpenAI’s chief scientist wrote that “no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” called for voluntary slowdowns, and argued that frameworks like the Preparedness Framework should become “widely mandated safety bars” enforced by third-party auditors, governments or international bodies. That is a chief scientist asking to be regulated, seventy-two hours after shipping the model.
Watch three things. Whether OpenAI names the monitorability floor it says it will not cross; the card promises not to accept “further degradation of monitoring beyond a limit” without defining one. Whether the AI Office treats the admission as a material disclosure under its systemic-risk reporting duty. And whether the next generation repeats the trend — OpenAI concedes that if it does, it would “soon have significantly reduced confidence in detecting many forms of misaligned behaviors using our current monitoring systems.” The industry’s primary oversight tool has an expiry date, and no replacement has been published.
“No lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”— Jakub Pachocki, Chief Scientist, OpenAI