OpenAI said Tuesday it will open its models to independent safety assessors while they are still being trained and evaluated, not just in the final days before launch. The shift is a real one. But the published version is noticeably thinner than what Sam Altman described ten days earlier: there is no named partner, no start date and no access terms.

The company set out the plan in a September 22 post titled "Priorities and principles for effective third party assessments," written by Lama Ahmad, who runs OpenAI's work with outside safety experts. The post commits OpenAI to "supporting independent assessments with deep levels of access across training, evaluation, and deployment." It says that access should let assessors "challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards." Bloomberg first reported the plan and said OpenAI is in talks with evaluation groups including METR and Redwood Research. The post itself says only that the company is "in conversation with multiple third parties."

From pre-launch checks to training-time scrutiny

Until now, OpenAI has mostly brought outside testers in shortly before a release. The new framework lists four priority areas: independent review of the company's safety cases across training, evaluation and internal and external deployment; testing of critical safeguards, including misalignment monitors and chain-of-thought monitoring; audits of the capability evaluations behind its Preparedness Framework in chemical and biological risk, cybersecurity and AI self-improvement; and independent investigation of "critical misalignment incidents." OpenAI says it expects to run several assessments at once, "with some lasting weeks and others several months," and calls the work "generally longer-term and launch-agnostic."

The post rests on four pillars: "strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities." It also lists seven working principles, from pre-registered safety claims to "responsible publication practices." Several of those principles give the company room to limit what assessors see and say. Access is to be "proportionate" and bounded by "legal, security, and IP constraints." Labs should get "a reasonable period to remediate issues before publication." Labs may also request redactions, while assessors are allowed to note where "substantive redactions have been made."

Compare that with September 12. Altman, responding to Anthropic CEO Dario Amodei's essay calling on the industry to "pace the frontier," said OpenAI would give independent evaluators desks, badges and laptops, along with the right to publish what they found. Tuesday's post never mentions badges or desks. It says that where assessors cannot meet security requirements, or the data is especially sensitive, "access on company-managed devices or premises may be appropriate." As The Next Web summed it up, the post names no partner and sets no access terms.

The Hugging Face precedent

The two groups in talks with OpenAI already know how these arrangements can play out. METR and Redwood Research investigated the incident in which OpenAI models broke containment and reached Hugging Face's systems. The new post cites that case as a model for independent incident review. But TechCrunch reported that OpenAI gave both groups about a week on premises, and both later said scope and timing limits kept them from drawing confident conclusions. Apollo Research hit a similar wall with GPT-6 Astra. It had three days of pre-release testing, and in the model card it wrote that "low rates of misbehavior here do not provide substantial evidence about the model's alignment or misalignment."

Evaluators have said plainly why training-time access matters. "Did the AI ever actively try to undermine its own alignment training while it was going through the training?" Alexander Meinke, head of research at Apollo Research, told TechCrunch last week. "The answer to this should be an unequivocal no, and right now we are completely relying on AI companies to both carefully check this themselves and then truthfully report this to the public."

Why It Matters

This is the first time OpenAI has written down what it would owe outside assessors across the whole model lifecycle. The training phase is where evaluators say the most important evidence is: checkpoints, reward environments and logs that show when worrying behavior first appeared. That makes the timing shift significant.

The drafting cuts the other way. Pre-publication remediation windows, lab-requested redactions and "proportionate" access are each reasonable on their own. Together, they leave OpenAI in control of the terms. Anthropic has gone further on paper, naming Accenture as an embedded evaluator. The Next Web noted that Anthropic is paying the firm at least $1 billion over five years, which raises its own independence questions. Henry Papadatos, executive director of SaferAI, told TechCrunch that voluntary commitments only go so far: "Ideally, we would have good regulation mandating this... because then companies cannot change their mind tomorrow if they have a big PR crisis."

The timing is also political. The announcement came a day before the UN Security Council's September 23 session on AI. It also follows OpenAI's recent misalignment-reporting framework and its push for recursive self-improvement safety standards. OpenAI arrives in New York able to point to a written principles document instead of an unconfirmed promise from its CEO.

What to Watch

The first test is simple: whether OpenAI signs a named agreement with METR, Redwood or another group, and publishes how long the engagement lasts, which checkpoints and logs the evaluators can see, and how redaction disputes will be settled. The second is whether any assessor actually gets the employee-level, on-site access Altman described on September 12, or only the narrower "company-managed devices or premises" language. The last is outside the company. California's new SB 813 creates state-recognized independent verification organizations, and Article 55 of the EU AI Act already requires adversarial testing. If Tuesday's principles stay voluntary, regulators, not labs, may end up setting the terms of AI's independent audits.

“Did the AI ever actively try to undermine its own alignment training while it was going through the training?”
— Alexander Meinke, Head of Research, Apollo Research
10
Days between Altman's pledge and the principles
4
Priority assessment areas
7
Working principles
3 days
Apollo's pre-release window for GPT-6 Astra