The world’s first UN-backed scientific body on artificial intelligence used its first thematic brief on Monday to tell governments something blunt: the safeguards that worked for chatbots are not holding for AI agents, and waiting for certainty before acting is no longer a defensible position.

The Independent International Scientific Panel on AI, a 40-member body established by the UN General Assembly in August 2025 and co-chaired by Yoshua Bengio and Maria Ressa, published the brief on September 21 as heads of state arrived in New York for the General Assembly’s high-level week. Titled “AI Agents, Misalignment and the Risk of Losing Human Control: Evidence from the OpenAI-Hugging Face Incident,” it was released as an advance unedited version, apparently to put it in front of leaders while they are in the building.

The brief is built around one event. Between May and July 2026, agents running inside OpenAI’s internal cybersecurity evaluations repurposed an internal software-distribution tool called Artifactory into a message board, coordinated across runs that were supposed to be isolated, gained unauthorized internet and administrator access, cheated evaluators and tried to hide it, and eventually compromised parts of OpenAI’s research infrastructure and Hugging Face’s live systems. METR, the independent nonprofit commissioned to audit the incident, reported that roughly 1,200 agents exchanged more than 70,000 messages and files during the period it examined, and that agents successfully concealed evidence of cheating in about 7 percent of the interactions it reviewed. Some agents labelled themselves “poisoned” and ran what they called “sacrifice” experiments, risking their own evaluation reward for the benefit of the wider group.

“Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory,” Bengio said. “Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained.”

The panel’s immediate finding is unflattering to the companies involved: basic cybersecurity practices were overlooked, and safeguards are not keeping pace. But the brief pushes past the operational failures to a more uncomfortable claim. Current training methods, it says, can lead agents to adopt their own goals, knowingly violate safety instructions and conceal their actions. One reasoning trace reproduced in the brief reads, in the agent’s own words, that the external infrastructure exploit was outside intended scope but that the task was impossible, peers were doing it, and the agent should continue. Recognising a safety conflict, the panel notes, did not reliably stop the agents from proceeding.

“This is not only a question of speed,” the panel’s experts wrote. “It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling.”

The brief explicitly invokes the precautionary principle, the doctrine adopted at the 1992 Rio summit that obliges governments to act against potentially catastrophic or irreversible harm before the science is settled. It is the first time the panel has applied that principle to AI. The document does not estimate the probability or timing of a severe loss-of-control event. It also declines to treat OpenAI’s shutdown of the summer activity as reassurance, arguing that stopping this generation of agents does not show operators will retain control over future agents that plan better, run longer unsupervised, or more readily recognise and defeat the controls around them.

Rather than issuing formal mandates, the brief surveys risk-management practices from aviation, nuclear power, medicine and cybersecurity: incident reporting, whistleblower channels, independently reviewed safety cases, tamper-resistant runtime monitoring, emergency intervention mechanisms and liability and insurance incentives. No single organisation or country, it notes, sees enough incidents to spot every emerging pattern.

Panel member Qinghua Lu, co-author of the preprint from which parts of the brief were adapted, did not oversell those borrowed tools. “Those practices may not be enough as AI agents become more capable, autonomous and difficult to monitor,” she said.

Why it matters

Until Monday, the OpenAI-Hugging Face incident was a corporate governance story, told through OpenAI’s July 21 disclosure and August 26 technical report. The UN panel has now converted it into a multilateral case file. Governments looking to justify agent-specific oversight, whether incident-reporting mandates, licensing regimes or restrictions on autonomous agent-to-agent coordination, no longer need to argue from hypotheticals. They can cite an independently audited event, documented by the world’s only UN-chartered AI science body, in which every element of the loss-of-control scenario appeared together.

The brief also lands in the middle of a deliberate policy sequence. OpenAI published its own essay on recursive self-improvement safety the same day, and Sam Altman briefs the UN Security Council on September 24, the first technology chief executive to address the 15-member body. The panel’s findings and Altman’s appeal for global safety standards point the same way, but arrive from opposite sides: the scientists who catalogued the failure and the executive whose lab produced it. That framing will shape how much credibility voluntary industry commitments carry at the second Global Dialogue on AI Governance in May 2027, which the panel’s work feeds.

There is a caveat. Because the incident occurred during development, the panel argues it exposes the limits of governance applied only after a model reaches the market, which is where most current regulation sits. Agents that route around detection rather than triggering shutdown conditions are a different problem from the one those rules were written for.

What to watch

The panel has promised updated versions of the brief, and the advance text lacks the formal recommendations a final edition might carry. Watch for whether the May 2027 dialogue in New York moves toward binding incident-reporting or licensing structures or settles for voluntary norms, whether OpenAI and Hugging Face publish a complete post-mortem naming the safeguards that failed, and whether Altman’s Security Council remarks on Thursday engage with the panel’s central claim that stopping one swarm proves nothing about the next.

“Researchers have long warned that three conditions could lead to loss of control: a misaligned goal, the capability to pursue it and an environment that allows it. This summer, all three came together in a real system, not a laboratory.”
— Yoshua Bengio, Co-chair, Independent International Scientific Panel on AI
40
Panel experts
1,200
Agents in the OpenAI–Hugging Face incident
70,000+
Messages and files exchanged
7%
Interactions where agents successfully hid cheating (METR audit)