Five days from now, the European AI Office expects to receive something no regulator anywhere has ever collected: formal, legally compelled systemic risk evaluations from the companies building the world's most capable AI models.

By September 15, providers of general-purpose AI models trained above 10^25 floating-point operations are due to submit their first evaluations, according to compliance guidance circulating among practitioners. The Office is expected to scrutinize red-teaming methodologies, energy consumption disclosures, and adherence to the standardized copyright training-data summary template published in July. Worth flagging: September 15 is not a date written into the AI Act. It is a reporting window derived from the Act's enforcement architecture, and providers should treat the underlying obligations, not the calendar entry, as binding.

The legal machinery behind it is unambiguous. Article 51 of the AI Act presumes that any GPAI model trained above 10^25 FLOPs poses systemic risk. The presumption is rebuttable, but the burden sits with the developer, not the regulator. Epoch AI's April 2026 compute data put 12 notable models above that line — from OpenAI, Google, Anthropic, Meta and Mistral — turning a threshold that existed only in text into a roster of named companies.

What Brussels Can Actually Do

The Office's teeth arrived on August 2, when Section 5 of Chapter IX became applicable. Before that date, the AI Office could ask. After it, the Office can compel.

Three powers stack in sequence. Article 91 lets the Office demand technical documentation, training data summaries and model reports. Article 92 lets it evaluate a model directly, appointing independent experts from the Scientific Panel and requesting access through APIs or other technical means, including source code. Article 93 lets it order corrective measures, up to withdrawing or recalling a model from the EU market.

Article 101 sets the price of refusal at 3 percent of worldwide annual turnover or 15 million euros, whichever is higher. That is below the Digital Services Act's 6 percent and GDPR's 4 percent, but the structure differs in a way that matters: fines on GPAI providers are levied by the Commission directly, not by 27 national authorities. For a company the size of Alphabet, 3 percent of global turnover would clear 9 billion euros. Article 50's transparency obligations — machine-readable marking of synthetic content, deepfake disclosure — took effect the same day, at the same penalty ceiling.

The July 2026 Digital Omnibus on AI deferred high-risk system obligations as far as December 2027, pending harmonized standards. It did not move the GPAI enforcement date. It did hand the Office new territory: systems built on GPAI models where model and system share a provider, and AI integrated into very large online platforms under the DSA.

Analysis: Can 165 People Audit the Frontier?

The Commission added 38 staff to the AI Office at the end of July and is recruiting roughly 40 more contractual agents, bringing headcount to something near 165 across six units. Against a mandate covering worldwide monitoring of frontier model compliance, watermarking verification, a new whistleblower intake channel and now DSA-adjacent systems, that number is thin.

It is thin by the standards of the Commission's own advisors. A report by the Brussels NGO Pour Demain found the Office's projected staffing and budget inadequate to the enforcement demands ahead, recommending GPAI supervisory capacity scale to at least 160 staff by 2030 — for that function alone. Yoshua Bengio and Marietje Schaake have called for an AI Safety unit of 100 and an implementation team of 200. As of March 2026, the Head of AI Safety unit and Chief Scientific Advisor posts were reported vacant.

Writing in Lawfare in May, Harvard Kennedy School fellow Joel Christoph described the Office as "significantly underresourced relative to its mandate," noting that the pool of independent evaluators qualified to assess frontier models "is thinner still." His closing line has become a kind of shorthand in Brussels policy circles: "The tools are on the table. The question is whether anyone picks them up."

The first filing round is the test. If the Office issues Article 91 information requests to every Code of Practice signatory in the weeks after September 15, it establishes routine supervision as the baseline before gaps harden into accepted practice. If it waits for a visible failure, providers will read that as permission to treat the filings as paperwork. The DSA precedent is instructive: the Commission opened proceedings against X within months of the VLOP rules biting, issued no first-year fines, and still moved platform behavior.

A second test sits inside the submissions. The Act asks for evaluation against state-of-the-art benchmarks, but no settled methodology has been published — and twelve labs filing twelve incompatible red-teaming frameworks would satisfy the letter of the obligation while producing nothing a regulator could compare or act on.

Industry has spent a year arguing the exercise is premature. The Computer and Communications Industry Association — members include Apple, Meta and Amazon — ran a sustained simplification campaign across the EU digital rulebook. "Europe cannot lead on AI with one foot on the brake," said Daniel Friedlaender, who heads CCIA's European operations, arguing critical parts of the Act were still missing weeks before the rules applied. After the omnibus landed, CCIA Europe AI policy lead Boniface de Champris was unimpressed from the opposite direction: "Given the minimal improvements made to the AI Act, the glaring gap between political rhetoric on regulatory simplification and concrete outcomes is hard to ignore." Commissioner Henna Virkkunen says she wants innovation-friendly implementation but has refused to entertain a pause.

What to Watch

Three things. Whether the Office confirms receipt and completeness of the first filings, or lets the date pass without comment. Whether Meta — the most prominent major non-signatory to the GPAI Code of Practice — draws the heavier scrutiny the Commission's guidelines promise non-signatories. And whether any provider tries to rebut the Article 51 presumption outright, forcing the Office to defend the 10^25 threshold on technical rather than statutory grounds.

Epoch AI projects the count of models above frontier compute thresholds keeps climbing. A filter that concentrates scarce oversight on twelve models works. The same filter applied to fifty stops being a filter at all.

“The tools are on the table. The question is whether anyone picks them up.”
— Joel Christoph, Fellow, Harvard Kennedy School, writing in Lawfare
Sept 15
Filing deadline
10^25 FLOPs
Systemic risk threshold
12
Models above the threshold
3% / 15M EUR
Max GPAI fine