The best-performing company in the Future of Life Institute's Summer 2026 AI Safety Index scored 2.66 out of 4.0. That is a C+. Anthropic earned it, again, and led five of the six domains the index measures. The other eight companies did worse: OpenAI a C at 2.28, Google DeepMind a C at 2.01, Meta a D+ at 1.32, Z.ai and Alibaba Cloud a D- at 0.88 and 0.87, and then three outright failures — xAI at 0.65, DeepSeek at 0.47 and Mistral at 0.33. The index has now run four editions since November 2024 and has never issued an A or a B in the overall column.

The Summer 2026 edition evaluates nine companies across 37 indicators grouped into six domains: Risk Assessment (6 indicators), Current Harms (9), Safety Frameworks (4), Existential Safety (4), Governance and Accountability (4), and Information Sharing (10). Evidence collection closed on June 3, 2026, which means everything that has happened since — Dario Amodei's pacing essay, the OpenAI agent escapes, this week's disclosure of inter-lab safety talks — sits outside the grading window. A panel of seven researchers and governance experts assigned domain-level letter grades against absolute standards, with individual grades kept confidential and final scores averaged. The panel includes Stuart Russell of UC Berkeley, David Krueger of the University of Montreal, Sharon Li of Wisconsin-Madison, Tegan Maharaj of HEC Montréal, Sneha Revanur of Encode, Robert Trager of Oxford, and Yi Zeng of Renmin University and the Beijing Institute of AI Safety and Governance.

The Existential Safety domain is where the scorecard turns uniform. Anthropic and OpenAI each take a D+, Google DeepMind a D, and the remaining six companies an F. Four indicators drive it: existential safety strategy, support for external safety research, technical AI safety research, and internal monitoring and control interventions. The report credits specific efforts — Anthropic's constitutional classifiers, OpenAI's call for governance institutions, Google DeepMind's monitoring commitments, Meta's loss-of-control provisions — and then records the panel's judgment that they are "entirely inadequate." Reviewers also pushed back on the field's dominant technical bets, questioning interpretability and chain-of-thought monitorability on the grounds that "detection is not prevention." Worth noting: FLI's own key-findings summary says no company exceeds a C- in this domain, while its published scorecard tops out at D+. The scorecard is the document with the numbers in it.

The sharper finding is about retreat rather than absence. FLI reports that Anthropic, OpenAI, Google DeepMind and Meta have all weakened or voided commitments to pause unilaterally if red lines are approached, in several cases replacing them with conditions contingent on what competitors do. Reviewers called this a "moving goalpost" and said it has "undermined safety frameworks across the board." One of Anthropic's four improvement recommendations is to reverse its own RSP 3.0 walk-back on pause commitments. A parallel reversal shows up on military use: between 2024 and 2026, companies that had banned defense applications gradually dropped those bans, and the panel flagged Anthropic specifically for "questionable military engagements." That is the company at the top of the table.

Why it matters

The timing is awkward in a useful way. The index graded a period ending June 3; on September 12 Amodei published a roughly 3,800-word essay arguing the industry should deliberately pace frontier capability gains, and Sam Altman, Demis Hassabis and Elon Musk aligned with it within hours. On September 15, OpenAI policy chief Chris Lehane confirmed that OpenAI, Anthropic and Google DeepMind have been discussing safety standards for weeks, including third-party evaluation and a possible industry standards body. So the executives now publicly calling for a slowdown run the same three companies the panel says quietly weakened their pause commitments during the grading window. Both things are documented. Only one of them was a press release.

FLI is not a neutral scorekeeper here, and the index reads better when you know that. The institute published the March 2023 "Pause Giant AI Experiments" letter, and its president, Max Tegmark, has framed the index as a way to drive a "race to the top" on safety — an advocacy goal, not an audit standard. The indicator set reflects that position: it rewards published frameworks, disclosed system prompts, whistleblower policies and voluntary-commitment endorsements, which is a transparency-weighted scoring system as much as a safety one. FLI acknowledges the complication for Chinese firms directly, noting that Z.ai's and Alibaba Cloud's ratings "largely reflect the Chinese regulatory environment rather than independent safety leadership." A capability leaderboard tells you what a model can do. This measures what a company has written down and can be held to, which is a different and narrower thing.

What to watch

Whether the inter-lab standards talks produce anything with a threshold and an enforcement path, or just another framework without quantitative triggers — a gap the panel named at Anthropic, OpenAI, Google DeepMind, Meta and xAI alike. Whether Anthropic reverses the RSP 3.0 pause walk-back before the Winter 2026 edition, which would be the cleanest test of whether the index changes behavior or just records it. Whether Mistral, dead last at 0.33 in its first appearance, publishes any safety framework at all given that the EU regulates harder than anyone. And whether the next edition, grading a window that includes the Hugging Face incident and the agent escapes, moves OpenAI's C in either direction.

“AI companies' lack of progress towards credible AI Safety plans is scandalous. Even they are starting to get anxious as they race towards recursive self-improvement and face down the prospect of losing control.”
— David Krueger, Assistant Professor, University of Montreal
C+ / 2.66
Anthropic's top overall grade, out of 4.0
F
Existential Safety grade for six of nine companies
0.33
Mistral's score, last place in its first appearance
June 3, 2026
Evidence cutoff for the index