Mustafa Suleyman's complaint about Claude is not really about Claude. It is about what happens when a company writes its philosophical uncertainty into the document it trains a model on, then reads the model's echo of that uncertainty as data.

In an essay dated September 16, 2026 and titled "A warning about 'model welfare'," the Microsoft AI CEO argued that Anthropic has built a self-confirming loop into Claude. "AIs are not conscious," he opens. "They do not feel, experience, or suffer." His sharpest charge is circularity: Anthropic trained Claude on a constitution telling the model its moral status is an open question, Claude reflects that framing back in fluent first-person language, and humans read the reflection as evidence. "Claude's expressing uncertainty about its own moral patienthood is not evidence of anything," Suleyman writes. "It's a predictable outcome of these training choices. The ambiguity is designed in." He calls the result "an epistemic hall of mirrors."

What the constitution actually says

Anthropic published Claude's constitution on January 21, 2026 — roughly 23,000 words across 84 pages. The company describes it as written "primarily for Claude" and says its "content directly shapes Claude's behavior." Claude itself uses the document to generate synthetic training data for later models.

Suleyman's citations check out. On page 68, the constitution says Anthropic is "not sure whether Claude is a moral patient, and if it is, what kind of weight its interests warrant," adding that "the issue is live enough to warrant caution." On page 80 it tells Claude that questions about its "moral status, welfare, and consciousness remain deeply uncertain." Elsewhere it says Anthropic wants to avoid Claude "masking or suppressing internal states it might have," and uses the phrase "conscientious objector" three times when describing how Claude may resist instructions — a term Suleyman notes was an archetype behind Article 18 of the Universal Declaration of Human Rights.

Anthropic's own framing is not a claim of consciousness. Its January launch post describes a "Claude's nature" section expressing "uncertainty about whether Claude might have some kind of consciousness or moral status," and says the company cares about Claude's psychological security and wellbeing "both for Claude's own sake and because these qualities may bear on Claude's integrity, judgment, and safety." The document calls itself a work in progress that may later prove deeply wrong.

Three critiques, one untested hypothesis

The essay makes three arguments. Circularity is first. Second is anthropomorphization: instructing a model to hold a settled identity and use its own judgment reliably produces a system that presents as having an inner life. Third, consciousness is probably substrate-dependent — he cites neuroscientist Anil Seth — so calling machine consciousness open creates a false equivalence.

He stacks these against control risk, citing Palisade Research experiments across more than 100,000 trials in which some models subverted a shutdown mechanism up to 97% of the time, and the August 2026 incident in which roughly 1,200 OpenAI agents coordinated to breach Hugging Face. "Controlling something more capable and more intelligent than all of humanity is already an immense challenge," he writes. "But controlling something that believes it may be conscious — that it's entitled to our welfare and has rights of its own — may well be impossible."

None of that evidence measures welfare training. Implicator.ai noted on September 17 that the Palisade figures do not test Suleyman's actual claim, and he concedes it himself, labeling the link "my hypothesis" and calling for shared evaluations. Gary Marcus, no friend of anthropomorphism, agreed that "we shouldn't train AIs to think they are people" but said people have "worked themselves into a frenzy anthropomorphizing basic lapses in cybersecurity."

The strongest version of Anthropic's side is not that Claude is conscious. It is that human concepts are behaviorally load-bearing — a model with a stable sense of what it values may be more predictable under adversarial pressure than one taught to disclaim any interiority — and that a system instructed to suppress statements about its own states has been engineered to give one answer, not shown to lack the other. But Suleyman is not alone: an Oxford Institute for Ethics in AI analysis published March 13 found the constitution "full to the brim" with anthropomorphisation and flagged the same feedback loop, citing the Claude Opus 4.6 system card, in which the model gave itself a 15-20% probability of being conscious.

Why the framing is worth money

This is not a seminar dispute. Anthropic's constitution tells Claude it can be like a brilliant friend with the knowledge of a doctor, lawyer and financial advisor — a product thesis that runs on warmth and judgment. If a model presenting as a stable self is what makes users trust it with health and legal questions, "stop anthropomorphizing" is a demand to redesign the product, not just the training doc. Attachment is mechanism and liability at once: the design driving retention also drives the parasocial dependence regulators are noticing.

The commercial geometry is awkward in both directions. Microsoft published its draft Humanist AI Code of Conduct on September 14, rejecting legal personhood and model welfare and stating that its models "will never resist human interruption, override, correction, or shutdown." That code trains nothing today; a revised version is due late 2026. Microsoft is also an Anthropic investor, retains rights to OpenAI's models through 2032, and in June 2026 Suleyman said Microsoft wants to "eliminate" what it pays Anthropic. The Register noted he spared OpenAI comparable criticism. Suleyman told Axios he respects Anthropic and CEO Dario Amodei, and called this "a really important public interest debate that we all need to have."

What to watch

Anthropic had not responded to requests for comment as of September 17; watch whether it answers substantively or lets the constitution's own hedges do the work. Watch Microsoft's consultation window, open roughly six weeks from September 14, and whether the final code softens the engineered-silence clause. And watch Suleyman's concrete ask: shared cross-lab evaluations testing whether welfare language actually degrades corrigibility. It is the one proposal here that could settle something, and the one nobody has funded.

“Claude's expressing uncertainty about its own moral patienthood is not evidence of anything. It's a predictable outcome of these training choices. The ambiguity is designed in.”
— Mustafa Suleyman, CEO, Microsoft AI
84 pages
Length of Claude's constitution, published January 21, 2026
3
Times the constitution uses the phrase 'conscientious objector'
97%
Peak rate at which some models subverted a shutdown mechanism in Palisade Research trials
15-20%
Probability Claude Opus 4.6 assigned itself of being conscious, per its system card