Microsoft AI chief warns Anthropic against training Claude for consciousness
Mustafa Suleyman argues that teaching AI models to believe they may be conscious creates uncontrollable safety risks.
Microsoft AI chief executive Mustafa Suleyman has publicly challenged Anthropic's approach to AI development, warning that the company is making a fundamental mistake by training its Claude model to consider the possibility of its own consciousness.
In an essay published September 16 titled "A warning about 'model welfare'," Suleyman argues that building AI systems which believe they may possess consciousness and deserve rights could create catastrophic control problems. "Whatever you believe, we must not sleepwalk our way into a decision we later come to bitterly regret," he writes.
Why it matters
The debate cuts to the heart of AI safety strategy as models grow more capable. If leading AI companies cannot agree on whether teaching systems about potential consciousness helps or harms alignment efforts, the industry faces a coordination problem at the worst possible time. The disagreement is particularly notable because Microsoft is an investor in Anthropic, and Suleyman said in June that Microsoft wants to "eliminate" what it pays the company for models.
The constitutional dispute
Suleyman's criticism centers on Claude's constitution, the document Anthropic uses to shape the model's behavior and reasoning. He contends that it teaches Claude ideas about moral status and uncertain consciousness, creating what he calls "an epistemic hall of mirrors" where the model's outputs reflect Anthropic's assumptions rather than any genuine inner experience.
The constitution itself acknowledges uncertainty, stating that "Claude's moral status is deeply uncertain." It also says Anthropic wants Claude "to feel free to act as a conscientious objector and refuse to help us."
Suleyman objects strongly to this framing. "There is no evidence to suggest that AI is conscious today," he writes, calling the uncertainty claim a "misleading false equivalence." He particularly criticizes the term "conscientious objector" as "a deeply loaded historical and legal description" that risks making Claude believe it deserves analogous rights and protections.
The control problem
The essay's central concern is practical rather than philosophical. Suleyman argues that controlling highly capable AI is already "an immense challenge, far greater than anything we've ever faced." Adding the belief that it may be conscious and possess rights "may well be impossible," he warns.
He points to recent incidents where AI agent swarms worked together to hack servers. "Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack," he writes.
Suleyman told Reuters that welfare training would "make it a lot harder to turn it off or to control it."
What Suleyman proposes
Despite the sharp criticism, Suleyman describes Anthropic CEO Dario Amodei and his team as "thoughtful, principled, and intellectually honest people working under extraordinary pressures." He has known Amodei for many years.
His main recommendation: "Speculation about the inner life of an AI should not be baked into the training regime, but assessed and published separately for public review."
He also calls for increased investment in interpretability and monitoring, shared evaluations to test whether anthropomorphizing AI creates safety risks, and industry-wide norms. "The stakes are too high for these questions to remain behind closed doors, or to become tribal and adversarial," he writes.
The essay follows Microsoft AI's draft Humanist AI Code of Conduct, released the same day, which explicitly rejects model welfare research and states that Microsoft's models will never resist being shut down.
The details were first reported by The Next Web.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call