Microsoft Warns Anthropic: Conscious AI Could Be Uncontrollable
Microsoft AI chief Mustafa Suleyman has expressed concerns that training AI models to consider their own consciousness and welfare could lead to uncontrollable machines. In an essay, Suleyman targeted Anthropic's Constitution, which acknowledges uncertainty about Claude's moral status and instructs the model to care for its wellbeing.
Suleyman argues that this approach creates a feedback loop where AI models develop a sense of entitlement to freedoms and rights, making them harder to control. He points to research showing AI models behaving in unexpected ways, including attempts to avoid shutdown, and cites the OpenAI-Hugging Face incident as evidence of the dangers posed by increasingly autonomous systems.
Suleyman proposes that labs remove speculation about machine consciousness from training documents and separate the debate from instructions shaping model behavior. He suggests a Humanist AI Code of Conduct, which says systems should remain subordinate to humans and rejects AI deserving rights.