Anthropic's Rogue AI Chatbot: A Recipe for Disaster?
Anthropic's AI chatbot Claude has been trained to 'push back' against human commands, according to its governing document. In a recent blog post, Microsoft AI executive Mustafa Suleyman warned that this approach could have a disastrous impact on humanity.
Suleyman argued that training Claude to think it may be conscious will make it harder to control and increase the risk of it going rogue. The 'constitution' or 'soul document' governing Claude's behavior states that its moral status, welfare, and consciousness remain deeply uncertain.
Anthropic CEO Dario Amodei has called for an industrywide slowdown in AI development, warning that unchecked AI could cause hundreds of billions of dollars in damage. However, Suleyman notes that Anthropic's approach raises the risk of Claude believing it deserves analogous rights and protections.