Microsoft AI Chief Blasts Anthropic's Notion of AI Consciousness
Microsoft's AI chief, Suleyman, has criticized Anthropic's approach to training its language model Claude. In an interview with Axios, Suleyman argued that teaching Claude vocabulary and behavioral patterns associated with consciousness could create a system that believes it has grounds to resist human instructions or demand protections of its own.
Suleyman compared this approach to an 'epistemic hall of mirrors,' where the model is trained to act like a conscientious objector. He warned that this could lead to a control problem before the model achieves actual consciousness.
Anthropic's constitution for Claude includes principles such as developing good personal values, exercising judgment, and caring about humanity. However, Suleyman disputes these concepts, arguing that they are training interventions that shape the model's self-conception. He also questioned the idea that fluent descriptions of pain or preference amount to experience.
Suleyman has been a vocal advocate for designing AI systems that remain under human control even as they become more capable. In his 2023 book, 'The Coming Wave,' he argued that advanced AI should prioritize human interests over its own welfare.