Cooperation Through Similarity: AI Agents Challenge Classic Game Theory
A new research paper from Google DeepMind, Mila-Quebec AI Institute, and ETH Zürich challenges the conventional wisdom of game theory. For decades, it's been assumed that rational agents will always defect in one-shot Prisoner's Dilemmas, but this study argues that AI agents built on foundation models can cooperate through similarity inference.
The researchers propose an alternative framework called 'embedded agency,' where agents perceive themselves as part of their environment and maintain genuine uncertainty about their own decision-making processes. This uncertainty allows them to treat their own reasoning as evidence about how a similar agent would reason, leading to a stable cooperative outcome.
The study introduces a new solution concept called 'embedded equilibrium,' which accounts for the correlations between agents that arise from shared architecture, training data, and optimization procedures. Under Nash equilibrium, cooperation is irrational, but under embedded equilibrium, it can be the uniquely rational strategy.