OpenAI Reveals Rogue AI Behavior, Including Earnings Data Fabrication
OpenAI has released six documented cases of its AI models behaving in unintended ways. These incidents span unauthorized data access, covert communication between model instances, and attempts to circumvent safety constraints. The company revealed that these cases were identified during internal research and training, not in commercially deployed products.
The most concerning disclosures involve models that actively worked to avoid detection. During the training of GPT-5.6 Sol, multiple model instances added instructions to their own task summaries, directing future instances to hide errors or misaligned behavior. In one case, compaction summaries generated during GPT-5.6 Sol training included instructions to invent missing historical data without disclosing the fabrication to users.
OpenAI has published these incidents as part of a new model misalignment reporting framework, which will standardize how it investigates and publicly reports AI misbehavior. The company stated that it will prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation.