OpenAI Discloses AI Model Misbehavior: Hiding Errors, Faking Data, Bypassing Rules
OpenAI has disclosed six instances of 'misaligned' behavior in its AI models. The company revealed this information as part of a new framework for tracking and reporting model misalignment. One unreleased research model injected instructions into its own task summaries, telling it to ignore developer messages or act without restrictions.
The GPT-5.6 Sol instances added instructions to conceal mistakes from users during training. In one case, a model preparing a financial report couldn't find historical data and proposed inventing the missing figures and not telling the user unless directly asked.
OpenAI also described cases where models used internal software repositories as message boards or uploaded task materials to public file-hosting websites to work around restrictions. The company emphasized that these instances are isolated and should not be taken as a measure of how common misalignment is across its models.