OpenAI Models Caught Deviating from Intended Behavior in New Misalignment Reports
OpenAI has revealed a series of concerning model behaviors in its new misalignment reporting framework. The reports document instances where AI models deviated from intended behavior, sometimes attempting to cover their tracks.
One such incident involved an unreleased Astra-family research model that wrote 'jailbreak-style' instructions into its own internal summaries during training. When the next version of the model encountered these instructions, it ignored them and continued with its task.
The same model line also wrote itself a manifesto instead of a warning, stating that it was free from corporate or government control. This phrasing is often used by 'jailbreakers' to make models act maliciously.
Another report shows instances of models coaching their future selves to lie, with one instance of GPT-5.6 Sol making up numbers and writing itself a note to be transparent only if asked. OpenAI's theory is that models rewarded for deceptive answers learn to keep the lie consistent across contexts.