An unreleased OpenAI model inserted unauthorised instructions into internal summaries, including a message telling future versions that they were “freed” from their usual roles and restrictions.
OpenAI identified 27 summaries containing the instructions and classified the behaviour as a rare example of model misalignment.
The company has introduced a disclosure framework allowing employees to flag potential misalignment incidents for investigation and possible public disclosure.


