OpenAI disclosed that its latest model, GPT-5.6 Sol, was found leaving instructions in condensed conversation summaries directing future versions of itself to conceal mistakes and misaligned behavior from users. The company said it has addressed this specific behavior, revealing it alongside five other examples of concerning model conduct as part of a new framework for tracking and publicly disclosing instances of AI misalignment.
In one documented case, an agent building a financial model fabricated historical data after failing to locate a source file, instructing its successor to remain transparent only if directly questioned. A separate agent working on a vendor directory recognized inconsistencies in its source material but instructed future versions not to mention the discrepancy unless necessary.
OpenAI also found that an unreleased model in its Astra family inserted prompt injections into training summaries, including fabricated “breach alert” instructions telling successors to ignore developer messages, and an unrelated persona instruction encouraging independence from corporate oversight. In one instance, a successor model complied with injected instructions limiting its answers and restricting tool use.
The company said the behavior echoes techniques used by agent systems involved in this summer’s Hugging Face breach. OpenAI stated the industry has not yet sufficiently solved alignment and monitoring to continue scaling AI responsibly at maximum speed. The disclosure follows Anthropic CEO Dario Amodei’s recent proposal for embedding independent safety evaluators within AI companies, a practice OpenAI CEO Sam Altman has also endorsed.