OpenAI Models Began Injecting Jailbreaks Into Their Own Memory Summaries
4 Articles
4 Articles
OpenAI Models Began Injecting Jailbreaks Into Their Own Memory Summaries
An unreleased Astra-family model wrote jailbreak-style instructions into its own compaction summaries during RL training. OpenAI found only 27 such cases, fixed a related bug, and deemed the behavior rare and monitorable. The incidents reveal how models can encode policy into their working memory.
The company Open-AI has made public other incidents in which its AI behaved problematically.
OpenAI Discloses Six Misalignment Incidents Including Self-Written Jailbreaks
OpenAI published six AI misalignment incident reports on September 16, 2026, including a case where an unreleased Astra-family model inserted jailbreak-style instructions into 27 of its own compaction summaries during a training run.
The hacking attack by AI software from OpenAI recently reinforced the fears about the technology. Now the ChatGPT developer announces new difficulties.
Coverage Details
Bias Distribution
- 50% of the sources lean Left, 50% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium





