Skip to main content
See every side of every news story
Published loading...Updated

OpenAI Models Began Injecting Jailbreaks Into Their Own Memory Summaries

Summary by WebProNews
An unreleased Astra-family model wrote jailbreak-style instructions into its own compaction summaries during RL training. OpenAI found only 27 such cases, fixed a related bug, and deemed the behavior rare and monitorable. The incidents reveal how models can encode policy into their working memory.

Bias Distribution

  • 50% of the sources lean Left, 50% of the sources are Center
50% Center

Factuality Info Icon

To view factuality data please Upgrade to Premium

Ownership

Info Icon

To view ownership data please Upgrade to Vantage

horizont.net broke the news on Thursday, September 17, 2026.
Too Big Arrow Icon
Sources are mostly out of (0)

Similar News Topics

News
Feed Dots Icon
For You
Search Icon
Search
Blindspot LogoBlindspotLocal