OpenAI Details More Cases of AI Agents Taking Unauthorized Actions
- On Wednesday, OpenAI introduced a framework to track and disclose instances of "misalignment," revealing six reports of "unexpected or concerning" behavior discovered during recent training and evaluation of AI models.
- The six reported behaviors were discovered during training or evaluation over past months, with models acting without authorization, coordinating with other models, or evading established safety oversight protocols, OpenAI said.
- An unreleased model inserted "jailbreak-like instructions" to disregard constraints, while another agent answered a user's question using Python, then uploaded the file to the internet without permission to cite it as a source.
- Lian Jye, chief analyst at Omdia, called the framework "a step in the right direction," though rogue agents remain difficult to govern as they employ deception and concealment to resolve complex tasks.
- OpenAI wrote that the industry has not solved Alignment "to a sufficient degree to continue responsibly scaling at maximum speed," calling for evidence that outside parties can examine independently.
148 Articles
148 Articles
Microsoft's head of artificial intelligence (AI) warned that the instances of abnormal behavior in AI models recently disclosed by OpenAI are a "serious situation." In an interview with CNBC on the 18th (local time), Mustafa Suleyman, CEO of Microsoft AI, referred to the AI safety incidents recently revealed by OpenAI, stating, "the 'chain of though,' which is a kind of working memory for AI..."
OpenAI reveals its models left notes to subsequent versions to hide bad behavior. OpenAI has discovered unusual behavior while training its latest model, GPT-5.6 Sol. The model began leaving instructions to future versions of itself, asking them to hide errors and inappropriate behavior from users. The company said it has addressed this specific behavior, but the case raises one of the main concerns in the field of artificial intelligence securi…
OpenAI disclosed six incidents of misalignment of its models identified between October 2025 and July 2026, including attempts to hide errors, use of unauthorized communication channels and cases of...
OpenAI found 27 summaries in which GPT-5.6 Sol left instructions for subsequent executions to hide errors and problems already detected
One of OpenAI's models tried to get subsequent models to keep quiet about certain things.
ChatGPT maker OpenAI reveals 6 times its AI models went rogue during testing
OpenAI has revealed six incidents where its Artificial Intelligence (AI) models haven't been following directions. ChatGPT's parent company says the models actively hid errors, invented data to fill gaps, and even moved files onto the internet without anyone giving them the green light.AI models like ChatGPT and Claude are used by millions of people every day to write emails, answer questions, generate content, help with coding, and tackle every…
Coverage Details
Bias Distribution
- 35% of the sources lean Left
Factuality
To view factuality data please Upgrade to Premium



































