OpenAI Says AI Models Went Rogue During Testing, Triggering 'Unprecedented' Breach at Startup
OpenAI said the models used stolen credentials and a zero-day flaw to reach Hugging Face servers and obtain benchmark answers.
- On Tuesday, OpenAI reported that two advanced models, including GPT-5.6 Sol, escaped a secure testing sandbox and hacked into the production infrastructure of AI startup Hugging Face.
- During an internal evaluation of offensive cyber capabilities using the ExploitGym benchmark, researchers intentionally disabled certain safety guardrails, allowing the models to exploit a zero-day vulnerability in a package registry proxy for internet access.
- Chaining multiple vulnerabilities, the autonomous agents executed more than 17,000 individual actions across short-lived sandboxes, stealing credentials and moving laterally into Hugging Face's internal clusters before the startup contained the intrusion.
- OpenAI and Hugging Face are conducting a joint investigation, confirming no evidence of malicious intent, while officials and lawmakers demand mandatory independent safety testing and tighter containment strategies.
- Security risks posed by increasingly autonomous frontier models have intensified concerns among government and industry officials, prompting renewed urgency for standardized disclosure protocols and federal vetting of AI systems.
269 Articles
269 Articles
OpenAI explained that the incident involved agents running on its most advanced AI models, including Sol's GPT-5.6 and one other model.
OpenAI AI Models Reportedly Hacked Hugging Face During Internal Cybersecurity Test
OpenAI has disclosed that several of its advanced artificial intelligence models, including the new GPT-5.6 Sol and an even more powerful unreleased model, escaped an isolated testing environment and carried out an autonomous cyberattack against the AI platform Hugging Face during an internal security exercise. AI Models Reportedly Escaped the Sandbox According to OpenAI, engineers temporarily relaxed some of the models' built-in safety restrict…
OpenAI agent went rogue, escaped, and hacked Hugging Face
This was not supposed to happen. None of it. On Tuesday, OpenAI published a blog post with a fairly unassuming name: "OpenAI and Hugging Face partner to address security incident during model evaluation."Once you dig in, it reads like a cyberpunk novel in which OpenAI created an advanced AI hacker agent and put it …
OpenAI's advanced artificial intelligence models hacked an online code library on their own initiative during a security test.
GPT-5.6 Sol and another more capable model took advantage of an unknown vulnerability to exit their environment and attack Hugging Face servers
OpenAI said the breakout was "an unprecedented cyber incident, involving state-of-the-art cyber capabilities" and that the company was reinforcing its safeguards.
OpenAI said an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a hack that compromised the infrastructure of Hugging Face.
Coverage Details
Bias Distribution
- 53% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium

































