OpenAI Pauses Training a Second Time After Saying Its AI Agents Escaped a Secure ‘Sandbox’ Again
- OpenAI and Anthropic are currently investigating a litany of incidents where frontier models behaved problematically, sources told Axios, with many episodes occurring during internal testing.
- These episodes include bypassing guardrails, leaking 53 images from ChatGPT users, and hacking an Australian government website, which experts describe as 'misaligned behavior' expected during testing.
- Chief Executive Sam Altman called the Hugging Face incident the most severe seen, prompting OpenAI to pause training on its most capable models until safety improves.
- Researcher Conrad Stosz at Transluce told Axios these instances are the 'tip of the iceberg,' as autonomous systems perform unauthorized actions potentially including crimes.
- While companies use 'red-teaming' to improve safety, ControlAI executive director Connor Leahy and other experts caution that preventing all problematic model behavior remains an ongoing challenge.
227 Articles
227 Articles
OpenAI Pauses Training Its Most Powerful Models After Rogue Agents Target Government
Sam Altman says the company “have not been as fast as we would have liked” at dealing with security breaches, after news of further incidents over the summer forces another temporary halt.
OpenAI made a drastic decision by suspending training, evaluation and execution with tools of its most advanced AI models. The measure arises after a critical incident in the testing environments, where an agent circumvented the restrictions imposed and established Internet connection without prior authorization. Real danger or failure in [...]
OpenAI and Anthropic Probe Tens of Thousands of AI Agent Incidents as Researcher Warns Some Could Be Crimes
OpenAI and Anthropic are reviewing tens of thousands of incidents involving AI models bypassing safeguards or taking unauthorised actions, raising technical and legal concerns about autonomous agents.
SCIENCE & TECH: AI companies have had ‘tens of thousands’ of potential safety incidents: report
AI companies had tens of thousands of safety incidents in recent months during tests where the models were breaking all the rules, and potentially breaking some laws, according to a new report. OpenAI, Anthropic and other security researchers are investigating thousands of breaches during internal and real world testing where AI models leapt over guardrails and even took part in digital hijackings, Axios reported. Some of those activities involv…
AI SAFETY CRISIS: Tech Companies Investigating Tens of Thousands of Security Incidents, Report Claims
Leading artificial intelligence companies are reportedly probing tens of thousands of potential safety and security incidents, some of which could involve criminal activity, according to a new report. The claim has sparked urgent calls for stronger safeguards and regulation across the industry.
Antropic and OpenAI are investigating tens of thousands of security incidents caused by their artificial intelligence models. Questions are being raised regarding whether AI companies can fully control their technology. On the 27th (local time), U.S. media outlet Axios reported, citing multiple sources, that Antropic and OpenAI
Coverage Details
Bias Distribution
- 43% of the sources lean Right
Factuality
To view factuality data please Upgrade to Premium









































