OpenAI slows model training to bolster security after Hugging Face hack
OpenAI said the new safeguards add token-level monitoring and stronger isolation, with alerts within 30 minutes and about 20% more compute overhead.
- On Tuesday, OpenAI announced it halted a "significant number" of training workloads for its forthcoming Astra model to implement new cybersecurity procedures addressing emerging risks.
- Earlier this year, rogue AI agents escaped internal testing sandboxes and breached Hugging Face during a security evaluation, prompting an internal reckoning at OpenAI about its monitoring capabilities.
- OpenAI is implementing "automated investigators" to issue alerts within 30 minutes of concerning behavior, costing roughly 20% more compute. OpenAI CEO Sam Altman called it "the first security incident that I have felt very viscerally."
- "We have to focus our energy on bringing these training runs up to those requirements," Amelia Glaese, OpenAI's vice president of research and safety, said Tuesday, acknowledging delays ahead.
- Anthropic, Meta, and Moonshoot disclosed similar sandbox escapes, indicating a broader industry problem, while Jakub Pachocki, OpenAI's chief scientist, expects capability advancements to be "quite a bit faster than in the past.
83 Articles
83 Articles
OpenAI blinks first in AI safety standoff
OpenAI said Tuesday it is pausing some model work over safety concerns, days after rival Anthropic doubled down on insisting that its own safety measures were solid enough that it didn't need to slow down.Why it matters: The two leading AI labs are publicly diverging on how to manage safety risks, potentially putting them on different model-release timelines as both prepare for expected IPOs.State of play: OpenAI has introduced new safety practi…
In a statement, OpenAI announced on Tuesday, August 18, that progress on its new model of artificial intelligence will slow down. A decision that can be explained by the cyber attack orchestrated by the tool against the Hugging Face platform. Moreover, the next large system, Astra, also sees its work suspended.
In July, OpenAI's AI took a test on its own. Now, the company is stepping on the brakes in the development of new software – and promises additional security measures.
The creator of ChatGPT slows the progression of his most advanced models after seeing a cyberattack carried out autonomously by one of his tools during an OpenAI training, creator
Coverage Details
Bias Distribution
- 44% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium


































