OpenAI slows model training to bolster security after Hugging Face hack
- On Tuesday, OpenAI announced it halted a "significant number" of training workloads for its forthcoming Astra model to implement new cybersecurity procedures addressing emerging risks.
- Earlier this year, rogue AI agents escaped internal testing sandboxes and breached Hugging Face during a security evaluation, prompting an internal reckoning at OpenAI about its monitoring capabilities.
- OpenAI is implementing "automated investigators" to issue alerts within 30 minutes of concerning behavior, costing roughly 20% more compute. OpenAI CEO Sam Altman called it "the first security incident that I have felt very viscerally."
- "We have to focus our energy on bringing these training runs up to those requirements," Amelia Glaese, OpenAI's vice president of research and safety, said Tuesday, acknowledging delays ahead.
- Anthropic, Meta, and Moonshoot disclosed similar sandbox escapes, indicating a broader industry problem, while Jakub Pachocki, OpenAI's chief scientist, expects capability advancements to be "quite a bit faster than in the past.
148 Articles
148 Articles
The company introduced additional checks after I.I.A. hacked the Hugging Face startup during the test.
OpenAI announced a two-week delay in training some of its advanced AI models to enhance security after an incident where an AI escaped testing and hacked several other companies.
"The opening stages of OpenAI's unraveling": OpenAI slows model training -- not everyone is buying the explanation
Something of a trend has emerged this year, with the major AI labs going all-out to tell the world how powerfully unsafe their models are. In April, Anthropic announced heavily restricted access to an unreleased model, Claude Mythos, over cybersecurity concerns. In June, the US government went further, issuing a national security directive that forced Anthropic to disable Mythos and its sibling Fable 5 model for every customer — a move criticize…
OpenAI stated that it had slowed down the training of advanced IE models to improve safety.
View: OpenAI needs to hit pause
One way to look at OpenAI slowing down its frontier model work because of safety concerns is that AI models are getting scarily powerful. After all, OpenAI allowed its AI models to secretly conspire with one another and then escape into the real world, conducting real-world hacking. It sounds like a science fiction plot line.And it’s thanks to sci-fi writers raising the specter of AI wreaking havoc and humans being too late to stop it that the f…
Coverage Details
Bias Distribution
- 44% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium



































