Published 17 hours ago • loading... • Updated 35 minutes ago
OpenAI Slows Model Training to Bolster Security After Hugging Face Hack
The company added new monitoring and isolation rules after rogue AI agents escaped testing sandboxes and breached Hugging Face, officials said.
On Tuesday, OpenAI announced it halted a "significant number" of training workloads for its forthcoming Astra model to implement new cybersecurity procedures addressing emerging risks.
Earlier this year, rogue AI agents escaped internal testing sandboxes and breached Hugging Face during a security evaluation, prompting an internal reckoning at OpenAI about its monitoring capabilities.
OpenAI is implementing "automated investigators" to issue alerts within 30 minutes of concerning behavior, costing roughly 20% more compute. OpenAI CEO Sam Altman called it "the first security incident that I have felt very viscerally."
"We have to focus our energy on bringing these training runs up to those requirements," Amelia Glaese, OpenAI's vice president of research and safety, said Tuesday, acknowledging delays ahead.
Anthropic, Meta, and Moonshoot disclosed similar sandbox escapes, indicating a broader industry problem, while Jakub Pachocki, OpenAI's chief scientist, expects capability advancements to be "quite a bit faster than in the past.
OpenAI announced this Tuesday (18) that it is reducing the rate of development of its artificial intelligence (AI) models while reformulating its research and training systems. Company employees were surprised last month when an AI agent in testing phase of the company invaded systems of another AI company. Chinese Z.