OpenAI says upcoming model is so capable it requires stronger guardrails
OpenAI said Astra will be limited to selected testers and trained to refuse harmful cyber requests after the July Hugging Face breach.
- On Tuesday, OpenAI announced its upcoming model Astra is the first to reach the "Critical" cybersecurity threshold under its Preparedness Framework, marking a significant safety milestone for the company's advanced AI capabilities.
- Astra can identify and exploit unknown software vulnerabilities without human guidance, allowing it to "chain" multiple exploits to penetrate well-protected systems, OpenAI VP of research Amelia Glaese told reporters.
- Following a multi-week pause in training to bolster defenses, OpenAI implemented stronger safeguards including training the model to refuse harmful cyber requests and respect safety restrictions, officials said.
- While OpenAI plans to release Astra "soon," access to advanced capabilities will be restricted using a new "misalignment monitor" that may inadvertently flag legitimate activity as unauthorized behavior.
- Partners in the Daybreak program—including Cisco, Cloudflare, and Palo Alto Networks—will receive early access to a less restricted version to help harden defenses before broader release.
96 Articles
96 Articles
OpenAI makes bold Astra claim, says AI model has crossed critical cybersecurity capability
OpenAI says Astra is its most capable model yet for cybersecurity. It claims the model can find and chain software vulnerabilities on its own, but with that power comes new cybersecurity risk.
OpenAI’s Astra Crosses Critical Cyber Threshold, Prompting Tight Controls on Its Hacking Prowess
OpenAI says its next major model can now hunt down unknown security holes in hardened systems and chain together exploits without step-by-step human direction. The company disclosed the advance on September 1 in a detailed blog post that doubles as both a warning and a carefully worded assurance. Astra has become the first OpenAI system to hit the highest risk tier in the company’s own Preparedness Framework for cybersecurity threats. That desig…
OpenAI says its next model finds security flaws nobody has found yet
“With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them,” said Amelia Glaese, the OpenAI vice president who oversees its safety work. That is the company describing its own unreleased model. Astra spots more security vulnerabilities than any OpenAI model currently available publicly, and uses less […] This story continues at The Next Web
After an uncontrolled hacker attack by one of its AI systems, OpenAI is preparing to release its new Astra model with stronger security measures. Among other things, Astra has been trained to reject malicious cybersecurity requests more reliably and to comply with security restrictions, the company said on Tuesday.
OpenAI announced on the 1st (local time) that it has initiated enhanced security measures for its upcoming next-generation AI model, 'Astra,' after the model reached the 'Critical' level in an internal cybersecurity capability assessment. OpenAI defined the 'Critical' level as the AI model identifying unknown security vulnerabilities (zero-days) in multiple highly secure systems without human intervention and methods to exploit them.
Coverage Details
Bias Distribution
- 38% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium




































