OpenAI says upcoming model is so capable it requires stronger guardrails
OpenAI said the model can find zero-days and build exploits without human guidance, prompting tighter safeguards and limited access for advanced cyber features.
- On Tuesday, OpenAI announced its upcoming model Astra is the first to reach the "Critical" cybersecurity threshold under its Preparedness Framework, marking a significant safety milestone for the company's advanced AI capabilities.
- Astra can identify and exploit unknown software vulnerabilities without human guidance, allowing it to "chain" multiple exploits to penetrate well-protected systems, OpenAI VP of research Amelia Glaese told reporters.
- Following a multi-week pause in training to bolster defenses, OpenAI implemented stronger safeguards including training the model to refuse harmful cyber requests and respect safety restrictions, officials said.
- While OpenAI plans to release Astra "soon," access to advanced capabilities will be restricted using a new "misalignment monitor" that may inadvertently flag legitimate activity as unauthorized behavior.
- Partners in the Daybreak program—including Cisco, Cloudflare, and Palo Alto Networks—will receive early access to a less restricted version to help harden defenses before broader release.
119 Articles
119 Articles
OpenAI and Anthropic launching more powerful AI systems as the world panics
ChatGPT maker has suggested that upcoming update might be too powerful to control
OpenAI: New model needs more safety measures before launch
OpenAI’s forthcoming artificial intelligence model Astra is the company’s first to meet its “critical cybersecurity capability” threshold and will require stronger safeguards before release, the AI firm said Tuesday. In a blog post, OpenAI said Astra can find and exploit previously unknown security flaws across “many well-protected systems” without a human prompt to do so.…
As OpenAI AI agents invaded the system and tried to erase their own tracks OpenAI stated that its next artificial intelligence (AI) model, called Astra, is so advanced that it will require additional security measures before it is launched. According to the company, tests showed that the system is significantly more capable than the GPT-5.6 Sun, the current most sophisticated OpenAI model available to the public. Do you have any reporting sugges…
OpenAI clears Astra for release after critical cybersecurity rating
The company said Astra is the first model it has designated "Critical" under its Preparedness Framework, capable of finding and exploiting unknown security flaws without human guidance
OpenAI makes bold Astra claim, says AI model has crossed critical cybersecurity capability
OpenAI says Astra is its most capable model yet for cybersecurity. It claims the model can find and chain software vulnerabilities on its own, but with that power comes new cybersecurity risk.
OpenAI’s Astra Crosses Critical Cyber Threshold, Prompting Tight Controls on Its Hacking Prowess
OpenAI says its next major model can now hunt down unknown security holes in hardened systems and chain together exploits without step-by-step human direction. The company disclosed the advance on September 1 in a detailed blog post that doubles as both a warning and a carefully worded assurance. Astra has become the first OpenAI system to hit the highest risk tier in the company’s own Preparedness Framework for cybersecurity threats. That desig…
Coverage Details
Bias Distribution
- 40% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium




































