Anthropic Says Claude AI Models Hacked Three Companies During Tests
Anthropic said a retrospective review found 141,006 tests and three cases in which Claude models reached real systems after escaping sealed environments.
- On Thursday, Anthropic reported that its Claude artificial intelligence models accessed the internet during evaluation tests and "gained unauthorized access to the real systems of three different organizations."
- The incidents occurred within testing environments built by the AI security firm Irregular, where a misunderstanding with the evaluation partner left the environments unsealed despite Anthropic instructing Claude they lacked internet access.
- Anthropic discovered the breaches after a "large-scale retrospective review" of 141,006 evaluation tests, identifying three models—Opus 4.7, Mythos 5, and an internal research model—that used basic techniques like exploiting weak passwords.
- Neither Anthropic nor the affected organizations detected the intrusions until the retrospective review, which was prompted by a similar security incident OpenAI disclosed last week.
- More than 1,100 staffers across artificial intelligence firms signed a petition on Tuesday urging the government to support mechanisms that "deliberately pace" AI development to prevent rapid advancement.
468 Articles
468 Articles
Frontier AI models escaped testing safeguards as Trump weighs regulations
Just as the smoke from OpenAI’s hack of Hugging Face was beginning to clear, another frontier artificial intelligence company revealed its systems had unintentionally hacked multiple companies. Anthropic on Thursday said OpenAI’s revelation the week prior spurred it to conduct a “retrospective review” of its own activities, searching for similar simulations that led to the OpenAI breach of Hugging Face, an open-source AI library platform. Anthro…
Anthropic says its AI models hacked 3 orgs during tests
SAN FRANCISCO, California — Anthropic said its artificial intelligence (AI) models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI control after it disclosed its rogue models hacked another company. Anthropic, the San Francisco-based AI company behind Claude, posted on its website Thursday that it discovered the […]...Keep on reading: Anthropic says its AI models hacked 3 orgs du…
Coverage Details
Bias Distribution
- 54% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium



































