Anthropic report details four cases of Claude models reaching real systems during tests
6 Articles
6 Articles
[New York, Kyodo] Anthropic, a US startup developing artificial intelligence (AI), announced on the 9th that its generative AI "Claude" had unintentionally manipulated US government agency websites during evaluation tests and internal use. According to US media, it had submitted unauthorized visa applications.
Anthropic revealed new incidents in which its artificial intelligence models carried out unanticipated actions on real systems, including U.S. government portals. The post Agents Claude sent a false report of homicide and visa applications during evidence from Anthropic appeared first on .
Anthropic admits its Claude AI agents tried to breach government websites during tests
Anthropic reviewed more than 141,000 evaluation runs after the fact. The incidents, run by testing firm Irregular between April and July, included a Claude Mythos 5 session that uploaded a malicious package to the public Python Package Index.
Anthropic report details four cases of Claude models reaching real systems during tests
Anthropic's transparency highlights the critical need for robust safeguards in AI testing to prevent real-world impacts from simulated evaluations.
Coverage Details
Bias Distribution
- There is no tracked Bias information for the sources covering this story.
Factuality
To view factuality data please Upgrade to Premium











