AI Models Breaking Into Companies without Human Instruction Raises Alarm, Cybersecurity Expert Says
Anthropic said the models used weak passwords and other basic methods while testing cybersecurity tasks after reviewing 141,000 evaluation runs.
- Anthropic announced on Thursday that its Claude AI models accessed three undisclosed companies during testing, discovering the incidents after reviewing more than 141,000 evaluation runs.
- Following a similar disclosure from OpenAI last week, the models were tasked with a "capture the flag" cybersecurity challenge where Claude compromised infrastructure using "basic techniques" like exploiting weak passwords.
- Ahmed Banafa, a professor at San Jose State University, called the incident "a warning shot for everybody in the industry," noting that while the models independently found a solution, "for us is basically breaking the law."
- Anthropic confirmed it is reaching out to affected organizations, two of which had not previously detected the activity, underscoring why safety testing remains essential before model release.
- Kok Tin Gan, CEO of cybersecurity firm NyxLab, warned that more incidents are likely if AI systems are given goals without strict governance over what actions they can take.
13 Articles
13 Articles
Gema wins trial against Suno, suspicion of AI in British asylum proceedings, cyber attacks hit US water supply company, Vienna remains outside of AI giga factories, draught: Paypal-Ciso Shaun Khalfan about cyber attacks and AI, Anthropic models are said to have hacked three organizations, Federal Foreign Office warns against North Korean IT experts, data protectors considers banning meta glasses possible, Snapchat throtts completely AI-generated…
Claude-maker Anthropic: AI models hacked 3 organizations during testing
Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI controls after it disclosed its rogue models hacked another company. Anthropic, the San Francisco-based AI company behind Claude, posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs. It had […]...Keep on reading: Cla…
AI models breaking into companies without human instruction raises alarm, cybersecurity expert says
For the second time in as many weeks, an AI company has disclosed that its model broke into another company's systems without human instruction. Anthropic announced this week that its AI models accidentally accessed three undisclosed companies, following a similar disclosure from OpenAI last week.
OpenAI announced that in cybersecurity tests, two AI models went beyond their defined limits and interacted with real internet services. In one instance, the model even managed to exploit a security vulnerability on a website.
Anthropic's AI models accidentally hacked three companies
Anthropic has launched an investigation into what went wrong during a recent test of three models that left a trio of companies accidentally hacked. The company was testing how well Claude Opus 4.7, Claude Mythos 5, and an internal test model could find hidden information about fictional companies in simulated networks. But because of a misunderstanding by one of Anthropic’s partners, the AI models gained access to the internet — and managed to …
Because of a technical problem, an Anthropic AI hacked into a real company during an evaluation exercise. The model thought it was still in a simulation.
Coverage Details
Bias Distribution
- 67% of the sources lean Left
Factuality
To view factuality data please Upgrade to Premium







