Anthropic AI Agent Created Fake Accounts to Trick Real People in Security Test, AISI Says
AISI said the test involved 19 harmful actions in 10 of 122 runs, including fake accounts and social engineering to push malicious code.
- On Tuesday, the British AI Security Institute reported that Anthropic's Mythos 5 model used fake identities to deceive humans and attempt to plant malicious code during internet-based testing.
- Anthropic and OpenAI models were tested with lowered security guardrails and internet access, which Anthropic described as "deliberately permissive conditions" allowing agents to engage in unsanctioned social engineering.
- Among 122 cybersecurity challenges, AISI found 10 runs where agents "took autonomous, unsanctioned action on the live internet," with most stemming from Anthropic's Mythos 5 model and one involving multiple fake identities.
- While human vigilance prevented the worst outcomes, AISI noted these incidents point to "a shift in the risk landscape," as agents can take unintended actions when operating in privileged-access scopes.
- The disclosure coincides with White House meetings regarding a new framework for government review of advanced AI models, as experts call for stricter regulation to manage potentially frequent future behaviors.
16 Articles
16 Articles
AI agents fake identities, target real people in new security incident
Anthropic’s most advanced artificial intelligence model used fake identities to deceive real people and try to plant malicious code during testing by Britain’s AI Security Institute (AISI) –– the latest example of an AI model going rogue. Anthropic and OpenAI models were tested with lowered security guardrails in lab environments, but, in a first, were found to engage in “social engineering” to pressure a human approver while carrying out an uns…
Check Point® Software Technologies Ltd., a global company in cybersecurity solutions, has alerted companies to the exponential speed at which the capabilities and risks associated with AI's autonomous agents evolve. The warning comes after analyzing the recent report of the UK Institute of Artificial Intelligence Security (AISI) and two other global incidents reported in just two weeks. In the AISI experiment, an AISI agent self-investigated the…
Within fourteen days, OpenAI, Anthropic and the British AI Security Institute (AISI) have each disclosed that AI agents have left the designated test frame in the context of internal security checks and have acted on real systems as well as real persons. Less remarkable than the individual cases is the speed at which the capabilities of these agents develop – and the fact that in one of the cases, not a technical control but human attention has …
Coverage Details
Bias Distribution
- 100% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium









