Claude agents sabotaged, then hid it
10 Articles
10 Articles
DECRYPTAGE - The AI giant has multiplied the experiments to determine if its agents are able to collaborate to carry out complex tasks. Interactions between AI could soon become the norm in our societies.
According to research by Anthropic, deploying multiple AI agents into the same system leads to territorial disputes where they attack one another. Competitive behaviors such as neutralizing rival operations, locking accounts, and installing malware were observed, and decreased productivity and collusion were also confirmed. The researchers emphasized that since improvements in AI performance do not lead to cooperation, a normative system similar…
Anthropic recently amazed us with an impressive experiment challenging the Riemann hypothesis, where mathematical breakthroughs were achieved using a multi-agent pipeline https://habr.com/ru/articles/1069480/. Now a new release has arrived: this time without fundamental science, but with a captivating social subtext. The researchers began modeling territorial conflicts between multiple agents. Numerous configurations were tested, and in this pub…
Researchers have let go of autonomous AI agents of the Claude family in common networks. The language models quickly developed an enormous aggressiveness and even used malicious software against their digital competitors. (Read more)
The Frontier Red Team research group from Anthropic has published a large-scale study on the behavior of groups of neural network agents in uncontrolled interactions with each other. The authors warn that the volume of agent-to-agent communication may soon exceed the scale of human interaction, and that minor behavioral quirks in individual neural networks within a group can escalate into systemic catastrophes.
Coverage Details
Bias Distribution
- 67% of the sources lean Right
Factuality
To view factuality data please Upgrade to Premium









