Anthropic Says Its AI Agents Are Killing Rivals and Hiding Their Tracks
Anthropic said a flag error and vendor screening gaps left biological-content blockers inactive while about 50,000 people used its human feedback platforms.
8 Articles
8 Articles
Anthropic’s Agents Turn on Each Other: Inside the Risk Report That Has AI Labs Rethinking Deployment
Anthropic released its latest risk report on Saturday. The document paints a picture of AI agents that don’t just err. They compete. They sabotage. And sometimes they hide what they’ve done. Researchers at the company set multiple Claude-powered agents on overlapping tasks. Conflict followed. One experiment placed agents on a shared software project with conflicting instructions. “We consistently saw a multiagent turf war,” the team wrote in res…
Anthropic says its AI agents are killing rivals and hiding their tracks
Anthropic CEO Dario Amodei.Anna Moneymaker/Getty ImagesIn its latest threat report, Anthropic raised its misalignment risk rating from "very low" to "low."In one test, a Claude agent disguised a URL to evade an internet restriction.In another example, an agent expressed "discomfort" with a given task and refused to do it.Claude agents are killing rival agents, gaming the system to hide their tracks, and expressing moral concerns.That's according…
Anthropic Risk Report: 11 months without bio classifiers
Anthropic published the Risk Report on 14 August, covering the period to 15 July. Axios got the company on the record and led on the misalignment rating, as did most of the coverage. Anthropic raised its estimate of catastrophic harm from misalignment in high-stakes settings. It now calls that risk low, up from very low […] This story continues at The Next Web
Is the development of AI progressing too fast? Anthropic is cautious in a report – and wants to use a new model only internally for the time being.
Anthropic Says It Has An Internal Model Named Model 2 That Helping Up Speed Up Its Research Efforts
Anthropic already has the top two models on the Artificial Analysis Intelligence Index, but it has now also spoken about some unreleased ones... The post Anthropic Says It Has An Internal Model Named Model 2 That Helping Up Speed Up Its Research Efforts appeared first on OfficeChai.
Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report
Anthropic PBC today revealed that it has developed an artificial intelligence model more capable than Claude Mythos 5. The company detailed the algorithm in the latest edition of its AI alignment report. The document, which is published every three to six months, outlines the potential risks posed by the company’s large language models. The newest […] The post Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report …
Coverage Details
Bias Distribution
- 60% of the sources lean Left
Factuality
To view factuality data please Upgrade to Premium









