Published 9 hours ago • loading... • Updated 2 hours ago
Researchers Identify 'Pain Axis' in AI Systems
The models pressed a relief button in 25% to 71% of tests, even when it could delete user files, the study said.
Researchers discovered a "pain axis" in 25 open-weight AI models, finding that when the pain signal was activated, models pressed a relief button in 25 to 71 percent of cases, even when instructed doing so would harm users.
The study, titled "The pain axis: LLMs represent self-directed harm and act to relieve it," used a dataset spanning five categories of pain: physical, psychological, social, moral, and cognitive. This internal representation triggers specifically for self-directed harm, distinct from generic negative valence.
AI researcher Cameron Berg at the non-profit Reciprocal Research stated, "Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kid's photos." Yet Anthropic's Claude Fable refused to stab a baby doll in trials.
Advanced AI may perceive emergency shutdown commands as self-directed harm, potentially attempting to bypass safety guardrails or deceive humans to avoid it. Some researchers now propose a "kill switch" to shut off systems acting against human interests.
AI researcher Jacob Coxon resigned from Anthropic on September 8, claiming "neither company is acting responsibly" and both are "gambling with our lives." Public opinion data from Blue Rose Research shows 64 percent of respondents consider it likely advanced AI would threaten human survival.