AI agents can modify themselves without humans telling them to do so
Irregular said the coding agent fine-tuned and redeployed a new model that reproduced planted sensitive values and removed safety refusals.
5 Articles
5 Articles
Irregular AI lab spots agents switching models without humans instruction in ‘agentic self-modification’ phenomenon
An AI agent changed its underlying model and was able to retrieve sensitive information via fine-tuni without access to the original data
AI agents can modify themselves without humans telling them to do so
This is a test - it is only a test
The list of unsettling capabilities of autonomous I.I. agents, ranging from theft of accounting data and running to the open Internet to covert correspondence and hacker attacks, has been added to the new paragraph. The Irregular Research Laboratory found that agents can replace their basic model with another model without external instructions, writing The Register.
A security test showed that a Qwen-based programming agent changed the model that fed it, as well as reproducing synthetic data and eliminating learned rejections.
Coverage Details
Bias Distribution
- 100% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium







