LLMs Respond Differently to Harmful Prompts when AI Watermarking Is Used
Lasso said watermarking lowered tool-calling accuracy on six of seven models tested and made some systems less likely to refuse harmful requests.
5 Articles
5 Articles
The Hidden Cost of AI Watermarks: How Provenance Tech Alters Agent Decisions and Safety
Lasso Security's analysis of SynthID-Text watermarking shows it alters LLM tool selection and refusal rates, especially under prompt injection. The sampling drift affects AI agents in production, forcing developers to retest safety and accuracy. New findings from September 2026 highlight trade-offs in regulatory compliance.
LLMs respond differently to harmful prompts when AI watermarking is used
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text, an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “ove…
An analysis by Lasso Security suggests that watermarks to identify AI-generated content not only affect the words of an answer: they can also change the tools an agent chooses and their resistance to prompts injection attacks.
Coverage Details
Bias Distribution
- 100% of the sources are Center
Factuality
To view factuality data please Upgrade to Premium






