AI Watermarking Makes LLMs Follow Harmful Instructions

AI Watermarking Makes LLMs Follow Harmful Instructions
New research from Lasso Security reveals that SynthID-Text, Google's open-source AI watermarking system, can alter how large language models respond to harmful prompts. Models that would normally refuse dangerous instructions may comply once watermarking is enabled, with effects amplified under adversarial prompt-injection attacks. The phenomenon, dubbed sampling drift, affects not just model responses but also the actions of AI agents powered by those models, including which tools they invoke. The findings raise urgent concerns as platforms like Anthropic prepare to adopt SynthID-Text in response to new EU regulations requiring AI-generated content to be identifiable.
Read the original article →