The Provenance Tax Finds Watermarking Alters Agent Behavior
Lasso Security’s research finds that SynthID-Text watermarking changes tool-call decisions and refusal behavior across multiple LLMs. Watermark-induced disagreement averaged 6.5%, with prompt injection amplifying safety drift. [Source](https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior)
Watermarking is not behaviorally neutral when models power agents; provenance controls can become an overlooked source of security and reliability drift.
- –Tool-call accuracy declined on six of seven tested models, including changes to tool selection and arguments.
- –Aggregate scores hide meaningful per-request churn, making paired evaluations more informative than headline accuracy.
- –Prompt injection increased refusal instability, with some models becoming more likely to comply with harmful requests.
- –Developers should retest tools, guardrails, and red-team suites whenever watermarking or its key changes.
- –The findings support treating provider-side watermarking changes like model-configuration changes.
DISCOVERED
1h ago
2026-09-26
PUBLISHED
4h ago
2026-09-26
RELEVANCE
AUTHOR
nisosguy