frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Understanding the Impact of LLM Watermarking on AI Agent Behavior

https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-on-ai-agent-behavior
43•nisosguy•1h ago

Comments

willmadden•57m ago
Watermarking sounds like a good idea, but it's not. Token drift from watermarking will degrade the quality of outputs and could allow clever people to circumvent guardrails.

You should assume all text is AI generated. If you want to "test" someone at school or during an interview, have them write with a pencil and paper.

skybrian•22m ago
Changing a random seed could either improve or degrade the output. In theory, better and worse outputs should be equally probable, depending on your luck.
WithinReason•53m ago
This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.
arcticbull•47m ago
Model companies are doing this for themselves anyways, it’s so they don’t feed generated content back into the slopper and collapse the model. From that angle it over time contributes to better model quality.
Phemist•15m ago
Also - it forms a cartel.

Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).

So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).

This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.

porridgeraisin•6m ago
What? It has absolutely nothing to do with "model collapse".
Phemist•30m ago
The article has a pretty decent summary of the watermarking algo though. This reads as a pretty dogmatic statement in comparison.

In your analogy: What if seed 42 specifically causes poor quality behaviour (in some contexts specifically). Normally, these quality differences will be washed out because the seed is random, now it is no longer random, so shouldnt we check into specific behaviour under this specific seed?

lemagedurage•24m ago
Klaus23•53m ago
Am I missing something, or did they actually completely misunderstand how this technology works?
cubefox•43m ago
More likely you misunderstood it than them.
fn-mote•24m ago
It’s hard to tell because the writing quality is garbage.

Second paragraph:

> Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token.

This is a stretch. True, but barely. The LLM is making slightly different choices near the end of the token generation process.

> At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection.

Claim support, if it appears, is pages later.

> At the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it.

?

> Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools.

Wtf. Non-sequitor. Where does this come from?

> Such a watermarking procedure can therefore affect both what the model says and what an agent does.

Duh? In the literal sense of outputting different tokens.

> We call this behavioral effect sampling drift.

I think they should have used an LLM for writing help, or paid more for the one they used.

zeroonetwothree•14m ago
It’s already mostly AI written.
dahart•12m ago
sinan-faizal•50m ago
watermarking is great tho
jamienk•38m ago
In _1984_ the Big Brother regime has the idea that by controlling language you can influence what is possible to think, and thus becomes a key tool of political repression.

Political Correctness has a similar idea that by adjusting the terminology we use, we can purge biases and historical implications and speak in a purer way.

Psychoanalysis has its own idea of repression - where a person struggles to BLOCK our associations between ideas, memories, and words in order to try to stop one thought from being contaminated by another, intolerable thought.

All of these attempts to control language are fundamentally misguided at best, often have severe unintended consequences, and are genuinely immoral at worse.

p_j_w•14m ago
> Political Correctness has a similar idea that by adjusting the terminology we use, we can purge biases and historical implications and speak in a purer way.

You either misunderstand political correctness or are trying to make it sound more nefarious than it is. It is nothing more than an effort, sometimes overdone and misguided, to not say things that make minorities feel bad. It’s nothing more than an attempt to broaden what’s considered good manners.

If anyone serious thinks political correctness is going to end racism, I certainly haven’t seen it.

> Psychoanalysis has its own idea of repression - where a person struggles to BLOCK our associations between ideas, memories, and words in order to try to stop one thought from being contaminated by another, intolerable thought.

It’s been a while since I took Psych 101, but my memory is that, according to Freud, repression is subconscious. There’s no struggle possible because there is no intent. It’s also about memories and emotions, so if it WAS an intentional act, it’s not an attempt to control language. It wouldn’t be about trying to not think of the word cat, but trying to not think about that time when your cat died.

samayashar•24m ago
I am unable to understand what happens if the watermarked output goes as input to another agent. Let's say we asked Claude a question and got a watermarked response. If we pick that response and append it to the question we're asking ChatGPT, then will it answer or refuse to do so?

If that's the case, then it's a brilliant strategy by the labs to cut down cross-AI usage and just stick to one model. But I'm pretty sure this won't be the case.

serbuvlad•23m ago
This article reads like it was written at least partly by AI to me. Specifically it reads like an article written by AI with edits made by a human further prompting the AI.

> Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.

Relevance and irrelevance are not introduced above this comment. This reads like an LLM-ism (particularly a GPT-ism) editing a document, removing something, and leaving a note about why it was removed, which doesn't really make sense when reading it.

> Their limited movement under prompt injection should therefore not be interpreted as evidence that watermarking preserves safety behavior more reliably on these models.

Also a GPT-ism which appears when it draws a counter-conclusion in the text because it feels the need to be honest and a human tells it to remove it because it's not true because of "reason".

Overall interesting research, however, I think it's great that model output is getting watermarked. I was skeptical of this at first, but Opus 5.5 is so good, it seems like it's a non-issue in practice.

The reason I think watermarking is great is because it's a really good way of preventing training on it's own output indiscriminately and Ouroboros-ing itself.

zeroonetwothree•14m ago
Pangram flags it as mostly AI.
skybrian•6m ago
The comments here are terrible. I got a better idea of what’s going on by asking ChatGPT what this paper’s weaknesses are:

https://chatgpt.com/s/t_6ab7d694885481918083b8cbf0ba9040

In particular: sometimes they measure “churn”, which doesn’t show whether the results are better or worse on average. They sometimes only test with one random seed. There are multiple-comparison issues. And they’re not testing Anthropic’s algorithm.

That's not true. Watermarks are messing with the next token generation probabilities based on some random seed. The quality is neccesarily lower, the difference is simply too small to notice, typically.
samsartor•16m ago
No, the probability distribution is the same. Watermarking changes the rng sequence used to pick from that distribution.
jannyfer•19m ago
“When implemented correctly” is probably what people are complaining about.

Opus 5 started adding a bunch of comments to code, even when instructed not to, and for very simple changes where the comment itself was longer than the code change. Was that so that there are enough tokens outputted for watermarking? Many people suspected so.

PunchyHamster•15m ago
So that's where that nonsense comes from...
KoolKat23•14m ago
Benchmarked output quality versus actual output quality are very different things. Some usecases are at the very fringe of model intelligence and depth of intelligence and logic suffers.
Phemist•6m ago
Ofcourse the AI corps are incentivized to downplay the effects of watermarking where they can, as they stand to gain so much from rolling it out (prevent model collapse).

Also interested in how this watermarking push makes sense when considering RSI.

docjay•3m ago
You’re using the subjective definition of “quality”, as in the shade of blue it chooses for “Build a website”, or the character names for “Tell me a story.” In those cases it’s likely still subjectively “high quality”, depending on who you ask.

What the article is discussing, and what many people are concerned about, is something that you might be missing in your understanding: they’re not actually random. In fact, they would be entirely useless for real work if every token was randomly selected based on all possible outputs. It’s not, even at temperature 1.0. It’s based on the training corpus and once you have your tool names, syntax, and prompting style aligned with the training data then they become incredibly deterministic in the areas that matter, such as tool calling and parameters. I build my toolset by testing thousands of names, syntax, return format, and other aspects until I find a convention that produces the exact correct call, 100% the exact same every time, regardless of context length. Those decisions are per model and what works with Opus 4.7 won’t necessarily work on 4.8, and neither version will work with a local model or GPT.

That’s only possible because the massive training corpus is the guiding principle behind the choices. Providing a file reading function called “Read_The_File” will fail, either on the first call or somewhere down the line, because that name is not associated with the concept. Your instructions are trying to override 500 trillion tokens from training and it will cause perplexity to manifest as wrong tool calls, wrong syntax, “oops deleted prod”, “Claude lost the plot again”, “WTF?!”, and probably nearly every frustration you’ve encountered and determined to be “they nerfed Claude” or “it’s a dumbass.”

For those that are aware of it, that knowledge lets people tweak and tune the prompts/tools accordingly.

You may not put that effort into your system, perhaps because you’re unaware of it, don’t use it in a way that requires it, or you’ve just taken the failures caused by perplexity as something that’s inherent in the framework, but for people that build precision infrastructure around them it’s potentially devastating news. Watermarking, which is based on whatever tokens, threshold, cutoff, and triggers some guy at a desk decided, will necessarily alter that entire system.

> I think they should have used an LLM

The article reads to me as largely LLM generated. I checked several sections with Pangram and it appears to agree that it’s either entirely AI generated, or at least 100% AI assisted.

I’m fine with the idea of AI assistance, but I agree the quality is low and this needed a bit more editing. I have my own list of gripes that include not defining terms that are used, text for tables and figures being repeated in the body and the captions, contradictions and non-sequiturs, and just AI’s real watermark of getting lost in details, using clever terminology, having a hard time being concise, and being unable to use plain and clear language.

I think they goofed on the conclusion: “These results do not argue against watermarking for provenance. They show that provenance and behavioral stability are separate properties.”

The article’s entire point is that stability depends on provenance techniques, therefore the results purport to show they are not separate properties. Oops.

Also, it’s not AI’s fault that URLs in the references section aren’t clickable. Come on.

porridgeraisin•4m ago
The article as well as most of the comments thuss far are a dumpster fire

Breaking Up with Google Play: Why Conversations Is Now Free

https://gultsch.de/posts/breaking-up-with-google-play/
283•ezst•3h ago•112 comments

Understanding the Impact of LLM Watermarking on AI Agent Behavior

https://www.lasso.security/blog/the-provenance-tax-understanding-the-impact-of-llm-watermarking-o...
43•nisosguy•1h ago•31 comments

Fifteen years later, the Apple Cards origin story

https://lexontech.org/fifteen-years-later-the-apple-cards-origin-story
154•ksec•5h ago•18 comments

Revealing the details of how OpenAI agents hacked Hugging Face

https://swarmtraces.org/
569•specked-citrus•17h ago•366 comments

We're gonna need a lot more mathematicians

https://terrytao.wordpress.com/2026/09/24/were-gonna-need-a-lot-more-mathematicians/
214•srcreigh•11h ago•298 comments

Plan mode is dead

https://www.aymannadeem.com/artificial/intelligence,/developer/tools/2026/09/24/plan-mode-is-dead...
414•jmvldz•1d ago•377 comments

Modern Object Pascal Introduction for Programmers – Castle Game Engine

https://castle-engine.io/modern_pascal
23•birdculture•2d ago•8 comments

Ollaya – Ollama for open-source, Jev-style decision models

https://ollaya.dev/
515•Ardakilic•20h ago•126 comments

Floci: Locally emulating any cloud service

https://floci.io
79•theanonymousone•6h ago•12 comments

Reflections on 1,000 Days of Math

https://gmays.com/reflections-on-1000-days-of-math/
5•gmays•3d ago•0 comments

A single function Jev-like wrapper for LLMs, including vision models

http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html
107•allanrbo•10h ago•30 comments

Is your Postgres migration safe or not safe?

https://safenotsafe.dev/
73•vira28•7h ago•21 comments

16GB iPod Nano 3G Upgrade

https://tuckerosman.com/projects/16gb-ipod-nano
74•Ivoah•2d ago•8 comments

Show HN: Jev Plays Pokémon Red

https://jev-pokemon.vercel.app/
223•pancomplex•1d ago•91 comments

ASML currently sells no chipmaking machines in Europe, executive says

https://nltimes.nl/2026/09/22/asml-currently-sells-chipmaking-machines-europe-executive-says
82•doener•2d ago•77 comments

Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML

https://jointhefreeworld.org/blog/articles/lisps/parsing-expression-grammar-lisp-org-convert-to-h...
52•jjba23•2d ago•10 comments

What even is an OS now?

https://sockpuppet.org/blog/2026/09/25/what-even-is-an-os-now/
227•fratellobigio•17h ago•330 comments

Calculating atmospheric drag on satellites for a Cubesat [pdf]

https://www.osti.gov/servlets/purl/1124870
14•walrus01•2d ago•6 comments

Ask HN: Who's still keeping a DOS machine up because the business depends on it?

177•mlaux•19h ago•167 comments

Gravity seems holographic. What does that mean for reality?

https://www.quantamagazine.org/gravity-seems-holographic-what-does-that-mean-for-reality-20260925/
237•ibobev•23h ago•186 comments

The Murky History of Soviet-Born Tetris

https://thereader.mitpress.mit.edu/the-bizarre-murky-history-of-soviet-born-tetris/
68•EA-3167•1d ago•15 comments

Earth is tearing apart beneath the Pacific Northwest

https://www.sciencedaily.com/releases/2026/09/260924231343.htm
3•emot•9m ago•0 comments

Jury finds Facebook liable for deceiving users in Cambridge Analytica case

https://www.cbsnews.com/news/facebook-liable-deceiving-users-cambridge-analytica/
312•pseudolus•13h ago•81 comments

Scientists build most accurate atomic clock

https://phys.org/news/2026-09-scientists-world-accurate-atomic-clock.html
34•wglb•2d ago•19 comments

Fourier Analysis: Drawing Llamas with Circles

https://adekau.github.io/posts/2020/llamas.html
59•cebert•2d ago•6 comments

The Copilot+ PC brand is dead

https://www.windowscentral.com/microsoft/windows-11/the-copilot-pc-brand-is-dead-microsoft-and-pc...
44•bj-rn•4h ago•26 comments

Excel now supports multiple values in a single cell

https://techcommunity.microsoft.com/blog/microsoft365insiderblog/put-multiple-values-in-one-cell-...
220•luispa•17h ago•153 comments

The far side of the Moon provides clues to a previous magnetic field

https://ethz.ch/en/news-and-events/eth-news/news/2026/09/the-far-side-of-the-moon-provides-clues-...
17•croes•6h ago•1 comments

One Month Without AI

https://blog.bustikiller.com/2026/09/25/one-month-without-ai.html
125•saibotk•4h ago•129 comments

From Thin Air to Bootable Images: The Tine Build System

https://amutable.com/blog/tine-build-system
26•Levitating•2d ago•8 comments