4 ms·
I've heard that this kind of watermarking process works by biassing the statistical sampling towards a partition of the set of possible next tokens (red set and
by benrow 2mo ago
I've heard that this kind of watermarking process works by biassing the statistical sampling towards a partition of the set of possible next tokens (red set and green set), at each position. It might only be a slight nudge each time, but over a sequence of tokens, the likelihood of repeating the bias by chance is increasingly improbable.
The bias is different for each position and follows a defined RNG, seeded somehow predictably.
Can be either an open algorithm, or not. If not open, then an API could be provided to determine if text is watermarked or not.
How it applies to code - maybe it could be a subtle nudge to symbol names, etc, I'm just speculating (I only read about this in passing very recently).
- cassianoleal 2mo ago> a defined RNG, seeded somehow predictably So, an NG?
- olmo23 2mo agoPRNG
- m-chrzan 2mo agoThere's a computerphile video (https://www.youtube.com/watch?v=XZJc1p6RE78 https://www.youtube.com/watch?v=XZJc1p6RE78) with Dr. Mark Pound explaining a paper by John Kirchenbauer, Jonas Geiping et al. (https://arxiv.org/abs/2301.10226 https://arxiv.org/abs/2301.10226) that described a method for watermarking LLM output like this. It's not directly stated anywhere in the Claude support article that this is what they're using, but the properties of the watermark described seem to point to this method.
- IshKebab 2mo agoIf it's based on position mod 2, wouldn't inserting or deleting (or splitting/merging) words every now and then trivially defeat it? If it is based on position mod 2 then wouldn't inserting/deleting (or splitting and merging) words every now and then defeat it?
- metalcrow 2mo agoBased on my understanding, it can only be applied to code in very limited ways: docstrings, variable names, string literals. The code itself can't really have tokens changed to another equally correct token (the foundation of the watermark) because then the code breaks! And the few places that you can do so are likely erased by formatters anyway.
- matherial 2mo agoYou have less latitude than in prose, but I think you're underestimating how much can be changed without causing breakage. For example, the most likely sequence might be "if (!foo) bar; else baz;", but you can also say "if (foo) baz; else bar"; substitute "foo == 0" / "foo != 0" for even more variety. Similarly, "foo = 1; <NL> bar = 2;" can be output as "bar = 2; <NL> foo = 1;". Keep in mind that the LLM "sees" the previous (tweaked) output and picks what makes sense based on that. There are few situations where a perturbation like that would be unrecoverable, and I assume these situations also correspond to a huge probability difference between the most likely completion and the second most likely one - in which case, the watermarking algorithm can choose not to touch the token.
- londons_explore 2mo agoThe bias has to be small enough that if you ask an LLM to repeat some passage of text like the national anthem, either from the training data or from the prompt it doesn't change random words. Gotta be hard to tune that.
- deleted 2mo ago[deleted]