4 ms·
Update: https://x.com/mrcslws/status/2089850106292162982 https://x.com/mrcslws/status/2089850106292162982 Pasted below: Finally convinced myself that non-dist
by mrcslws 1mo ago
Update: https://x.com/mrcslws/status/2089850106292162982 https://x.com/mrcslws/status/2089850106292162982
Pasted below:
Finally convinced myself that non-distorting watermarking is real, a la Anthropic / Google SynthID. It's not just spin / marketing. This really is a "Monty Hall" like problem. (h/t @random_walker for that analogy.)
Sharing here in case it helps anyone else.
On one hand, watermarking is "obviously" damaging to the output. The LLM does all this work to compute token probabilities... then you essentially perturb the probabilities? Of course that's bad! Every single token is using perturbed probabilities!
On the other hand, you obviously can freely inject a layer of randomness into sampling. Take any categorical distribution. For each sample, perturb the probabilities, then draw the sample. Done correctly, over many samples the results will match the original distribution.
It still kinda feels like magic, but it's not surprising that you can take that core trick and shape it into a non-distortionary watermarking scheme.
(i.e. it gives you sequences that were just as likely to be generated by the original model as any other sequence)