2 ms·
I think you're correct. It does alter the distribution for each output token, implicitly giving each candidate token a different probability. Maybe that's fine,
by mrcslws 2mo ago
I think you're correct. It does alter the distribution for each output token, implicitly giving each candidate token a different probability. Maybe that's fine, but it's not as magical as Anthropic [and a lot of commenters here] are making it out to be.
- mrcslws 2mo agoUpdate: https://x.com/mrcslws/status/2089850106292162982 https://x.com/mrcslws/status/2089850106292162982 Pasted below: Finally convinced myself that non-distorting watermarking is real, a la Anthropic / Google SynthID. It's not just spin / marketing. This really is a "Monty Hall" like problem. (h/t @random_walker for that analogy.) Sharing here in case it helps anyone else. On one hand, watermarking is "obviously" damaging to the output. The LLM does all this work to compute token probabilities... then you essentially perturb the probabilities? Of course that's bad! Every single token is using perturbed probabilities! On the other hand, you obviously can freely inject a layer of randomness into sampling. Take any categorical distribution. For each sample, perturb the probabilities, then draw the sample. Done correctly, over many samples the results will match the original distribution. It still kinda feels like magic, but it's not surprising that you can take that core trick and shape it into a non-distortionary watermarking scheme. (i.e. it gives you sequences that were just as likely to be generated by the original model as any other sequence)