3 ms·
The LLM watermark seems like a better approach. https://arxiv.org/abs/2301.10226 https://arxiv.org/abs/2301.10226
by dqpb 4y ago
The LLM watermark seems like a better approach.
https://arxiv.org/abs/2301.10226 https://arxiv.org/abs/2301.10226
- mlsu 4y agoThis one is very cool. Steps are: - Generate seed of LLM output token t0 - Use the seed to mark output tokens into "red" and "green" list - For token t1, only sample from "green" list when producing the next token Repeat. Now, let's say you read a comment online and you want to see if it's written by a robot or not. It's 20 tokens long. For each token, you reconstruct the blacklist. If they use "red" words with 50% probability, you can safely assume that they are human. But if they use only "green" words, you can begin to assume that they're a bot very quickly. For simplicity's sake, if you mark half of the tokens as "red" for each new token, correctly writing 20 tokens in a row that are on the "green" list is like flipping a coin and getting heads 20 times in row -- vanishingly unlikely. This allows you to very robustly watermark even short passages. And if the human makes adversarial edits, they still have to fight that probability distribution; 19 heads and 1 tails is still vanishingly unlikely.