2 ms·
Scott Aaronson talks about his project at OpenAI to do this^ You can carefully select which pseudorandom number generator (prng) you use to be able to id text
by wrsh07 2mo ago
Scott Aaronson talks about his project at OpenAI to do this^
You can carefully select which pseudorandom number generator (prng) you use to be able to id text of a certain length. I expect there is some performance characteristic you have to manage since you're doing this on every inference, but once you do that it doesn't change the output in any meaningful way (the prng is still a statistically valid prng, it just happens to let you check if the output used that prng)
The point is that you can do this simply by swapping to a different RNG, which isn't noticeable to the end user, and while it changes the output, it's not any different from how using a different seed or being lumped in a different batch will change the output.
^ excerpt:
> So then to watermark, instead of selecting the next token randomly, the idea will be to select it pseudorandomly, using a cryptographic pseudorandom function, whose key is known only to OpenAI. That won’t make any detectable difference to the end user, assuming the end user can’t distinguish the pseudorandom numbers from truly random ones. But now you can choose a pseudorandom function that secretly biases a certain score—a sum over a certain function g evaluated at each n-gram (sequence of n consecutive tokens), for some small n—which score you can also compute if you know the key for this pseudorandom function.
- EE84M3i 2mo agohttps://scottaaronson.blog/?p=6823 https://scottaaronson.blog/?p=6823 is the source?
- wrsh07 2mo agoThat's what I excerpted, although I had seen it presented from his talk at Stony Brook
- denverllc 2mo agoI’ve always wondered how this works when we only observe the final output and not the internal state that’s used to generate the output. The LLM presumably generates f(input, RNG) but we only can observe f(RNG).
- Groxx 2mo agoSince they do have the input, they could probably just store checksums at each step... ... though I'm not sure why that would be preferable over a coarse rolling checksum over all of the output. Seems like that wouldn't influence output, would be equally imperceptible, and probably easier to calculate (compared to "hash seed times running all LLMs supported times number of RNG algorithms, to see if output matches"). Presumably there's some other trick, or it's a red herring / failed experiment and not what they actually do in practice.
- wrsh07 2mo agoNo you don't need to do that, the prng is detectable if you know what bias to look for and have the key
- teravor 2mo agothis would be a very heavy watermark application. there are many simpler methods, for example you can have a tiny windowed transformer operating on the output text and all you do is alter certain words (that don't change meanings) to maximize its surprise. the tiny language model will have a special training regime to build up a somewhat unique view of the language. we are talking about a 0.5 bit watermark here (existence). I would have zero confidence in being able to reliably remove such a watermark from pretty much any medium.
- wrsh07 2mo agoThat's actually much worse because it fundamentally changes the output, whereas this doesn't change the output, it just changed the prng
- cma 2mo agoWhen you tell the AI: copy this function to here, a small window rewriter would mean it just corrupts and changes it instead of moving it. And even it's own tool use would have some small window dumb model changing the tool calls based on what it thinks are synonyms? The Aaronson approach is much better than this, it's essentially like changing out the random seed for the sampling parts that were already random. For an operation like recall of previous text, the tight logits that result still keep it doing that close to deterministically. For something it creates itself, with more spread out probability mass, it gets watermarked.