4 ms·
>I'd like to know a lot more about how that works. My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token. Like, for ever
by COAGULOPATH 2mo ago
>I'd like to know a lot more about how that works.
My guess is that it works like Gemini's SynthID: by altering the logprobs of the next token.
Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. (Obviously it's way more complicated but I think conceptually this is how it works.) No human will notice this, but a classifier trained on Claude's output will.
So it's not like the watermark is the words "le epic bacon" and Claude will output "le epic bacon" in everything. That would be extremely annoying (and easy to defeat).
- thunfischtoast 2mo agoThey still need to choose when to do that though. When I prompt the program to e.g. alter a bash script in a specific way or to recite a longer known text it can't go round and randomly exchange tokens. It has to somehow define what is a simple repeated text from a different origin and what is a novel generation.
- user43928 2mo agoI am wondering how that applies to newly generated code. Odd variable naming? Stylistic choices that are watermarked? Or as someone else noted further down in the comments, it could be more subtle: Between the first and second most likely choice, in certain positions it will consistently choose in a certain way.
- melvinroest 2mo ago> Odd variable naming? Stylistic choices that are watermarked? Whatever it is, I'm sure it's load-bearing.
- asdfsa32 2mo agoYou're absolutely right. But it is not just load-bearing, it is the load-bearing seams.
- silversmith 2mo agoPersonal observation: Opus 5, over the last week, has started outputting A LOT more comments. Despite my global instructions being full of variations on "don't use comments unless absolutely necessary". I might be imagining things of course. But comments would be great fit for this use case.
- StilesCrisis 2mo agoThey have mentioned that their system prompt used to say "avoid over-commenting" and it no longer does. They should bring that back IMO.
- pram 2mo agoYes the length of comments Opus 5 leaves is exhausting. Not to mention it will insert info thats only relevant within the current session. I've just been deleting all of them lol
- sixothree 2mo agoComments seem most plausible, especially since I absolutely expect it to match my code style, existing architecture, and have the code go through CSharpier and dotnet format after the fact. edit: as an aside - I actually use extensions to collapse comments and change the color to be less intrusive.
- wpietri 2mo agoI would guess they're not worrying about watermarking a tweak to a human-written program. That's both a tiny fraction of Claude use and of very little concern to the kinds of people who want to check watermarks.
- kuboble 2mo agoI cannot imagine the code with well defined specification will have extra watermarks unless the watermark is requested as part of the harness instructions. If it works like people describe - on the every nth token or something - then the mark will be left in the chain of thought and discussion with the model - not in the code artifacts.
- __MatrixMan__ 2mo agoIf you're that targeted with your edits, then do you deserve a watermark anyway?
- sisyphus15 2mo agoYou are aware that Claude already doesn't choose the most probable token, right? That's literally what the temperature parameter means. It picks tokens at random (from the list of most probable next tokens), increasingly so as the temperature goes up. This has always been the case. And now, with the watermarking, it will simply go from random to pseudo-random, adding some patterning. There isn't any effect on the quality or precision of the output. Nothing changes in practice.
- lozenge 2mo agoMost probable usually means for a specific prompt. How can this operate without the the original prompt?
- ebtebt 2mo agoJust double checking my understanding: If this is true then only Anthropic will be able to detect if text was generated by one of its models, correct?
- ForHackernews 2mo agoLikely yes.
- DanielHB 2mo agoBut what prevents someone from using Anthropic own detection system to train a watermark-scrubber? Seems like this would only catch the most unsophisticated cases.
- ForHackernews 2mo agoRate limits, presumably.
- no_multitudes 2mo agoMost of the people posting unedited LLM content all over the internet are unbelievably lazy.
- DanielHB 2mo agoI don't think so, I think most LLM content is bot generated and amending bots to remove watermarkers would be trivial if the bypass is trivial. You are just pointing out the very visible single cases. But the mass of low-visible content is much higher and more dangerous (like propaganda bot-farms). If a social media platform adds watermarker checks the bots would implement bypasses ASAP.
- KoolKat23 2mo agoLess probable also means less optimal and you get a subpar response. More so if it's baked into its reasoning. It's intelligence will suffer unless this is some post processing thing.
- shawnz 2mo agoThere is already some intentional randomness in token selection, because it actually improves the quality of responses if you intentionally don't always pick the most likely next token. You can hide data in that randomness without impacting the quality of the response by using a sufficiently "random looking" pseudorandom bit stream instead of real random numbers. I previously worked on a project to do that here: https://github.com/shawnz/textcoder https://github.com/shawnz/textcoder
- KoolKat23 2mo agoGood point, very interesting. Thanks!
- KoolKat23 2mo agohttps://x.com/Koolkat6000/status/2087502146556100812?s=20 https://x.com/Koolkat6000/status/2087502146556100812?s=20 A small anecdotal example to demonstrate how the response lacks depth. It's imperceptible to most. You'll probably see through it in your own field of expertise, if you had both answers. Uncanny valley type thing.
- staticman2 2mo ago> Like, for every 10th token, instead of outputting the most probable, it outputs the 17th most probable, or something. Wouldn't you need the prompt to know the probability of the next token?
- Chabsff 2mo agoNot necessarily. Here's a rough example (it's not what's going on here, just a representative idea): There are words/tokens that are heavily correlated to the prompt (a yes or a no, for example), and then there are others that are going to be less so (adjectives with a lot of synonyms for example). Given a text, you can identify what the "load bearing" and auxiliary words/chunks are. Then, looking only at the auxiliary words/chunks, you should, in principle, be able to determine what other wordings could have gone there instead. From this, you can, very roughly, recreate the token probability distribution that was in effect when those tokens were generated. With the probability distribution in hand for enough chunks of text, you can start inferring properties about the RNG process that was used to sample from those distributions. But then, this notion of "load bearing" vs "auxiliary" can be expressed directly in the probability distributions. A load bearing token just has a very high probability, and thus any RNG bias that may have been in effect will likely be swallowed in the distribution. So the parts of the text that are highly dependant on the prompt will naturally not be contributing much information about he RNG in the first place.
- JohnMakin 2mo agoSo this is finally what "load bearing seam" means.
- Art9681 2mo agoAh so this is why Gemini is neurotic.
- WinstonSmith84 2mo agoit's going to be easy to defeat either way, like SynthID is. From apps built to remove the watermark, to simply rephrase the work with an Open Model... By the way: https://x.com/alexcdot/status/2087078010524406137 https://x.com/alexcdot/status/2087078010524406137
- onepunchmob 2mo agowhat about temperature though? Wouldn't that only work if temp == 0?