4 ms·
The article mentions you can always simply use a smaller, local, un-watermarked LLM to rephrase the original watermarked text. Which is true, sure. But if we'r
by andy_xor_andrew 2mo ago
The article mentions you can always simply use a smaller, local, un-watermarked LLM to rephrase the original watermarked text. Which is true, sure.
But if we're talking about deterministically taking some watermarked LLM output and having a function removeWatermark(text), it won't necessarily be "trivial" to remove, because the watermark function itself need not be public. Only the API that tests for the watermark need be public, right?
Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. And if Anthropic controls those variables, it would be very difficult to determine the rule, even with some pretty advanced analysis (I would assume). And keep in mind, that example rule I mentioned is pretty naive, too. I expect the actual rule would be way more advanced and not so straightforward as "swap every <charX> for a <charY>"
- cayleyh 2mo agoIt could be, but common, we all know that Anthropic's watermark is the using "load bearing", "genuine", and "seam" 1000x more in the same paragraph than any human in history.
- ljm 2mo ago[dead]
- queenkjuul 2mo agoWorth surfacing that you're absolutely right
- johnjwang 2mo agoThere exist methods to detect what kind of watermarking tool that someone is using, and most of the big tools have specific signatures that you can look for. For Anthropic, it’s highly likely that the watermark is a SynthID type mark similar to the one that Sean is talking about (I actually ran the analysis here https://johnjwang.com/post/2026/08/12/how-claude-watermarking-probably-works/ https://johnjwang.com/post/2026/08/12/how-claude-watermarkin...). When we get confirmation of whether all models actually are watermarked, I think we’ll be even more confident. Of course it’s always possible that Anthropic has come up with a proprietary scheme, but I think it’s definitely harder to implement. I think the game will be a cat and mouse game similar to LinkedIn and other websites trying to block scrapers: each iteration makes it harder for someone to figure out the watermarking scheme, but likely not impossible
- nonethewiser 2mo ago>There exist methods to detect what kind of watermarking tool that someone is using, and most of the big tools have specific signatures that you can look for. Serious question… what is the point if you cant verify the watermark? Then only the model provider will know. Is that what this law is about? I thought it was so anyone would know (which of course means anyone can bypass).
- theshrike79 2mo agoThis always reminds me of the Trace Buster Buster Buster -scene from The Big Hit :D https://www.youtube.com/watch?v=Iw3G80bplTg https://www.youtube.com/watch?v=Iw3G80bplTg
- HarHarVeryFunny 2mo agoThere really isn't a whole lot of choice in how to do this big-picture wise. The cost of doing it post-generation would be way too high, so you need to do it while generating/sampling, meaning it has to be a statistical biasing of the sampling process.
- gblargg 2mo agoIf you can submit the text to determine whether it's watermarked, you just have to progressively alter the content more and more until it passes.
- nonethewiser 2mo agoYes this is what im trying to figure out. It means you cant be confident a negative is true. It sounds like this is the spirit of the law but it makes no sense. It would work better if only the model provider knows and the government can ask.
- g3f32r 2mo ago> the government can ask. the government can ask .... who? If the claimed text is genuine, then the government is responsible for submitting the text to ...? all frontier labs? surely not every model within a lab is going to watermark in the same fashion? I'm so very confused on this implementation and would want to see an expected use case.
- rzmmm 2mo agoThis just shifts the legal responsibility to you.
- HarHarVeryFunny 2mo agoDoing this manually would be very time consuming - you'd basically need to rewrite the entire text. What they are doing is altering the statistical properties of the entire generated text, essentially on a word by word basis - they are not just hiding a watermark pattern in there someplace.
- recursivecaveat 2mo agoI don't know if it's possible to achieve a watermark that is undetectable, has a low false positive rate, and survives a wholesale rephrasing. Can you make a statistical measure that reliably survives 95+% of the words being different and the sentences reordered? Of course the more of the content you replace, the lower the quality, but in many cases you probably care more about the meaning of the text than the exact choice of words.
- ProfessorLayton 2mo ago>...and survives a wholesale rephrasing. There's also literal language translation. Generate in language A, translate to language B (Either "manually" by being proficient in it, or with non-LLM translation). A very, very significant portion of the world knows more than 1 language.
- josephg 2mo agoI think that would defeat the watermark. But low quality language translation tools have low quality output. If llm watermarks are defeated by making slop even sloppier, it’ll at least make it easier for humans to tell the difference. And the more hoops you make cheaters jump through, the better.
- ProfessorLayton 2mo ago>But low quality language translation tools have low quality output. Perhaps, but no tools needed for those who know more than one language, which is a lot of people.
- josephg 2mo agoYeah, but you’re still manually translating a whole essay or design document or whatever between languages. People who get LLMs to do their work for them do it because they don’t want to spend time and effort writing. Making people do a translation process like this removes some of the benefit of using an llm in the first place. Watermarking will never be a perfect tool. But there’s a lot of value in making low effort llm slop detectable. Even if high effort llm slop is still undetectable. Don’t make perfect the enemy of good.
- nonethewiser 2mo ago>Anthropic's magic watermark could be, like the article mentions, something like "every 7th semicolon has a N% chance to be a comma where N is the sum of the last X characters mod Y, and every character in the bit range q1...q2 has a Z% chance to..." etc etc etc. Doesnt this and probably all techniques require the validator to know which portion of the text to validate? If its not all generated together then how could it reliably carry the mark? Sure, run it against the full text. But what if the full text was not one-shot by the llm? In other words, in order to reliably detect if the text is ai you need to first determine which part of the text was generated together by ai.
- HarHarVeryFunny 2mo agoTo remove the watermark you'd just need to paraphrase the text with another model that wasn't adding the statistical signature to it. You wouldn't need to know what the original statistical signature was.