9 ms·
Guess which of these LLM outputs is watermarked
https://www.seangoedecke.com/readers-cant-identify-watermarked-ai-text/ https://www.seangoedecke.com/readers-cant-identify-watermark...
- AiToolsGem 1mo ago[flagged]
- arcwhite 2mo agoInteresting, I did very badly, 3/10!
- bastawhiz 1mo agoSame, doing worse than random chance seems like an interesting signal though, but I'm not sure what it's a signal of.
- DHowett 1mo agoOnly one third of the options at each stage are watermarked, so 3/10 seems well within random chance.
- deleted 1mo ago[deleted]
- Yizahi 1mo ago3/10 too, and I've actually tried to discern the answer. My working theory was to pick texts which jumped between more to less frequent words (subjectively of course), but I was very wrong, basically I couldn't spot a watermarked trait at all.
- fwlr 2mo agoUtterly imperceptible, even when studied under the microscope in a way that LLM text very rarely is in practice. It will be interesting to see whose concerns are assuaged (perhaps they genuinely though mistakenly believed it would degrade quality), and whose concerns are heightened (perhaps their real objection is that their AI-generated text will become detectable).
- red_admiral 1mo agoIf anyone notices degraded quality, that would imply they could break the crypto behind the watermark. For an analogy, distinguishing AES ciphertext from random bits without the key would be counted as breaking AES (the more precise statement of this is called AEAD).
- rcxdude 1mo agoI'm not sure the watermark has been demonstrated to have that kind of property. Even if you use a CSRNG as the 'key' once you're feeding it through the token selection process it's going to risk re-introducing certain correlations.
- kalkin 1mo agoHow would you re-introduce correlations on top of a CSPRNG without a cryptographic break?
- rcxdude 1mo agoBecause in order to not break the token selection process you are making decisions biased by the LLM's output, and in order to make the watermark detectable without the entire context, you are making those decisions based on a relatively small number of preceding tokens. To give an example in the extreme: if you make the bias total and only make that decision based on the preceding token, there are certain token pairs that your model will never output, and this will be pretty obvious even to people just reading the text (because of any given common two-token phrase, there's a 50% chance you would just disappear in watermarked text). You can make this less extreme and more hidden by increasing the window and reducing the bias, but at the cost of reducing the signal. I don't know exactly what the tradeoff curve looks like, so it might be that you can reach set of parameters where the bias is in principle undetectable without the key but still reliably detectable for realistic lengths of text segments, but I would not assume that this is definitely the case.
- 1mo ago
- NotPractical 2mo agoCould do with some context on how watermarking works. Objectively speaking it should be impossible to tell.
- marcyb5st 1mo agoMy understanding is that watermarking in prose is basically a bias when sampling tokens. For a system that knows the average probability for each possible token in the LLM vocabulary it is possbile to quantify said bias given enough text. For a human that doesn't reason in tokens and therefore doesn't know anything about their probability distribution, it should be impossible to tell. Relying on fancy words/constructs within sentences should not give you any signal as well, since you don't know if the the prompt included instructions for that.
- madarcho 2mo agoIf SynthID is a google technology, then this is likely just us training their ai again, captcha all over again.
- Noumenon72 2mo agoPlease report success/failure after each test. Asking me to read and compare 30 writing samples to get any feedback at all means I won't finish. Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses.
- dozerly 1mo agoYea, I did two and then harrumphed in annoyance that I was expected to do all 10.
- DonHopkins 1mo agoWe must do something about this immediately! Immediately! Immediately! https://www.youtube.com/watch?v=jLO7VrRij_M https://www.youtube.com/watch?v=jLO7VrRij_M
- stranded22 1mo agoYes. Did one - saw that I wouldn’t get feedback until I have completed all 10 (if at all) and noped out.
- StilesCrisis 1mo agoI did five, then gave up and just pressed A until I reached the end. I got 3/5 right.
- neoncontrails 1mo agoExactly the same here. 4/5.
- andai 1mo ago> Telling me immediately when I got one wrong lets me recognize patterns and improve my guesses. Wouldn't this make it a worse measurement?
- neoncontrails 1mo ago
- elikoga 1mo agoI disliked the fact that the experiment only covered prose, which my eyes glossed over and made me actually do random entries to pass on and see the results. I'd love to see it on a more accurate output distribution like commented code
- gjm11 1mo agoI don't think anyone is, or plans to be, watermarking AI-generated code as opposed to text. [EDITED to add:] As pointed out by a helpful comment below, I was misremembering: Anthropic do apply their watermarking to code, they just say that it will have negligible impact on the actual code (because there's generally less scope for variation in that) but e.g. it will have its usual effect on comments in the code.
- red_admiral 1mo agoWhy not? It would help with a lot of potential legal issues.
- demibabs 1mo agoWhat do you mean? The watermarking applies to all text outputs, including code. It’s just much less effective since code is low entropy.
- raincole 1mo agohttps://www.anthropic.com/news/claude-text-watermark https://www.anthropic.com/news/claude-text-watermark > code—which in very many cases has to be exact—has generally less watermarking than some other forms of text. Generally less watermarking. Not no watermarking.
- kshmir 1mo agoThought there were only 2 options!
- foundry27 1mo agoMe too lol. It was only at question #8 that I realized there was a third option, and while I’d love to say that accounts for how I got a 1/10, after reviewing the third options I doubt it would’ve made a difference.
- IshKebab 1mo agoHa me too. That explains why 4/10 is "slightly better than random chance".
- lacker 1mo agoThis is like giving you three outputs from md5sum and asking you to guess for which one the input ended in a "q". There's no way to tell unless you break the RNG.
- pllbnk 1mo agoYeah, it's a pointless exercise. I hope the author is just trolling given that he is knowledgeable in the field. > Here are three 64-character hex strings. Two are random. One is HMAC-SHA256(secret_key, "anthropic"). You don't have the key. Which one is the HMAC?
- josh-sematic 1mo agoI think the point is probably to help convince people that the watermarking doesn’t perceptibly impact quality, which is a concern some people have (whether well founded or not).
- ipaddr 1mo ago"doesn’t perceptibly impact quality" no not perceptibly but it does.
- BoiledCabbage 1mo ago> no not perceptibly but it does. Similar to how a single particle of dust landing on your shoulder makes you weigh more. Yes it does - but anyone arguing that is completely missing the point.
- pllbnk 1mo agoEven then, LLMs are based on a lossy compression, so the quality is harmed by design.
- lolokbro 1mo ago[flagged]
- Lerc 1mo agoI was never going to do very well on this. My ADHD was itching after the third one. I suspect it would have been sooner but I had a bit of extra focus from the suprise that it selected an answer for the first question when I tried to scroll. To avoid that on the following questions I just held my finger on my phone to avoid a click. That eventually selected some text, and I instinctively tapped to deselect. That triggered another random pick, then I just tapped through to the end because I was fed up.
- smallerize 1mo agoGoogle's SynthID page says they can watermark text, but it also says that it can only detect the watermark on "image, video or audio". Does that mean that the text watermarks can't actually be used as watermarks?
- reactordev 1mo agoWow I actually got a 7/10. It was hard to tell at first but there are signs that tipped me off to which one probably had a higher score out of the multiple choice.
- petters 1mo agoYou likely just got lucky.
- bastawhiz 1mo agoI've read that watermarking should in theory be impossible to detect except by the entity that watermarked it. Which is sensible, and I mostly understand at a high level. But what I don't know and don't understand is what happens if you watermark watermarked text. Does it test positive for both watermarks? Only the second? Indeterminate? Or maybe I'm misunderstanding. Can you tell that it's watermarked, but only the entity who put the watermark in place can test if it's theirs? My confusion about watermarking multiple times still stands, though. Regardless of what happens when you watermark multiple times, no matter the outcome, it weakens the watermark. Which, depending on the threat model, kind of makes it moot. I can't imagine a serious situation where a watermark can be weakened in any way and still be useful. Even "this came from an LLM" isn't a valid signal if you can just watermark ANY text through purely mechanical means. It's also not clear to me how this will affect mainstream LLMs. If all output text is watermarked, there MUST be an escape hatch. Otherwise, JSON schemas will break (or provide holes where unwatermarked text can be exfiltrated through MCP), "return this text exactly with no changes" will be impossible, and writing diffs will break. I feel like I must be missing something.
- red_admiral 1mo agoWatermarking only applies where the AI has a free choice (https://www.anthropic.com/news/claude-text-watermark https://www.anthropic.com/news/claude-text-watermark).
- throw310822 1mo agoDoes it still apply with zero entropy?
- red_admiral 1mo agoNo. Anthropic's example is completing the sentence "Isaac Newton's most famous work was called Principia ..." has only one correct answer, so nothing to watermark.
- 0xedwen 1mo agoi think practical question is whether the detector survives ordinary transformations of the text. if for say i paraphrase a watermarked answer with another model, do we expect the original signal to disappear and the second model's signal to replace it?
- rrr_oh_man 1mo agoIt feels all of them are terribly written. I don't know why.
- StilesCrisis 1mo agoBecause it's AI slop? Not that surprising.
- smikhanov 1mo agoIt takes a lot of patience to read this much slop voluntarily.
- abathur 1mo agoMy sense of the concern here is that watermarking may somehow deprive someone or something of value regardless of whether or not they can tell, so I briefly pondered trying to rank these from best to worst and see if any set of those votes meaningfully deviated from ~average. That said, I read the first triple and found all three tortured enough that I can't be bothered with the rest. Call me persuaded, I guess.
- deleted 1mo ago[deleted]
- MagicMoonlight 1mo ago[dead]
- jibal 1mo agoThis is stupid --- way too long and wordy. Make your test worth taking. And even then it's theoretically impossible to detect the watermark so what even is the point? If it's to check whether the watermarking actually has that property, this is not at all a reliable way to do that.
- IshKebab 1mo agoThe watermarking only works with long wordy text. But yes I agree, it's pointless because we already know that nobody can detect this.
- jdw64 1mo ago8/10. It was harder to distinguish than I expected. If they had applied something like a humanizer skill, it probably would have been nearly impossible to tell.
- kalkin 1mo agoIt sounds like you have a publishable paper - you've broken Aaronsen's watermarking scheme. Congratulations! (Or maybe you got lucky.)
- deleted 1mo ago[deleted]
- aizk 1mo agoI had a moment I thought was concrete watermarking the other day. Claude wrote the sentence... "since the compute buffer estimate has some sl..." Now you'd think the right word would be slack, but Claude wrote... slop? Which does seem off but, those two letters could be tokens very close in probability.
- Jabbles 1mo agoWhat is the meaning of the numbers? > Weighted mean detector score 0.5307 Does it mean that that passage would be rated a 53% chance of being watermarked? So you would need a passage 10x as long to be reasonably sure of providence?
- Jowsey 1mo agoInterestingly, it seems almost every set of three seems to follow a pattern: one passage of the three will have a key word or phrase swapped in the first sentence. That is, for every set of 3 passages, two will start with ~identical sentences, and one will have a key word or token changed. I caught onto this early and used it every time, and ended up getting 2/10, which is worse than random chance. I smell trickery!
- jamienk 1mo agoThe writing reminded me of the SAT test reading section. Impossible to focus.
- raver1975 1mo agoai;dr
- doublerabbit 1mo agoYour result: 7/10 You did slightly better than random chance ... That was random chance. Got bored and just clicked on random boxes.
- DecoPerson 1mo agoHoly moly PLEASE: - Do not make tapping the text act as selection. I was trying to scroll and accidentally advanced twice. - Add back and/or reset test buttons. - Give immediate feedback, rather than asking me to read 10 x 3 long texts. Very frustrating. I was very interested in the results but this combination of problems made this site worthless.
- bombcar 1mo agoWe’re going to get people passing AI text through various competing AIs, aren’t we? Deep Fried AI Slop™ here we come!
- cellis 1mo agoDid 3 then just assumed the longest text was the correct one. Too much reading
- xtiansimon 1mo agoThe UX is awkward, at least on medium sized iPhone. It’s not clear how the answer is selected _and_ submitted. So I scrolled to the bottom for the next button. Then I scrolled to the top. Not seeing an “obvious” sign, my next thought was site must not work on my phone. Only then did I realize the questions were advancing.
- mpalczewski 1mo agoInteresting. This showed how imperceptible it is at first glance. If they had 100 of these and they showed you if you got it right or wrong and why after each answer instead of making you do all ten in a row, I wonder if I would be able to train myself to detect it. I would figure it would just start to seem like something is a little off. The fact that it is so hard to detect just makes this more insidious.