4 ms·
Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that the
by levocardia 2mo ago
Crazy how a smart person like this fails to understand the gumbel softmax technique. It does not affect writing quality at all, provably. The very fact that there is generally no "best next token" with 100% certainty is precisely why the trick works (you cannot watermark a response to "respond with the To be or not to be soliloquy from the first folio Hamlet", for precisely this reason).
- reader9274 2mo ago[flagged]
- brookst 2mo agoI think he’s still generally good on business, UX, and hardware design. That’s all subjective and taste I suppose, but his taste works for me. On deeper tech stuff, like this utterly nonsensical misunderstanding of watermarks… yeah, classic case of a guy who is smart, and has lost the ability to realize when they’re not knowledgeable in a domain.
- tapland 2mo agoMaking blog posts about AI that make it apparent that the tech is going whoosh is a choice.
- selectively 2mo ago[dead]
- docjay 2mo ago[dead]
- Art9681 2mo agoIf this is true then the probability of the detection tools flagging completely human generated text as AI generated is non-trivial. Let's say I write a completely original piece and the detection tool says there is a 36% probability it was generated with Claude. What then? Now it's up to the person looking at the score to cast a subjective judgement. Maybe to me, anything over 25% is unacceptable. Maybe to someone else, it must cross over the 50% threshold. This is the problem. Cognitive surrender.
- pessimizer 2mo agoI don't think that's true. I think it's a binary 0% or near 100% probability of a watermark having been detected; the more changes to the text having been made after the text was output by the LLM and the less leeway the LLM had for probable word choices, the longer the passage necessary to see it. The "problem" is that seeing the watermark doesn't mean that the person claiming to be the author didn't make extensive changes to the output of the LLM, or that the LLM wasn't simply the final editor of something that the author had put a lot of work into. > Cognitive surrender. I don't know what this means. It's just drama. Don't let the LLM write for you and this is not a worry. I'm not worried about the poetry of LLM output being subtly adulterated.
- beering 2mo agoNo, watermark detection is not binary, you get a real number. You decide on a threshold when looking for the watermark. This is the problem - by random chance, some human text will be detected as watermarked. You can turn the detection threshold up until it guarantees <0.001 false positive rate at the expense of higher false negatives, but seems inevitable that someone gets wrongly flagged.
- fwipsy 2mo ago> the probability of the detection tools flagging completely human generated text as AI generated is non-trivial How does that follow? AI-generated text is already not a perfect emulation of human writing. There's lots of room to affect it laterally without changing the level of quality. As I understand it, LLMs with temperature >0 can select from many possible outputs. All they're doing is limiting the possible outputs to ones that contain this pattern. I don't see any reason why the quality of that subset should be lower than average. The very best outputs will likely be eliminated, but so will the very worst.
- eru 2mo agoTo sample from the probability distribution you already need random numbers. If you get your random numbers from a cryptographic PRNG, then to notice the difference between that and 'real' random numbers even in theory, means you need to break the cryptography. In practice, your gut feeling about how good some text is won't break modern cryptography.
- wpietri 2mo agoIt seems to me like he started out mad and looked to justify it. I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists. If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experiment harness and rated a few hundred blinded examples and found a measurable difference. But getting this upset in advance of any demonstrated problem? It really seems to me like the point isn't the point
- beering 2mo agoGoogle has A/B tested watermarking on millions of responses. They say they observed no difference in user behavior.
- gwd 2mo agoIt says they observed no difference in people clicking thumbs up or down. There are loads of other behavior that they didn't observe; like, say, switching to a different LLM.
- iainmerrick 2mo agoHow exactly do you propose they should keep track of quality, then, if not by A/B testing?
- gwd 2mo agoThe question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable. The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
- conartist6 2mo agoI couldn't be happier that people are mad about it. To quote Calvin, "nothing helps a bad mood like spreading it around"
- krackers 2mo ago>fails to understand the gumbel softmax technique I think in this case it doesn't help that there are multiple watermarking schemes, and the easiest for people to understand is the red/green scheme by Kirchenbauer et al. (https://arxiv.org/pdf/2301.10226 https://arxiv.org/pdf/2301.10226), which does technically distort the logits (but I'd argue only in cases where you wouldn't notice it anyway). I wasn't aware of this gumbel softmax scheme, it seems you're referring to https://simons.berkeley.edu/talks/scott-aaronson-ut-austin-openai-2023-08-17 https://simons.berkeley.edu/talks/scott-aaronson-ut-austin-o... ? That's really clever as it doesn't even distort the logits, basically cryptographically indistinguishable from a "real" random sample unless you have the key. The actual scheme Claude uses seems to be neither of those two though, they say it is SynthId-text which seems to be tournament sampling based.
- zahlman 2mo ago> The very fact that there is generally no "best next token" with 100% certainty Indeed. It's frankly bizarre to see the assumption to the contrary being made by someone who's been passionately blogging by hand for years, who also happens to be responsible for the notoriously vague, humanistic, DWIMmy Markdown standard.
- neuroticnews25 2mo ago>It does not affect writing quality at all, provably Prove it, then? It's not a claim that GumbelSoft paper makes: "Regarding generation quality (perplexity), GumbelSoft shows relatively low perplexity" https://arxiv.org/html/2402.12948v3 https://arxiv.org/html/2402.12948v3
- yorwba 2mo agoThey do claim that the per-token output distribution remains unchanged, but the proof is relegated to Appendix B.1. The perplexity comparison includes methods that do change the output distribution.
- CuriouslyC 2mo agoBy definition watermarking narrows and biases the response distribution. Clever algorithms might reduce the perceptual impact and minimize some cherry picked metrics, but it's still worse.
- adamgordonbell 2mo agoIt's a writer perspective versus a reader perspective maybe? Sometimes when you're trying to write something, it really seems like the exact words matter a lot. Suggestions made to be more direct or use a more common word here or whatever seem to really impact the thought that you're trying to communicate. Certainly we've all had times when trying to communicate clearly when the specific words seem very important.
- urams 2mo agoGruber is always happy to lead with his emotions and backfill justifications for them. See his recent debacle with App Store review: https://daringfireball.net/2026/08/retraction_app_store_rejection_of_the_week https://daringfireball.net/2026/08/retraction_app_store_reje...
- ragazzina 2mo ago> a smart person like this Can you give an example of something smart John Gruber has said or written? Because I can't think of one, but I can think of many dumb ones.
- jgruber 2mo agoAgreed!
- sva_ 2mo ago> The very fact that there is generally no "best next token" with 100% certainty This is not entirely accurate. Sure, there's never a token with 100% certainty, but there are often tokens with 99.9% probability, but this technique of course does not change how such a token is sampled.