5 ms·
It seems to me like he started out mad and looked to justify it. I'm skeptical that anybody generating LLM text is really all that concerned about optimal word
by wpietri 2mo ago
It seems to me like he started out mad and looked to justify it.
I'm skeptical that anybody generating LLM text is really all that concerned about optimal word choice. Or even particularly good prose. But let's pretend that person exists.
If that person tried, say, an open model and that same model with watermarking applied, I'd be eager to hear their thoughts on the prose quality. Especially if they built an experiment harness and rated a few hundred blinded examples and found a measurable difference.
But getting this upset in advance of any demonstrated problem? It really seems to me like the point isn't the point
- beering 2mo agoGoogle has A/B tested watermarking on millions of responses. They say they observed no difference in user behavior.
- gwd 2mo agoIt says they observed no difference in people clicking thumbs up or down. There are loads of other behavior that they didn't observe; like, say, switching to a different LLM.
- iainmerrick 2mo agoHow exactly do you propose they should keep track of quality, then, if not by A/B testing?
- gwd 2mo agoThe question here is not to what product management advice to the Gemini team. The discussion here is whether the watermarking is noticeable. The easy thing to do here would be to have 1000 questions, randomly assigning one half to an LLM with a watermark, and the other half without. Then show people pairs and say, "Which one seems watermarked?" (Or, "Which text seems more natural" or "Which is a better answer" or something like that.) If they come out equal, the watermark really is indiscernible, at least to most people.
- iainmerrick 2mo agoThey do exactly that "which is the better answer?" test -- I've seen it pop up a few times.
- summarybot 2mo agoIsn't "which one is watermarked?" a different question than "which one is better?" "Which diamonds are shinier, the blood diamond sourced ones or the ethically sourced ones?" ... that's not the same question as "which diamonds are blood diamonds" (to employ an extreme analogy) Concluding that no one could detect which ones were blood diamonds because they were "equally shiny" is not really correct now, is it?
- iainmerrick 2mo agoThat's true, but you don't typically explain what you're testing in this sort of (presumably) randomised trial. And the Daring Fireball article does complain that watermarking will reduce quality. If that's what you're trying to check, "which is better?" is the right question.
- Fogest 2mo agoAgreed, if I simply didn't like the style or words an AI was using in something it wrote, I would switch to a competitors and see what it can come with. I probably wouldn't hit the thumbs down on the Gemini response as it's not that the response is wrong, I just didn't like it. I usually reserve the thumb down for when the AI is wrong. Also, depending on what I am asking it, I often don't want to use the thumb down or up, as this may mean my conversation is going to have some kind of human review and depending on what I am asking for, I may not want to bring attention to my stuff.
- robomc 2mo agoI assume they are just passing off AI prose as their own and don't want anyone to be able to tell. Which is surprising for someone who's been blogging for a thousand years. But I don't really see any other reason for this amount of heat and FUD.
- wpietri 2mo agoIt's a reasonable theory for sure. From what I read, two big problems for long-running columnists are getting tire of the work and running out of things to say. I have no knowledge of Gruber, but I can certainly see why people who are expected regularly to have something to say would turn to "AI". It doesn't get tired and is always ready to spew infinite words.
- boredhedgehog 2mo ago> It seems to me like he started out mad and looked to justify it. Yes, but that's neither surprising nor a reason to dismiss the anger. People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible. They're angry -- and Gruber acknowledges that factor too -- because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner. It might be another instance of consequentialism vs. honor ethics. Many consequentialists don't seem to understand that something that doesn't have demonstrable consequences can still have moral implications.
- zahlman 2mo ago> People get angry about DRM schemes in video games, even if the slowdown these cause is practically imperceptible. Any potential "slowdown" doesn't even come close to making the list of top reasons people get upset about DRM.
- pjc50 2mo agoEven if honor were real, LLMs do not have honor, and people using LLMs to write without disclosing that fact (or indeed at all, to some purists) do not have honor either. > because a stranger manipulates what they regard as their own domain, without consent by or benefit to the owner. It's LLM output! It's not your domain, it's the LLM owner's!
- KaiserPro 2mo ago> People get angry about DRM schemes Yes but thats a thing that degrades something in a catastrophic way, as in I can use the thing one day, and not the next. A different randomisation system on something that is a text generator which is designed to be unperceptable sounds like the people who are annoyed at FLAC vs MP3[1] Done right you won't know the difference, done badly and you will. [1] ex audio engineer, try me.
- porridgeraisin 2mo ago> try me Good luck explaining that one
- 2mo ago
- weego 2mo ago> It seems to me like he started out mad and looked to justify it. that's been his thing since it was just a blog about apple product speculation and update. It's always been tedious.
- Kwpolska 2mo agoConsidering Gruber's always comically butthurt about regulation, especially EU regulation, your theory seems accurate.
- throwawayffffas 2mo agoI care about optimal word choice when generating LLM texts. Because my use case is almost exclusively reading the generated text not posting it. I use LLMs to summarize, translate and review other texts. When using LLMs in that way, as a research tool watermarking is a pointless and should not get in the way of "optimal" results.
- wpietri 2mo agoAgain, given the limits of LLMs (stochastic, rapidly changing, everything's a hallucination, widely known prose issues) I am skeptical that you really care that much about optimal prose. I could believe it's one of the things that you care about, but at a pretty low priority level. Taking you at your word, though, I'd be interested to see what you think of the watermarking technology in a blind A/B test.
- throwawayffffas 2mo agoIt's very important on translations at least. Watermarking will result in poorer results. Do they also do it with code? Do you think deliberately picking tokens that are not the highest probability in code is acceptable for the consumer?
- wpietri 2mo agoWhat's your evidence that it will result in worse translations? I'm skeptical that such a thing as a universally optimal translation exists in cases beyond the trivial. But if it does, I see no reason to think LLMs are anywhere close to it, so I think nobody will be able to tell the difference with watermarking. That's certainly true for code. LLM code is at best mediocre. There is oceans of room to subtly watermark generated code without practical impact.