4 ms·
Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having be
by smallerize 2mo ago
Translation: No one can ever again use Claude for proofreading their own prose unless they’re willing to risk that the whole thing might be flagged as having been generated by Claude.
I think that was intended, yes.
- ButlerianJihad 2mo agoIt is quite just, if you think about it. Human works are copyrighted and protected at the moment of creation. All rights reserved. Yet, LLM outputs are uncopyrightable. Therefore, if Claude or any AI has processed my copyrighted work, the end result is uncopyrightable and in the Public Domain. The public has a right to know: is this a human copyrighted work, an LLM PD work, or is the human falsely claiming authorship in order to retain copyright? A point of confusion for me, however: is every watermark unique? Is every algorithm for watermarking going to vary amongst models and amongst model versions? Will each model publisher keep this watermarking as a trade secret, that they alone can detect? If so, this can't scale! How do you detect "JoeBob 4.3 LLM" output? By querying every single model's watermark-detector? And if they all work by re-running the model and using tokens anew? That is extraordinarily wasteful. If a watermark is not self-evident, or universally detectable, then it is no good. Take, for example, US currency. The security measures are published and well known. Any count-out room in retail has a big poster indicating how you can detect authentic US bills. Nobody has to accept non-US currency in the US, and so the only authenticity you need to worry about is your US bills alone. LLM watermarking has none of this in common. Currently sounding like a shitshow, if you ask me.
- fwipsy 2mo agoPerhaps LLM outputs are uncopyrightable, but derivative works of copyrighted works are not automatically in the public domain.
- ButlerianJihad 2mo agoThat's an intriguing twist, isn't it? It could lead to a tug-of-war. Working backwards: if it is possible to confirm 100% confidence that a chunk of text is LLM output, then it is "PD until proven otherwise". How can a human reliably assert human authorship of their source text? When all watermark tests fail? Is that proof of humanity now? If a human proves human authorship, and LLM watermarking tests positive, then is that going to be considered a "derivative work" or not? What if there is an applicable license for the source work, such as "CC-BY-ND" that prohibits derivative works? This has not been court-tested, and I expect that it will need testing at that level before we can have any assurances.
- deleted 2mo ago[deleted]
- pessimizer 2mo agoThe watermark, if I'm understanding correctly, only proves that something has been touched at some point by a particular model, not that it was entirely generated by a particular model. As for it being "uncopyrightable" if it were the output of an LLM, I think computer people are making a very aspie interpretation of a single decision. I think it's more that the LLM (and thereby its owners) cannot itself hold a copyright on its output, that output has to be touched by a person before it is copyrightable. A particular view from the top of a mountain can't be copyrighted, for example, but a photograph of that view can be. I'm not sure it at all precludes a "robosigning"* sort of situation, where machines generate output, hired temps sign and claim that output, and immediately sign it over to the people who hired them (as a work-for-hire.) Copyright is stupid, artificial law, not logical. ----- * https://www.mortgageauditsonline.com/what-are-robo-signers/ https://www.mortgageauditsonline.com/what-are-robo-signers/
- demibabs 2mo ago> The watermark, if I'm understanding correctly, only proves that something has been touched at some point by a particular model, not that it was entirely generated by a particular model. For the watermark to be detectable, the text needs to be like 75% AI generated. If you have an LLM “touch” one section of the article, it’s not gonna be detectable.
- inigyou 2mo agoHow do you assert it now? I post some text on the internet, you claim you have copyright, how do you prove that?
- pizzly 2mo agoSome camera pointing at you while you work. Guessing the camera needs to be watermarked itself, which can be done by adding some watermark on the chip level (each camera will have different watermark), manufacturer can confirm the watermark.
- dare944 2mo agoAs I understand it, the current watermarking methods rely on a secret key, making the detection schemes a black box to anyone not in possession of the key. This means organizations like Anthropic are free to make any claim about authorship they want, true or not, and no one can call them on it.
- demibabs 2mo agoPart of the legislation requires them to make a public AI text detector (ala GPTZero I assume). Wouldn’t having that be enough to eventually reverse engineer the key?
- dare944 2mo agoNot if they designed the algorithm right.
- inigyou 2mo agoProbably not to get the key, but you could certainly use it adversarially to remove the watermark. Removal may come down to changing every third token to a different one.
- NewsaHackO 2mo agoYea, they are going to have to monitor how often similar writing pieces are being submitted, or limit the access to the checker somehow.
- inigyou 2mo agoEurope trusts its institutions way too much. It may well be that only the police can access the checker.
- demetrius 2mo agoI'm not sure the quoted statement is true. Proofreading like "point to problems in the text", if you fix the problems yourself and don't copy-paste the solutions given to you, should still be safe, shouldn't it? So, human-written text should not be falsely flagged if you use LLM for proofreading. And if you copy-paste the answers from LLM, I think it's only fair the end result gets flagged. You're not writing it yourself.
- 190n 2mo ago> And if you copy-paste the answers from LLF, I think it's only fair the end result gets flagged. You're not writing it yourself. I wonder if it would even get flagged in that case, because wouldn't the probability distribution of a token when the LLM is suggesting an edit to your writing be different than the distribution of that token once it is in the context of the text it's editing?
- beej71 2mo agoI don't even ask LLMs to go that far. Tell me if I've made a spelling, punctuation, or grammatical error, period. Don't rewrite a thing, because LLMs suck at that.
- jleyank 2mo agoRands made this point a few days ago as I recall. Worries about having his tool corrupt his writing during editing, etc.
- skew-aberration 2mo agoCan't the LLM just generate e.g diffs? Or some other intermediate language. Then the watermark is lost when the translation step is applied.
- croemer 2mo agoThat's a good one. Problem is that SynthID uses only roughly the last 4 tokens, so any literal test 4 tokens or longer that is generated has the watermark. If you fix typos, insert punctuation, rename variables then yes this should be undetectable if produced from a diff.
- epihelix 2mo agoYou know, back in the era when proofreaders were human, I never one met a proofreader who rewrote my text afresh, rather than annotating the text with a pen. It's still possible to use Claude to proofread - highlight grammatical, flow, structure, logic errors and make simple suggestions for you to pick and choose or adapt as you wish. No watermarking will flag your text. No flaw accusations of LLM authorship will haunt you. All will be fine. But if you want an LLM to rewrite your text, that's (a) not proofreading, and (b) should be flagged as LLM generated ... because it is.
- richardatlarge 2mo agoNonsense. Forget proofreaders. Think editors. In publishing some editors practically wrote the books. And then theres ghostwriting ! Think of that!
- Apocryphon 2mo agoOkay, then it should be acknowledged if a work was AI-edited-written, or AI-ghostwritten.
- kalleboo 2mo agoPart of me wishes we had the same regulation for ghostwriting etc. Nobody should be claiming to have written a book they didn't.
- wuschel 2mo agoWhere is the problem with using LLM generated text? You could use your own hypothetical house elf to do it for you, or pay someone to do it. LLMs are just cheaper for a certain set of problems. People will find ways to circumvent this, so this limitation will only hit the technically less adept people.
- zahlman 2mo ago> Where is the problem with using LLM generated text? In the fact that you didn't write it. > You could use your own hypothetical house elf to do it for you, or pay someone to do it. Yes, and those would be similarly problematic (and more expensive).
- zmmmmm 2mo agoThe question is, when is the "pro writer" version coming that lets you control this behaviour but costs more? Like night follows day, this will happen. They will need to dodge around the EU requirements but it will probably just come down to an alternative method to watermark or a contractual assurance you won't mis-represent the source of the text.
- wuschel 2mo agoI had the same thought. I hope the other providers will add a geographical limitation on this EU rule. (On a side note, I wish they would replace those EU beauracts with LLms).
- smb06 2mo agoThe thing that could change is interpreting "the whole thing as generated by Claude"
- inigyou 2mo agoWell yeah, if it's output from Claude it's likely to get detected as being output from Claude.
- beering 2mo agoIt’s a strawman argument because if the LLM is really just “proofreading” for you, there will be little or no watermarked text in your writing. Not enough to trip the watermark detector.
- KaiserPro 2mo agowait what? but proof reading is a linter, not a writer. the proof reader will say "I think this is clumsy can you try x,y & z"