7 ms·
My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might ha
by ghrl 2mo ago
My biggest concern is that checking any text for watermarks requires sending the entire text to Anthropic. And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, Mistral, ...
So every check requires sending the text to as many AI providers as offer a watermarking detection API, almost all of which have a very dubious track history with obtaining training data through illicit means.
Any university using AI detection in their submission pipeline, or lawyers, editorialists, proofreaders that check for AI marks will be sending significant amounts of text like unpublished research, books, potentially internal documents and more, most of which is high quality human written, to dozens of AI companies, blindly trusting they won't train on any of that.
- josephg 2mo agoI think this is a very real concern. But I’m not sure of any way around it. Any stenographic system that you have the code for can be trivially defeated. I wonder if this would be a good use for homeomorphic encryption. There might be a way to let anthropic check some text without actually giving them access to the source text. Any experts around? We could use your skills!
- piker 2mo agoWon't we just be able to fine tune OSS models to detect these patterns across providers? It will be cat-and-mouse but my bet is it converges to a central detector that isn't affiliated with any model provider.
- josephg 2mo ago> Won't we just be able to fine tune OSS models to detect these patterns across providers? A good fingerprint should make use of cryptographic signatures. Without knowing the keys, the fingerprint should be indistinguishable from noise (or just random token selection)
- piker 2mo agoWouldn’t those hashes be trivially defeated by tweaking the language?
- zrm 2mo ago> Any stenographic system that you have the code for can be trivially defeated. They're giving you an oracle regardless, which is almost as good. Take LLM output, make some modification, ask the detector if it's LLM output, repeat until you learn what kind of changes you have to make to defeat it. Or don't even bother learning what to do, just make arbitrary changes until it says it's not, so when the person they're submitting to does the same check it says the same thing.
- sebastiennight 2mo ago> it says the same thing Reference needed? I think it remains to be proven whether those detectors can be considered deterministic.
- kevincox 2mo agoI assume this oracle will be behind 20 layers of anti-bot protection, CAPTCHAs and hardware attestation challenged. It will be incredibly painful to use. It won't stop the motivated attackers, but will make it too annoying for the average person.
- teeray 2mo agoWhen all else fails, you can hire a lot of folks cheaply to effectively Mechanical Turk it with their home internet connections.
- _flux 2mo agoThey could (..and probably will..) store that version and then refuse the check if this attack is detected, i.e. the version is too close to a known LLM output. Alternatively they could also just keep saying "yes" if it's close enough to a version that was close enough.. Although that would enable the attack to allow arbitrary text to be "proven" AI, by slowly morphing close-enough generated material to the desired text. But perhaps this is not a problem they are not concerned with. To satisfy the letter of the law I expect it's enough to just provide the oracle, without any mitigations.
- nativeit 2mo ago
- chii 2mo ago> blindly trusting they won't train on any of that being allowed to train on any data that you can legally obtain ought to be a right for anyone. After all, i am allowed to learn off anything i can legally read (and perhaps even illegally read). The only thing not allowed (rightly so) is to produce a copy with enough similarities that it can be replacing the original.
- Gud 2mo agoWhy would that be a legal right? Why should we hand over even MORE power to the owner class? In a fantasy world this could be possible yes.
- juggle-anyhow 2mo agoMake it a right, then companies/universities will think twice before using said APIs. Instead of this grey area where we will never know.
- rcxdude 2mo agoCopyright (or any other such restriction on free use of information) creates power for owners by the simple fact that it turns information into something that can be owned.
- m12k 2mo agoWe don't hand over more power to the owner class by making fewer things ownable.
- bonzini 2mo agoIt's a bit different when "training on any data" means basically storing a lossily-compressed copy of that data, that could be spit out years later if the model decides to do so.
- TeMPOraL 2mo agoIt's exactly the same problem as with humans, though. It's part of why we sign NDAs, and why their duration is measured in years (and that's not even targeting the human retention - just duration after which information ages enough that its disclosure is not likely to negatively impact anyone who cares).
- Topfi 2mo agoWhy not switch it around? Solves problems of privacy (no data upload), is far more doable and it's far more important to be able to verify content has not been tampered with after originating from a human or capture device (like a camera). We also have decades of cryptographic experience in reliably signing text, images, video, etc. and it doesn't break apart because a text is too short. Having proof that content (especially images and video evidence) is unmodified (whether via Photoshop, Paint or a model) is far more valuable then having evidence that an image was manipulated or generated fully by a model (which still leaves other forms of manipulation), I feel the same goes for human authored vs generated text. Free to admit that using models to generate any kind of media whole-cloth is still unappealing to me and I still pay for commissioned artwork or make it with my limited abilities for what that's worth. Do like to (poorly) write my musings too and see UX as something were thoughtful contributors (like the opinionated, sometimes controversial, but certainly talented GNOME Gitlab contributors) can make a major impact. Code can be beautiful, interesting and serve purpose beyond execution, of course, but for most people, in most cases, it does not in the same way as audiovisual content (not limited to art). Having code just to execute and resolve a problem can have value all in itself, the code being a means to an end whose quality, let us be honest, was barely a concern in most corporations long before LLMs. Also have rarely (honestly never) before LLMs fully owned all parts of any code base, always relied in part on someone's prior effort in (Flutter/Dart mostly) packages, whereas when writing, drawing, etc. I have far more situations where I make something from scratch and everything there is only there because of my conscious decision. Even simple marketing mockups that, quality wise, any modern model would beat feel different when I was fully in control, where to place what, etc. Objectively worse (at my skill level), probably, but still never the same. Knowing something was made from scratch by a human has value to me, beyond misinformation prevention. Knowing for a fact that LLMs were used instead of importing a library, using a template, or something similar that leads to expending similar amounts of effort, I don't see that being nearly as valuable. Heck, with all the importing and my experience back then vs now, I am spending more effort actually fully reading any LLM output in my code then I spent back then auditing Flutter/Dart packages. Then again, LLM output fails far more unpredictable then those messy packages that simply got Gradle to take down my system... Happy to admit, I have been skeptical of watermarking LLM output being feasible for quite some time and having looked into SynthID Text and proposals being researched, I am convinced that it is challenging to impossible beyond the lowest common denominator and less important then proofing human authorship. It will catch people just copying LLM output into their replies without thought, which is not a negative in my book, especially if it is not discernibly affecting output quality in regular use cases. Anyone who wouldn't copy Wikipedia into their dissertation will, in my opinion, be able to bypass text watermarking as proposed however, I feel we need to be honest there. Thing is, if that's the case and text watermarking will only ever catch LLM created slop, is that a bigger problem then the misinformation, harm to creators due to authorship questions and accusations, making it harder to use evidence in proceedings, teachers not trusting students even when they did the work themselves, etc.? Signatures for all such cases will be difficult to implement, yes, but I feel are going to be of greater value in the not to distant future and I equally feel are not impossible, not least because idiots will always want to hide their LLM usage, whereas human authorship is something they take pride in and want to proof.
- deleted 2mo ago[deleted]
- ThePhysicist 2mo agoWould really like to know how their watermarking technique works, and if they can use it to store arbitrary information in the text (I assume they can if only in a limited way). I assume if the text is long enough they could add all kinds of metadata that would then be undetectable as the model that generated the text is the "cryptographic" key that encodes the data. I wonder if you can have another model run over the text and destroy the watermark. I predict an interesting cat and mouse game to develop.
- ddalex 2mo agoProbably they bias the RNG for selecting the next token. This can be done practically in a lot of ways, including during training. I suspect the signal will be significantly under the noise floor, so it's not detectable if you don't know exactly what to look for, but certainly you can submit more information then the textual contents.
- vintermann 2mo agoBut if the user's prompt is in the context, you don't know exactly what the RNG chooses between. I don't know what trick they use to get past that, but it seems impossible to get by it in the general case (i.e. if the prompt can be anything) and you'll probably quickly compromise quality if you try.
- Majromax 2mo agoA reasonable guess about the algorithm is 'A Watermark for Large Language Models' (https://arxiv.org/abs/2301.10226 https://arxiv.org/abs/2301.10226). The idea is that each generated token (or bigram) seeds a strong PRNG that splits the vocabulary into a 'green' and 'red' set. The sampler then tries to select a 'green' next-token for generation. After-the-fact checking only needs the vocabulary splitter, which is independent of the LLM. Over a sufficiently large text non-watermarked text would expect to use green and red tokens with the baseline probability, and that difference can easily become statistically significant over sufficiently long texts. The basic algorithm has obvious knobs to tune, among them the initial ratio of red to green tokens and how hard the sampler tries to pick a green token. These would balance fidelity to the original distribution against watermark detectability (minimum required content length for statistical power).
- saidnooneever 2mo agoits lovely training data. no detection? add to training set -_-. its also kind of laughable that somehow people are trying to prevent the outputs not to be altered. Asif you cannot manually paraphrase anything you can read. So the only solution would be, to make it utterly unreadable (which is not possible, it obviously defeats the purpose of the thing). Not to mention local models ofcourse :-)
- taneq 2mo agoI think what we’re testing for here is LPM output that hasn’t even been skimmed by a human, let alone paraphrased.
- ShinyLeftPad 2mo agoThat's why detectors like Pangram exist too I think.
- saltwatercowboy 2mo agoPangram doesn't work.
- adamgordonbell 2mo agoLike their fp rate is a lie? It works well in my limited testing.
- saltwatercowboy 2mo agoIt's idiosyncratic to the point of uselessness in mine.
- redsocksfan45 2mo ago[dead]
- ShinyLeftPad 2mo agowhat do you mean by "work"? I think it works perfectly as UGC honeypot.
- andy_ppp 2mo agoDoesn’t this mean Anthropic can accuse anyone of using their AI to write for them?
- mapt 2mo agoYour favorite anti-AI political candidate turns out to have not written their thesis, with a 73% confidence level.
- MagicMoonlight 2mo ago[dead]
- mbreese 2mo agoIs watermarking really watermarking if it can’t be independently verified? I mean, all of these text content watermarking schemes require the company to assess if the text was AI generated or not. They aren’t going to tell us where the toss-up tokens are or what is in the red vs green pools of words.
- marianobayu 2mo ago[dead]
- SatishPophale 2mo agoyes, why would someone will use this tool for writing then?
- jjk166 2mo agoSeems like exactly the sort of problem the threat of defamation lawsuits are meant to solve.
- pjc50 2mo agoEvidence is a massive problem here. As well as the extremely high threshold for US defamation; political candidates routinely tell the most absurd lies about each other.
- SadErn 2mo ago[dead]
- applicative 2mo agoI’m worried that until such tech is perfected, pervasive and uniform across all models, education will be dead, as it certainly is at the moment. Flat out dead. I read stacks of term papers all year and it is a reality that, apart from such schemes, we are in an extinction event for civilization.
- logicchains 2mo agoEducation is in the best place ever for people who actually want to learn, they can have 24/7 access to a tutor with a wide breadth of knowledge and infinite patience for stupid questions. Education is in a terrible place for people who just want to get a degree and don't care about actual learning, but such people generally don't contribute much anyway, and are the easiest to replace with AI, so no big loss.
- PunchyHamster 2mo agoI expect them to work as reliably as AI text generators do now...
- applicative 2mo agoAnthropic had absolutely nothing to do with this. The Chinese models will soon be adopting such devices as well. It is an overwhelming force coming inter alia from educators worldwide.
- LtWorf 2mo agoThat's just storing what they output in a database and then checking, not a watermark.
- kergonath 2mo ago> And even that is not sufficient, as the text might have been generated with ChatGPT, Gemini, Grok, MistraL I would think that these things would eventually converge and we’d get one watermarking algorithm as an industry standard. That way, all major provider would follow it and we’d get independent software for checking. This would partly limit the efficacy of the watermarks, but on the other hand if it’s done correctly, removing the mark could still be enough of a pain that casual users would not bother. That would obviously depend on a lot of factors. It would at least add significant friction in the production of daily slop. Of course it wouldn’t do much for thing like foreign propaganda but that’s a whole other discussion we need to be having. > Any university using AI detection in their submission pipeline, or lawyers, editorialists, proofreaders that check for AI marks will be sending significant amounts of text like unpublished research, books, potentially internal documents and more, most of which is high quality human written, to dozens of AI companies, blindly trusting they won't train on any of that. Isn’t it already what they are doing right now with some of the plagiarism detection tools? Not every university is going to have a representative corpus, and yet they are all using the software. So I guess the provider is doing the work of feeding all that data to their algorithm.
- ComputerGuru 2mo agoYou can’t check it with only the algorithm, you need the secret seed key. Which will be different for each provider (and they’ll probably have and use multiple). And you need the llm itself, to generate the potential tokens at each step.
- Alex3917 2mo ago> checking any text for watermarks requires sending the entire text to Anthropic Couldn’t it be checked in the TEE using confidential computing to keep Anthropic’s algorithm secret?
- ghrl 2mo agoI can't speak for Anthropic, but with Google's SynthID, the algorithm is actually public. However, checking (or creating) the watermark requires a symmetric key, and the providers likely wouldn't share that key.
- atroon 2mo ago>Any university using AI detection in their submission pipeline, or lawyers, editorialists, proofreaders that check for AI marks will be sending significant amounts of text like unpublished research, books, potentially internal documents and more, most of which is high quality human written, to dozens of AI companies, blindly trusting they won't train on any of that. Uh, yeah, why do you think it's setup this way? The frontier companies desperately desire more high quality human text and this is how they are planning to get it for free.
- stillpointlab 2mo agoSounds like a business opportunity.
- apparent 2mo agoAnd I'm sure students will use tools to have every other paragraph written in the style of a different AI, in an attempt to defeat this fingerprinting.
- bluecalm 2mo agoYeah, wait for LLM "scrambles" that put every paragraph and then the whole text through multiple re-write/edit style cycles.
- solid_fuel 2mo agoHow would that change anything? The proposed watermark is applied while the output tokens are being chosen, taking that text and running it through an LLM again would just repeat the process.
- apparent 2mo agoIt says it only applies to passages over 200 words, so you could use a different LLM to rephrase every other paragraph and undermine the watermarking.
- bluecalm 2mo agoYou take the output and run it through another LLM with "please re-write this in xxx style". Then you repeat that a few times on different part of text and glue it all together at the end.
- hirvi74 2mo agoI saw that someone already created a Github repo for a Python script that strips the watermark out of Claude generated text. It was released, I think, within 24 hours of the announcement. I cannot attest to how well it works, but I found it humorous nevertheless.
- deleted 2mo ago[deleted]
- j45 2mo agoWhile Anthropic is sharing this publicly, there’s isn’t much reason that other models could quietly be doing this or start. Local models could probably catch some patterns.
- rolandog 2mo agoExactly, so even if you are avoiding AI, any interaction with society is now being structured so you have to submit to the digital surveillance equivalent of a cavity-search machine. What a dystopia awaits the budding generations.
- micromacrofoot 2mo agoYou've discovered the perpetual mutually assured destruction money generator — guess what the best defense against LLM spam also is? LLMs cause a wide number of problems, the good news for our investors is that they're all solvable with LLMs.