3 ms·
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may s
by Dilettante_ 2mo ago
I notice the "Limitations" section talks about how content only at some point touched by Claude may return a positive, and content that returns a negative may still be Claude generated. But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept.
- sureMan6 2mo agoMaybe it has no false positive rate
- suddenlybananas 2mo agoThat's essentially impossible, unless you mean they didn't measure a false positive rate.
- Filligree 2mo agoFor watermarked long-form text, it is actually possible. Makes the watermark more fragile, but the math is considerably more forgiving than usual.
- embedding-shape 2mo ago> For watermarked long-form text What about "watermarked long-form code"? I'm having a hard time understanding how a model could watermark not prose, but functional/semantic text like code, that actually has meaning. You can't switch our the characters, you can't use various types of whitespace, you can't add arbitrary code comments, and a lot of other restrictions. Is there any state of the art methods for watermarking code without affecting the quality/correctness?
- wpietri 2mo agoAs anybody who has put together a coding standard knows, there are a lot of options for individual expression, meaning a lot of room for things like watermarking. And of course you can add arbitrary comments; my Claude-generated code is very verbose.
- embedding-shape 2mo ago> there are a lot of options for individual expression, meaning a lot of room for things like watermarking The way I use LLMs (and I'd advice everyone to do the same) there really isn't, the agent implements things exactly how I want them, or I use the agent to massage it into the exact bit-by-bit version I imagined when I first sent the prompt afterwards. I honestly don't know what the point would be to let the agents compose worse code than what I'd do manually, although I know it's a popular approach taken by many. > And of course you can add arbitrary comments; my Claude-generated code is very verbose. So watermarking for all users who allow code comments from agents, no watermarking for us who force the agents to never write a single code comment? Alright, I'd be fine with that.
- wpietri 2mo agoFrom what I've seen, your approach to LLMs is exceedingly rare, so I suspect it's one the people who care about watermarking aren't very concerned with. And the reason to let Claude make worse code than a professional would by hand is basically suppressed demand. Since programmers are expensive, previously code mostly got written when a large number of dollars were on the line, or when an individual programmer did something not economically optimum (e.g., hobby project). That left a whole lot of somewhat less valuable software unwritten. It's the economic space that no-code tools have been nibbling on for years. One way to think of things like Claude Code is as effectively no-code tools. Pre-LLM no-code tools would produce data structures that got executed by special environments without ever being seen or tuned by a human. Claude Code can be used just like that, with text as the input and python as the intermediate representation that nobody ever looks at. That approach probably isn't sustainable for what we professional programmers would call a serious project. Claude can easily get in over its head and I expect that its code decays over time, in a fashion similar to how many human teams get in a state where they just have to rewrite everything. But faster, I'd expect. But there are a lot of unserious projects that previously would have never been created. E.g., a quick app to manage your little league team, or a bit of in-house business stuff in the "a little hard to do with a spreadsheet" range.
- OGWhales 2mo agoAnd yet, it remains possible that a human could write the same sequence of characters.
- no_multitudes 2mo agoHow often do you add seemingly-random zero-width unicode characters to the text you write?
- OrangeMusic 2mo agoThat's not how it works (that would be trivial to erase).
- embedding-shape 2mo agoRead said section yourself perhaps.
- mysterydip 2mo agoIf LLM training data is human-written, and LLM output mimics that input, how could you not have false positives?
- SkyBelow 2mo agoBecause it won't be in the training directly. It is applied after a model generates its distribution of likely tokens, biasing each token randomly based on a random key and unrelated to any meaning of the words. So half the time, the most likely token becomes more likely and half the time it becomes less likely, and the same for every other token (when temperature is above 0). You then look at the tokens actually picked to see how closely they follow this pattern that isn't connected to the meaning of the tokens. With enough text, you can then analyze the chance of it happening by chance verses being because the generation of the tokens was done using the algorithm, and you can save a positive result until you are arbitrarily sure. There is a chance of a false positive, but the chance of a false positive approaches the chance that the murderer happened to have fingerprints that matched your and both forensics labs happened to have mixed up the dna tests and the eye witness happened to misremember the face and your phone gps happened to glitch out and put you at the murder scene at the time of the crime all happening. It is theoretically possible only in the same sense that quantum teleporting a cat is theoretically possible. The real question is how much text do they need for a given level of certainty and what do they check for. If they flag a positive at a p value <.01, that's a problem. If they can reasonably get a p value of < 1e-12 in only a few paragraphs of text, that is effectively no false positives (but a lot of 'too short to analyze' outcomes).
- ricericerice 2mo agoHow do you verify in practice then? Wouldn't you need the original prompt so you can reobtain the likely token distribution to validate again the random key(s)?
- Eisenstein 2mo agoA token is hashed and used to seed a random number generator, which produces the red list for the token after it. Paper: * https://arxiv.org/abs/2301.10226 https://arxiv.org/abs/2301.10226
- wrsh07 2mo agoSomewhat trivially, if I ask Claude to transcribe an image and then check if that transcription is ai generated it will likely say yes. Many users are not smart enough to realize that the transcription step is where the ai (watermarks) were necessarily injected.
- basch 2mo agoHow is a perfect transcription of an image watermarked?
- FeteCommuniste 2mo ago"Perfect as far as human perception can tell" is a weaker standard than "bit-to-bit copy." Maybe it's that?
- wrsh07 2mo agoIt depends on how it does watermarking!! Note, there are many ways to represent words visually on computers that look identical
- basch 2mo agoIf they were substituting glyphs for identical ones people would be able to reverse engineer it. Theres no way that’s what they are doing.
- wtfwhateven 2mo agoWhy would you say something so ridiculous?
- jobigoud 2mo agoI think they mean it like this: imagine you ask me a random number sequence. I give you a random number sequence. Little did you know, I used a very specific PRNG to generate it, so later I can prove with certainty that your number was generated by me, and you can't say you came up with it yourself. There is no room for false positive here in the same way you can't randomly find a collision in a hash function if it's strong enough. Like the rate is so infinitesimal that it is effectively zero. Now replace random number sequence with prompted string of words. And instead of using the PRNG on every word I use it every n words. If the generated text is sufficiently long I can tell by matching the expected deterministic pattern. You can defeat it by changing the words yourself and triggering a false negative but there isn't really any room for a false positive if the text is long enough and the pattern matches perfectly. If the pattern doesn't match then I can compute a probability.
- bufbupa 2mo agoSorry you're getting downvoted, this interpretation doesn't seem that far fetched to me. Here's the strawman: The text-based watermarking is going to be done procedurally instead of generatively. Maybe they add some sequence of zero-width Unicode characters to all generated text at certain intervals. Then, there is effectively no false positive possible (because humans would [effectively] never type such sequences of unicode naturally). It may survive some editing (depending on how you select/edit the characters), and it's possible to be stripped (false negatives).
- unprovable 2mo agoThis. FN rates are cute, but FP rates will ruin an academic career or a student's work/further study choices if their content gets marked erroneously. Surely the answer is a sequence of marks? Keen to see if they are doing something SynthID-esque?
- WD-42 2mo agoDo they care about false positives? As long as it’s even somewhat reliable that’s enough for them to prevent training on their own slop. I think this is a big reason to do this that’s overlooked.
- DennisP 2mo agoGood point. But if that were their only purpose, there'd be no need to share it with anybody. In fact, they'd get the best results by not mentioning it.
- WD-42 2mo agoThat’s true. But they probably want to be able to identify other models slop as well. And with the laws popping up, it makes sense to do it the way they are.
- WiSaGaN 2mo agoMy guess is that they will later "reveal" some "violations" but provide little evidence citing proprietary algorithm.
- akersten 2mo agoIf there is any false positive rate (which, because text will naturally and by chance include tokens from the green and red sets in some pattern, there will be), tools making promises like "detect AI-generated text" are unacceptable. They are going to turn innocent people into pariahs on some unsubstantiated "this content is 37% likely to be AI" claim that the user has no way of verifying or inspecting more deeply, we just have to trust the statistical box and assign some meaning to whatever that number means. 37% of my phrases are AI? There's a 37% chance my entire text is AI written? Part of the fun is not knowing! This is scripture homeopathy and it's irresponsible.
- wrsh07 2mo agoI'm curious about your thoughts on pangram. I only really see posts on Reddit claiming it falsely labels their content as ai generated but nobody will actually post examples of "textbook from twenty years ago" or upload screenshots of a journal (also those posts usually feel deeply ai generated without an ai detector) Do you think this is an impossible task and we shouldn't try to solve it? Or do you think it's doable and that some ai detectors might be better than others?
- nemomarx 2mo agoThis feels testable - you could go to fanfiction or similar sites with billions of words of writing from before 2016 or so and run them through it. I tried a chapter just now and got human doing that, but I'm not invested enough to run a hundred samples today. But it sounds like it would be an alright way to audit it? I will confess I'm pretty skeptical you could ever eliminate false positives here though. I can often get an ai sense from some writing on my own but I doubt it would be better than 90% accurate, and "ai plus human editing" might screw with that anyway, stuff like that. I would have preferred we just never developed this kind of thing so I wouldn't have to guess.
- StilesCrisis 2mo agoIt's been done and showed up on HN recently. Older content was quite consistently marked as not-AI.
- nunez 2mo agoI think Pangram is way ahead of Anthropic on this with their custom dataset.
- whack 2mo ago> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated Are you talking about pieces that were fully human-written with zero AI editing/rewriting etc? If so, what makes you think that false positives will happen there? They aren't looking for "writing styles" or emdashes etc. They are using watermarks and metadata. If you're talking about people using AI to copy-edit text they manually wrote, this was explicitly called out in the article: > A detected mark provides a signal that content was processed by Claude, but is not fully conclusive. Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content. For example: Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source; The content may have changed after Claude processed it. Marked content may be modified, excerpted, or combined with other material after Claude processed it.
- Dilettante_ 2mo agoThe former. I'm not sure what you mean by metadata, but my expectation was that anything that Claude could put into the plaintext to identify itself may plausibly also accidentally be produced by [a million monkeys on typewriters/one in a million human writers], since in the end, the writing is using the same language and symbols that humans use. How unique could the LLM possibly make it while still retaining its usefulness?
- derefr 2mo ago> How unique could the LLM possibly make it while still retaining its usefulness? They could be doing invisible and vaguely-harmless Unicode stuff. Insertion of zero-width joiners and non-joiners, replacement of regular spaces with non-breaking spaces, building spaces from multiple hairline spaces, intentional use of non-NFC-normalized codepoint sequences for accented characters, etc. Text with all this junk in it still reads the same; it just might wrap a little strangely, or not byte-match / collate correctly in a database (and Anthropic has never made a guarantee that their models would be capable of emitting text with these properties, so that’s fine.) And, importantly, no regular text or document editor would insert these things (especially in the useless places you could insert them for watermarking.) You only really see them in text that’s been explicitly typeset for a specific layout (e.g. in text-containing SVGs, website mastheads, or game HUDs) or for print publication. Of course, if this is the technique they end up using, then it’s very simple to strip it out by canonicalizing the text (i.e. Unicode-normalizing it + stripping out invisible layout characters + replacing “weird spaces” with regular ones, etc. Essentially the same thing many sites already do to user-generated content to prevent users from using Unicode features to break the page’s layout.
- m00dy 2mo agoDeepwalker once cracked Gemini's watermarking system, I'm sure they will also work on this [0]. [0]: https://deepwalker.xyz/blog/evaluating-synthid-watermark-robustness https://deepwalker.xyz/blog/evaluating-synthid-watermark-rob...
- jrflo 2mo agoI think that false positives are inevitable due to the method of watermarking being embedded in the text itself. The output is intended to mimic human writing, therefore it's entirely conceivable that a human could by chance write text that contains the watermark. The odds may be extremely small, but it's not something you could ever guarantee.
- andai 2mo agoI keep hearing how humans are thinking and writing more and more like AI. I think in this case I think it's some kind of cryptographic signature smeared across the token IDs, so I don't think the risk is very high.
- xena 2mo agoYou're absolutely right! Humans have been slowly thinking and writing more and more like AI. As people get more and more exposed to the stochastic patterns of large language model tools, it's normal for them to emulate the styles of communication they are exposed to. This is commonly called "brainrot" by those in Gen Z and younger cohorts. If you find yourself getting to be afflicted by this "brainrot", be sure to go outside and take a moment to ponder what's around you. The grass is there and will be there long after we are all gone. Consider this for a moment as your organic thought processing unit starts to slowly munch away at its internal context window.
- jrflo 2mo agoI read the original paper they're basing this off of and I think you're right. I do wonder how much of a quality tradeoff there is with perturbing the next token probability distribution. My intuition tells me that a more "prominent" watermark will necessarily degrade output quality. If they are trying to balance quality and watermark prominence, I wonder if that affects the FPR.
- ed_elliott_asc 2mo agoIt’s worse than that, false positives are possible but someone generating text should be able to get ai to change some words and formatting to break the watermarking, then ai detectors can tell them how well they did. I don’t know what the answer but I absolutely know it isn’t this.
- dbqpdb 2mo agoI think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board. But each intermediary or source (optionally) cryptographicaly signs a piece of content that it either generates, edits, or passes along, and the end result at a destination, is that content is either 'trusted' if its cryptographic chain is solid, or un-trusted otherwise.
- dragonwriter 2mo ago> I think we need a chain of custody system for content, but that would require browsers, software, websites, operating systems, phones, camera manufacturers, etc to all get on board. It would also require the individual humans you are trying to control to get on board otherwise the analog hole breaks the chain, absent mindboggling levels of physical surveillance on top of the the total monitoring of all electronic data flows that this idea requires.
- normalaccess 2mo agoI think that's the end goal.
- normalaccess 2mo agoThis is a meme video but I think it hits the nail on the head. TLDR: AI will force global online digital ID for everyone that uses the internet for the exact reason you mentioned. And that would forever change free speech forever allowing the powers that be to put the genie "back in the bottle" so to speak. link: Raiden Warned About AI Censorship - https://youtu.be/-gGLvg0n-uY https://youtu.be/-gGLvg0n-uY
- dragonwriter 2mo ago> But I really would have liked for them to state explicitly that entirely false positives where a piece is fully human-written may still be marked as generated, because too many institutions with the power to ruin someone's life over that have trouble understanding the concept. This is marketing material aimed, in part, at encouraging the usage you are concerned about, which is why they do not highlight that problem.
- anon373839 2mo agoBut I thought Anthropic was an altruistic organization devoted to the betterment of humanity…
- AustinDev 2mo agoIt would appear that their altruism isn't very effective.
- hojinkoh 2mo agoSure, there are things you could do legally when falsely accused; and there are things authorities and companies should do. But ultimately, when you are powerless and can't afford to do the fighting: I'm convinced the only way to protect yourself is to be very mindful about your writing style, and to deliberately corrupt the language through objectively wrong "stylistic elements".