6 ms·
Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (pa
by akersten 3mo ago
Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading.
Images, absolutely, there are tell-tale artifacts from today's generators that simply aren't emitted by "natural" paths to create them, and you can "detect AI" with high confidence (for now). Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.
- jgalt212 3mo agoIt depends on how much text. For example, chardet often falls down on short strings, but 1K characters it nails it.
- stymaar 3mo ago> but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Hard disagree. LLMs (especially base ones, that only received pre-training) can produce output that is undistinguishable from human writing (because that's what they were trained to do). But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy to spot when reading AI slop, and that's not surprising that it's easy to spot automatically either. And I don't think that's going to change anytime soon, unless their incentives change. (We can say exactly the same thing about man-made stuff optimized for a specific purpose, like stock photography, clickbait titles or industrial food: they aren't stereotypical because their creator lacks the skill to make them otherwise, they are like that because that's what works best).
- empath75 3mo agoIt does mean that this will have a drift problem if it's just trained on the idiosyncrasies of model fine tuning. That's fine! But it is something to be aware of.
- ravenstine 3mo agoThey're also designed to not offend anybody, so their output tends to be very bland even compared to the most milquetoast of human beings. I was only surprised once when ChatGPT responded with an enthusiastic "hell yes" seemingly organically, but 99.9% of the time these AI services clearly are instructed and trained to provide flavorless word vomit. I don't think there's a technical reason why an LLM couldn't produce totally convincing output, but internet grifters don't need to go through that trouble. It's like how most phone, email, and social media scams come off as completely transparent to most of us, but that's the whole point; we're not the target audience of the scams. Readers looking for substance, nuance, and real opinions aren't going to notice if something with written by an LLM – unless there are some cliche punctuation tells.
- pixl97 3mo agoWhen DANmode bypasses were a common thing the LLMs would drift significantly far from corporate speak. But that's the point of corporate speak, you tend not to say thing that may offend your clients and deprive the company of future revenue. Of course there are some companies that make their living being 'counter-culture' and saying what they want, but they are a small percentage of all revenue.
- BoredomIsFun 3mo ago> especially base ones Did you actually try them? I did.They generated even more "slopey" text than instruction-tuned ones.
- skinfaxi 3mo agoBut was it content indistinguishable from someone learning the language being used, for instance?
- AnthonyMouse 3mo ago> But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy to spot when reading AI slop, and that's not surprising that it's easy to spot automatically either. There are two problems with this. The first is that it would still misclassify human-authored text written under the same incentive, and most people have various incentives to "maximize engagement". And the second is that then people would just make other models that are tuned for defeating that sort of classifier, which would be used whenever the classifier is being used.
- OneManyNone 3mo agoAll of that may be true, but pangram currently has a false positive rate of about 1 in 10000, and this has been tested by feeding in thousands of texts written before 2020. That may not last if AI companies start trying to build models that fool it, but for the time being at least, modern models do have strong tells.
- AnthonyMouse 3mo agoYou can get an arbitrarily low false positive rate by sacrificing against false negatives. It's trivial to make it zero, just classify everything as human-generated. Meanwhile a false negative rate of even 1% is a pretty big problem since someone can easily use LLMs to generate 100x the volume of text and then use whichever ones make it through the classifier. And that's before anyone even tries to get the LLM to generate a different style of text. Or for that matter creates a "style model" that rephrases text.
- vidarh 3mo agoYou don't really need a style model - current models are very good at doing "style transfer" of a model text onto whatever it has written if you just have it do it chunk by chunk. It takes more to prevent it from being detectable by good detectors, but it does remove a lot of the worst tells.
- 3mo ago
- zmjone2992 3mo agoi think one thing overlooked by this perspective is that many of a detectors adversaries are not that sophisticated. so despite this i think it is a useful thing to try to do. particularly when people are trying to do fraud which will often having to use abliterated models and generally trying to be as economical in their efforts
- cyanydeez 3mo agoSure it is; we do it all the time, and then we modify each other's etc, etc; english we speak today was spoke yesterday waspake the same in yesteryears; we have no trouble dating english or other languages to a time. A better argument is people themselves are just too influenced by reading that they'll sound like LLMs in a couple of years.
- driverdan 3mo agoThere are two problems, false positives and changing the LLM's pattern. It's really easy to have a false positive and false positives can be very harmful if the person using the detector isn't aware of that risk. It's also very easy to change the pattern of LLM output. You can provide basic prompting that will significantly change the structure of the output. For example, having it utilize the Wikipedia article on signs of AI writing and avoid everything it describes. https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
- WhitneyLand 3mo ago"It's really easy to have a false positive" Not really. The false positives for the SOTA detector are very very low. "It's also very easy to change the pattern of LLM output." Not in a way that can reliably avoid detection. The problem is the patterns are baked into the distribution itself. It's smoothed over, so it becomes difficult to prompt your way out of that.
- Der_Einzige 3mo agoWrong. Effective sampling (I.e high temperature like temp 10) with the corresponding sampler stack that enables this coherently destroys all attempts to detect it. There are many more ways like this involving manipulating the logprobs
- WhitneyLand 3mo agoI’m not sure what you’re saying I’m wrong about. The comment I was responding to, about changing LLM output, referred to prompting, not temp/sampling tricks. I’m not aware of Pangram being beat by clever prompting. There’s some interesting work on creative writing using contrastive prompt techniques, but I haven’t seen it tried as evasion. Even if you control temp and sampling, they’re not magic. If you raise the temperature too much writing can go to hell, so you may beat the detector but end up with junk. There are some ways to mitigate such a quality drop like raising temperature in conjunction with min-p, but still, I haven’t read any research that shows it getting good results at anything close to 10. Now you want to get more clever and manipulate logprobs…well ok, you could come up with elaborate strategies designed to evade specific detection methods. But I don’t see that getting done as a weekend project while maintaining writing quality. And if it does happen there’s no guarantee the detector can’t train on its characteristics and start an arms race.
- Retric 3mo agoSignal is easier to detect with more data to work with. Largely AI generated books are a vastly different situation than a one paragraph homework assignment. But multiple rounds of homework assignments would change the accuracy.
- yorwba 3mo agoWhether a text was written by a human or not is just a single bit of information. So you can't rule out its detectability a priori, since even the shortest text contains more information than that. As long as LLMs are used to write texts humans wouldn't want to write if they could help it (that's why they're getting an LLM to do it, after all), they'll remain detectable. Even if the reasoning might end up equivalent to "This looks like spam; no human in their right mind would write this spam by hand if they could get an LLM to write it, therefore it's most likely written by an LLM."
- onestay42 3mo agoNot all humans are in their right minds, unfortunately.
- cwmoore 3mo agoIt is much harder to tell one from the other, and for oneself, than it often seems on the surface.
- skinfaxi 3mo agoThis is exactly the point I saw in a recent x post, that building anti-bot detection was incredibly difficult because some people exhibit bot like behavior. Blizzard employee once told me anti-botting in WoW was extremely challenging due to the number of real people that acted identically to bots. Every assumption was invalidated: - unbelievable # of consecutive hours played - consistently repetitive patterns of movement and clicks - farming patterns that aren’t considered fun (“why would anyone do that”) - solo, no external engagement - goes on for months The problem with botting is many humans ARE bots https://x.com/IceSolst/status/2076372992959959493 https://x.com/IceSolst/status/2076372992959959493
- inigyou 3mo agoHow do they know those were real people? Were they livestreaming their face and talking about what they were doing the whole time?
- 3mo ago
- WhitneyLand 3mo ago"Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it...it's a bad fiction to perpetuate that any of this is anything more than tarot card reading." Not true at all. Pangram is highly effective and has a very low false positive rate. The post here is impressive for a small project, it looks like they independently thought of one of the core ideas Pangram uses of creating twins to compare. You can see how it works here: https://arxiv.org/pdf/2402.14873 https://arxiv.org/pdf/2402.14873
- hgoel 3mo agoSo, if the decision from Pangram determined, on every assignment, if you would be expelled from university for plagiarism, would that be acceptable to you regardless of how you actually did the work? If you would not be okay with that, what level of consequence would be acceptable for the output from this tool?
- WhitneyLand 3mo agoThat’s a different point. I’d want detectors to be as accurate as possible, false positives of 1 in 10000 seems like a good starting point. I believe their results have been independently tested. And as a separate matter, any tool for evaluating students should be applied fairly, safely, and with adequate human review and due process. You need good tools and good oversight.
- cwmoore 3mo agoDue process should never just become a checkbox item. To deal with lives and livelihoods justly, you need appeal pathways and meaningful liability exposure for the processors. Plagiarism and cheating sucks for everyone. Worth solving.
- hgoel 3mo ago>And as a separate matter, any tool for evaluating students should be applied fairly, safely, and with adequate human review and due process. Agreed, that's a fair and reasonable stance. The reason I asked is that I have a hard time understanding the point of these tools. When it comes to education, it can be a matter of learning objectives. But outside that, what's the point? The prediction from the tool is pointless for deciding on copyright or contract issues, and other text should be judged on its correctness or applicability to the task. If all the tool is good for is "maybe this student cheated, but only an in-depth investigation would maybe prove it", it isn't a very useful tool, because it's more straightforward to just mandate that evidence is submitted regardless of what the tool says. On top of that, even the lack of evidence of manual work isn't good proof of using LLMs.
- onecomment1 3mo agoThis sounds like it was edited by an llm.
- overgard 3mo agoI don't know, the thing about most text slop is how little effort goes into disguising it (for now, anyway). I'm sure anyone dedicated can go undetected, but it's the really low-effort stuff that's generally the problem. If you can catch some of it, that's something at least.
- jaco6 3mo agoThe best method is, as always, an anti-privacy method. Simply track all citizens' writing patterns throughout their life, from cradle to grave, then diff with any given text's signature--you'll know if it was human written or not. Better--opt in--install a "personal text signature" on your devices, sign things that you wrote yourself with it. But I suppose that's just like the image provenance chips on cameras. Either way father fascism is more with us than ever, praise him!
- energy123 3mo agopow(n,m) where n is alphabet size and m is number of characters is very dense.
- TacticalCoder 3mo ago> ... but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading Most people's issue with AI-generated llmish however is not that it's AI-generated. It's its insufferable tone. So if we get to a point where we have to read tea leaves (an image you seem to appreciate) to determine if it's llmish or not, we'll have won by then. Really: it's that full-on asshole tone I (and many others) want to see disappear from blogs, comments, LinkedIn, etc.
- dofm 3mo agoMy broad feeling is that if we generally independently identify the insufferable tone, and we absolutely can, so can a sufficiently trained machine-learning model. It's an aside, but my biggest problem with trying to get up to date and learn about LLMs is how much of the documentation, blog writing, and tutorial material has obviously been written by no-one. It is just so much harder to read (and, like generative AI slop generally, curiously much harder to recall later).
- bob1029 3mo agoWith sufficient information you can derive a signal even in the presence of overwhelming noise. Assuming the noise is not perfectly correlated with the signal this is always possible. Schemes like GPS, CDMA and DSSS are based upon this concept. GPS in particular is quite impressive in its ability to recover information that is received below the thermal noise floor.
- Ferret7446 3mo agoThere has to be a signal to detect it. Take this sentence: Bob went to the store to buy milk. Was that AI generated or not? There simply isn't a signal there. The problem isn't noise, the problem is, is there even a signal to begin with. Sure, you might be able to recognize the quirks of a specific LLM just as you recognize the quirks of a particular person, but as the number of LLMs proliferate, then the signal turns into noise. (The signal isn't buried by noise, it becomes noise. The signal no longer has any discriminating power.)
- bob1029 3mo agoStatistical power comes from having many samples. I agree that having just one sample doesn't take you very far.
- sdenton4 3mo agobut.... the LLMs are actually all trained on approximately the same stuff, and tend to have similar quirks. In the human world, writers develop recognizable voices, which are detectable and classifiable (as in the article we have all supposedly read). Furthermore, we don't necessarily care about telling one LLM from another, just that they aren't human. That's different from trying to identify one human amongst a sea o fhumans, or one bot form within a sea of bots.
- Tade0 3mo agoThe article itself explained that it was much easier to classify text as human or LLM generated than to have "human" as just a category along with all the different LLMs as it's likely the LLMs are distilled from each other, creating a unique footprint. If a signal is weak, it might not even appear in every sentence, but that doesn't mean it doesn't exist. For instance, I don't recall ever consciously using an em dash, but you'll probably need an entire paragraph to find one in LLM-generated text. My own sense of whether text is generated is partially based on its sheer length - humans typically don't bother writing so much.
- itemize123 3mo agoobviously, a universal model doesn't exist since the signals are non-stationary but it's way better than what tarot reading
- ekelsen 3mo agoSee the discussion on https://news.ycombinator.com/item?id=48837460 https://news.ycombinator.com/item?id=48837460 You can absolutely still tell.
- Jolter 3mo agoSo you’re saying that the linked article’s findings are implausible? Is the article fake, then, in your opinion?
- seanmcdirmid 3mo agoIf you have access to the detector, you can formulate a generative solution that avoids being flagged. Which gets me wondering why don’t model providers do that? There must be something about that that destroys semantic weights somehow.
- lemagedurage 3mo agoWhy would sounding human be a goal rather than a byproduct of trying to communicate efficiently?
- SiempreViernes 3mo agoOne of the big use case is cheating on your essay assignment.
- seanmcdirmid 3mo agoThat is manager/executive/manager speak, real people don’t speak like that (unless they are in the aforementioned roles).
- lemagedurage 3mo agoI just feel like, from the POV of AI companies, that reducing the amount of em dashes they use to "blend in" more and talk more human-like, for the sake of being less detectable, wouldn't be a big priority.
- wzdd 3mo agoThe article discusses a technique by which the author achieves high accuracy at detecting AI written text. Unless you have a problem with their experimental method, this is the opposite of tarot card reading. > we are well into undetectable sophistication with today's models The article directly contradicts this, as do you, in your previous paragraph: "Sure you might be able to detect today's tells". The article is literally about a technique that detects today's tells. Your comment is mostly expressing doubt that this technique will work reliably in the future, but it's framed as opposition to the article, which it's not: the article is about detecting today's AI-written text, at which it seems to be quite successful.
- mschild 3mo agoIt does achieve high accuracy but I think given the context when one wants to know this information, plagarism for research papers and college/highschool essays and work, it's unfortunately not good enough. My neighbour is a teacher. She has a really good idea which of her students uses AI to do their homework but 80% accuracy is not good enough. She'd need to be able to prove it with certainty.
- SiempreViernes 3mo agonot really, a strong suspicion is enough to motivate assigning an extra paper and pen in person test to a student, and then you can fail them on that result.
- CalRobert 3mo agoThat seems pretty unfair. Why not make the original test pen and paper then? (Or at least a typewriter, offline computer, etc - my handwriting is awful)
- projektfu 3mo agoOriginally the idea was that a take-home essay would allow students to work at their own pace, study their own way, and produce something interesting. But if more than half will just prompt an AI and learn nothing, then I agree you should proctor all assessments.
- lelanthran 3mo ago> Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. This is simply untrue, and completely divorced from reality. Tarot card readings have literally zero predictive success. Last I checked, LLM-detection had a +90% success.
- dofm 3mo ago> Text is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. This does not sit well with personal experience and I wonder if it is just one of these questions of AI people being unaware of the level of skill that exists in domains they think have been automated. It is of course possible that my tendency to spot LLM-written text has much to do with the way that it sounds like an averaged Californian college student to my British grammar-school-educated ears, as so many of the situations where I am encountering AI text are Brits using it without apparently realising they are giving themselves away. But I know people who don't have particular technical skills in this sphere or a grammar-school background who also have an uncanny knack for pointing out LLM-written text. > Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's. I especially don't think this is true. Will they be able to do it in the future? Maybe. Is it possible to prompt a current cloud LLM to write in a way that is obvious? Yeah. (IMO Gemma 4 writes less detectably than most of them!) But my instinct is that someone with any facility for language is going to be better than chance at spotting LLM-written text once it is three or four paragraphs long. So I think it should be possible in principle to train machine learning systems to detect those patterns.
- CalRobert 3mo agoIf you can train a system to detect these patterns, presumably you can train systems not to generate text which matches them? I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA…
- dofm 3mo ago> If you can train a system to detect these patterns, presumably you can train systems not to generate text which matches them? I don't know. I mean, it feels like the systems that would detect them are likely qualitatively different to the machines that make them. One of the things that feels obvious to me is that LLMs are always going to write in a new way, because words do not get all that close to perfectly conveying the inner thoughts of competent writers. Competent writing is always a battle to find the better word, or even to create it. So sure, you could add another adversary that the generator has to satisfy, but "this sounds like a machine wrote it" is only an observation; it's not a prescription for not writing like a machine. Maybe it's never going to be possible. > I do struggle at times with thinking my own writing looks like AI. But I’m an average Californian who went to college half way between SF and LA… :-) You guys do just sound a certain way, in the same way Brits sound a certain way to you I expect. But I think the reality is that the final stage of training LLMs was largely done in a Californian voice and with rather Californian communication objectives. (Though equally I think much of what I am detecting is more Madison Avenue than Palo Alto)