7 ms·
Detecting LLM-Generated Texts with “Classical” Machine Learning
- cyanydeez 3mo agotoday, sure. Tomorrow, the LLMs will be training the humans thought patterns that will directly start skewing their natural writing. Generation alpha is going to have a lot of trouble if we keep perpetuating the myth that you can really interpret text in an ongoing fashion.
- pixl97 3mo agoI think you're about a year late for this revolation. https://www.washingtonpost.com/opinions/2025/08/20/chatgpt-claude-chatbots-language/ https://www.washingtonpost.com/opinions/2025/08/20/chatgpt-c...
- cyanydeez 3mo agoI'm not late if people constantly put effort into finding LLM text, or every other comment on hacker news is either about something being LLM generator.
- pixl97 3mo agoAfter seeing comments on hacker news attempt to call an article from 2015 as generated by an LLM, I have very little faith in commenters having any ability in actually detecting AI written text. And that's just one particularly egregious case I remember. Posters that are technical writers or use English properly get called bots quite commonly when their post history shows a writing style going back over a decade. But now that LLMs are causing a language drift in English users our filters of "that's an LLM" will become even more useless.
- unfocso 3mo agoI had done the same for classifying and generating bookmarks of thousands of datasheets, along with a very naive yolo-based classificator (to detect pages made out of diagrams and pictures mostly). Done with GLM-OCR, I had to watch text sloooowly crawl out of the llm and still have to live with hallucinations and the model not following the schema
- Krssst 3mo agoThe classifier does not seem so big, I wonder if something like it for English could be used in a browser extension to run against every single paragraph being displayed ? If the internet is going to drown in LLM text it would be nice to have tools to detect that automatically just like we have adblockers today to avoid wasting time on ads. (the article was a good read, thanks!)
- xiaoyu2006 3mo agoI assume different models will have different distribution, so it has to be kept updated?
- Krssst 3mo agoThe article mentions that AI texts are often caught by multiple models, so hopefully text from newer LLMs could still be caught without updating the model?
- pixl97 3mo agoYou know what GAN is, right? In training all you have to do is take their model as the adversary and then it's useless.
- cygn 3mo agoI built a browser extension that does this, well for posts on twitter, hackernews, reddit etc. If you want it for all text, it would also be feasible. I use a quantized mini-LM model that runs very fast and classifies eg your whole twitter feed in a couple of seconds. Check it out: https://slopsieve.com/extension https://slopsieve.com/extension Accuracy is also much higher than this approach here. 0.9944 AUC, 0.966 acc@.5, 0.971 F1@.5
- Anoian 3mo agoWent to the website and inserted the blogposts I wrote in the last year. I had a pretty good understanding which of my blogposts had more reworks and which ones had entire passages being generated using AI and then left as is because I was happy with them. None of my articles were over 50% according to your model but the ones that I know took me a long time to write, even though I used lots of AI in the creation of them, hit below 10%, probably because I hand-edited them a lot. Overall a nice website, thanks for sharing :)
- aberoham 3mo agoI wonder about this technique vs simple SVM classifiers: https://x.com/rosmine/status/2056406399471558872?s=20 https://x.com/rosmine/status/2056406399471558872?s=20
- janalsncm 3mo agoThis article is about training a classifier to detect synthetic text. The link you sent is for generating text which attempts to defeat those classifiers.
- akersten 3mo agoText is simply not information dense enough to be able to decode some arbitrary signal of provenance from it. Sure you might be able to detect today's tells (particular sentence structures preferred by Claude, phrases, etc) to get you some arbitrary chance percentage it was machine generated, but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Images, absolutely, there are tell-tale artifacts from today's generators that simply aren't emitted by "natural" paths to create them, and you can "detect AI" with high confidence (for now). Words, no, the signal is far too sparse and we are well into undetectable sophistication with today's models, let alone tomorrow's.
- jgalt212 3mo agoIt depends on how much text. For example, chardet often falls down on short strings, but 1K characters it nails it.
- stymaar 3mo ago> but it's a bad fiction to perpetuate that any of this is anything more than tarot card reading. Hard disagree. LLMs (especially base ones, that only received pre-training) can produce output that is undistinguishable from human writing (because that's what they were trained to do). But commercial chat models are specifically tuned in a way that maximizes user engagement. It's that specific tuning that is very easy to spot when reading AI slop, and that's not surprising that it's easy to spot automatically either. And I don't think that's going to change anytime soon, unless their incentives change. (We can say exactly the same thing about man-made stuff optimized for a specific purpose, like stock photography, clickbait titles or industrial food: they aren't stereotypical because their creator lacks the skill to make them otherwise, they are like that because that's what works best).
- empath75 3mo agoIt does mean that this will have a drift problem if it's just trained on the idiosyncrasies of model fine tuning. That's fine! But it is something to be aware of.
- 3mo ago
- XiphiasX 3mo agoAnything too “clever” and “snappy” = instaLLM
- hasteg 3mo agoThis is also how I pretty much filter LLM generated text in my head.
- teeray 3mo agoThe problems are simply too great if an LLM detector has any false positives at all. Imagine how soul-crushing writing an entire dissertation by hand and having it rejected because some “good enough” LLM detector decides you write too much like an AI.
- dmurvihill 3mo agoIt depends on the application. Dissertation? Hell naw. Blog post? Absolutely, run it through that thing.
- teeray 3mo agoThe problem is that ed-tech is absolutely ravenous for an LLM detector and would rather use snake oil than accept that it might not be possible.
- rayval 3mo agoAs I recall, a few years ago (in the era of first generation LLMs), a professor in Texas used an anti-plagiarism tool that flagged more than one-third of the class using AI in an exam, and used that finding to give them a failing grade. If memory serves, one student objected strenously and ran the professor's own work (published 10 years earlier) into the same tool and it flagged that work as AI-generated. EDIT: HN item from June 2023 https://news.ycombinator.com/item?id=36215823 https://news.ycombinator.com/item?id=36215823
- pixl97 3mo agoExactly. The more corporate and proper you tend to speak, the more likely it's to classify you as an LLM. It's like the classifiers want us to talk like trash at their current rate. This seems to be really problematic for ESL speakers/typers that may have been trained on a smaller, more proper subset of the language.
- Krssst 3mo agoWe can measure false positive rate. The detector in the arricle is 85% accurate (not sure about false positives, but let's assume) which is too low to make conclusions, but enough when browsing the web and skipping reading likely-slop withiut accusing anyone. If the false positive rate becomes <1% then it's better. The alternative is the world drowning under slop so I'd rather have imperfect detectors and have users aware they may fail in rare cases to avoid witch hunts. The general issue is that people only realize they're reading slop halfway through which is frustrating. If you know it from the start thanks to a detector and move on without commenting, no time waste, no frustration, less negativity towards LLM users.
- gleenn 3mo agoI think the fundamental problem is that training current SOTA AI models is very expensive. If a simple "classical" model can detect them, presumably at much lower algorithmic cost, then why wouldn't the model trainers use these same tools to feed back into their models to improve them at low cost to make them better? It's an arms race. Any cheap pattern can and presumably will be used to retrain if it becomes and effective way to catch AI.
- Retric 3mo agoIt’s an arms race where the AI companies are at an extreme disadvantage due to relative training costs.
- arjie 3mo agoIt’s simply not a priority. The labs can do many things. Making text non-LLM is not really that useful. Analogous to Facebook not picking up the obvious $20 bill in front of them. It’s because they’ve got $100 bills at their feet they’re picking up.
- pixl97 3mo agoNot a priority currently. Selling services to spammers... I mean marketers is still big money and eventually someone will pick it up. If training costs ever drop, then it's one of the first things that will happen.
- OtherShrezzing 3mo agoIn part because model vendors specifically prefer when people think that lots of content is produced by their model. The more Claude-like writing appears on the internet, the more signal there is to investors that people are using Claude for a greater number tasks.
- t-writescode 3mo agoCould also be a problem of the form of P=NP. Validating might be very easy, but writing might be hard. Like the traveling salesman problem. It’s very easy to tell whether a specific path takes N units of time, but it’s hard to figure out if there’s any path, among all possible paths, that takes N units of time.
- docheinestages 3mo agoI think figuring out if a text is AI-made is a losing battle. What could work is gauging how much effort went into writing the text, regardless of who the author might be. What's easy today is generating mountains of text that are extremely hard to read. What requires effort is knowing how to engage the reader, how to keep out extraneous information, and how to keep the text as short as possible without losing details. That needs effort, with or without AI.
- calebh 3mo agoThe easiest way is to keep track of the text's edit history, keeping a block of edits over time and having them signed by a timestamp authority. The final edit history can then be inspected by some external authority, then signed if the edit history looks human. I have a blog post from 2023 on this topic: https://helbl.ing/Written-Proof-of-Work/ https://helbl.ing/Written-Proof-of-Work/ For Google Doc users, you can already inspect the edit history over time to verify that text is written by a human.
- visarga 3mo agoThat human might have used AI. You can never know. Hand fixed AI output, human just polished the corners? Light rewording of a full text written by hand, because the author is not confident in their writing? Actual human text, but after researching with AI?
- theoreticalmal 3mo agoExactly. Detecting AI writing is an arms race that can only end with detection coming in second place.
- warkdarrior 3mo agoI am working on a browser extension to help with that. Basically it interposes on any text field and canvas and if user pastes a large amount of text (copied form example from a chat bot), the extension will "replay" that text at normal, human-editing pace, and introduce typos that are fixed through later edits.
- metalman 3mo agothere is not much point in detecting LLM generated text, in that humans are useing info from LLM's, but obfusicting it's origin, with there own garble, along with purely human garble, and almost(but not quite) human LLM product meaning that the threshold for rejecting "data" must be lowered, which personaly means a very very low tollerance for wierdness, except where it can yield imediate possitive cash flow for the rest I do my own research and verification thank you very much
- mike_hock 3mo ago2 misplaced apostrophes, 8 spelling errors — definitely human output
- 40four 3mo agoI could be wrong, but I just don’t see how trying to “detect” LLM generated texts is ever going to work. The only thing that makes any sense if you truly want to have confidence a human wrote it is some type of “proof of work“ system. I think there’s a lot of interesting ways to approach the proof of work problem with different pros and cons, but that is where our energy should be focused if we seriously want to solve this problem.
- IshKebab 3mo ago> I just don’t see how trying to “detect” LLM generated texts is ever going to work He literally demonstrated a working system in this post. Do you mean you'll never get to 100% accuracy? Clearly, but you don't need that.
- 40four 3mo agoI just mean no matter how hard anyone tries, I don’t see how useful these systems would be in practice. Sure they demonstrated a “working” system. Plenty companies sell products that “work” to one extent or another. But how useful is it really to get a result of “This is 80% likely chance of being LLM generated”? Or 75%, or 95%? What if the text is a mix of human written text and LLM text? How would you even begin to test that? I suppose a text that is half human half LLM would theoretically score in the 50% range, but do you see the problem? You can slap a confidence % score on a test run, but interpreting the results leads to a whole other can of worms. Point is there are so many variables, and it’s not clear that the result from any of the systems is even valid or applicable to help you make a decision in a real life situation.
- IshKebab 3mo agoDepends on the application. I would love to have that percentage next to HN submissions so I don't waste time reading (or starting to read) obvious slop. Doesn't really matter if I occasionally skip something that isn't actually slop.
- jaco6 3mo ago
- richard_chase 3mo agoAm I the only who largely enjoys the output of LLMs more than most stuff written by humans? I find myself coming back to old chats with ChatGPT frequently because the output is amazing.
- therealdrag0 3mo agoI wouldn’t go that far… but it can be kinda like Wikipedia, clean and readable.
- arjie 3mo agoNeat. I will implement something like this for myself. I just need to reduce the spam a little. Imperfection is okay for a social network context like HN.
- pixl97 3mo agoIt will work for a bit, but as people start speaking more like LLMs and LLMs start training using said classifiers as a GAN, it will become useless.
- Krssst 3mo agoIf we get precise detectors and LLM posts don't get shown by social networks recommendation algorithms as a result, the chances of people starting to talk like LLMs get lower.
- cygn 3mo agoyou can try my browser extension which does this for hackernews: https://slopsieve.com/extension https://slopsieve.com/extension
- connorboyle 3mo ago> Eventually, I faked my way through the thesis, and life moved on. This is a very startling admission! I checked the Chinese (original?) version of the post, and saw the author uses the word "糊弄" (in the place of "faked"); I'm not a native speaker but I think this may come across more as a self-effacing comment on the low quality and/or effort behind their thesis, whereas the English version implies fraud. May be wise to change this!
- jshmrsn 3mo agoWell cheated would definitely imply fraud. “Faking it” as in “fake it till you make it” is more like pretending you know about a topic until you learn enough on the job to participate competently.
- hgoel 3mo agoI don't know if the Chinese text implies something different, but I think even in English it's pretty normal for people to claim they 'faked' their way through something without referring to fraud. E.g. "I faked my way through the interview!" = "I did my best to respond to questions I did not feel fully prepared for, and managed to get through the interview"
- woadwarrior01 3mo agoSmall encoder-only transformers are excellent at classifying LLM-Generated Text. I built an on-device iOS app using a custom small encoder that achieves an AUROC of 99.81 on RAID-bench.
- throaway54321 3mo agoI don’t think it actually matters (and it’s a losing strategy as others have noted). This issue with AI generated stuff is that that it’s sometimes asymmetric: either the author worked very little to produce a lot of slop and now the reader(s) all have to do the heavy effort of reading it OR the author puts a little extra work in once and resolves all future readers’ burden. If it was possible to boil down an artifact into a prompt + some resources that would be an interesting tool, or at least some way to tell if some artifact is “worth my time to read”
- moxza 3mo agoThe thing I find most encouraging is that the best AI detector is still humans. Don't write the Turing test off yet. From what I understand, your approach is clever, it's like an accent detector. Known models tend toward a specific median approach. Humans have a much richer degree of randomness. Riffing on Anna Karenina... All models are alike in that they present predictable patterns. Humans inevitably write in unique ways. I gave a lot of thought to the idea that humans will devolve to the median led by volume of AI interactions, but in the end, I think we're still interacting with each other when not at work/on machines, and the fact that we even have a genetic heritage is always going to differentiate us.
- gdiamos 3mo agoas soon as you release a way of measuring it, you give LLMs a signal to optimize
- maxspero 3mo ago> Sounds promising, right? I spent some time trying [perplexity], but results were disappointing—plenty of false positives and false negatives, and no reasonable threshold could be set. Perplexity was widely considered SOTA in 2022. One part of it is because everyone was evaluating on open models or closed models that were still close (i.e. GPT-2 vs. GPT-3.5). Today, the gap is so much wider between the models you can use to compute perplexity and the frontier models people actually use. Also so many AI text detection papers used a strawman RoBERTa baseline that was very undertrained for the task. The synthetic mirrors method for data generation used here is the same as what we use at Pangram. Good blog post, thank you for sharing!
- sMarsIntruder 3mo agoAm I wrong or it doesn’t seem to detect the em-dashes as clear warning signal?
- hacker-matrix 3mo ago[flagged]
- m00dy 3mo agoReddit is doing this really good.
- clickety_clack 3mo agoIt looks like the text that this classifies might be Chinese, is that right? Do Chinese speakers have the same cultural aversion to AI-generated text? I’m wondering if there might be a different level of effort put into making text seem human-generated in English v. Chinese.
- poisonfountain 3mo agoMy theory is that the labs are RL'ing models to output easily classifiable text so they can avoid model collapse when training the models on data scraped from the web. Of course, you can skew the distribution with some effort and generate text that avoids even the best classifiers out there (like Pangram), but even tech-savvy people aren't usually doing it (see the amount of AI-written posts that end up in HN and get tons of comments complaining about AI mannerisms), so I guess they're successfully avoiding like 99% of the slop using such classifiers. I don't think it's in the interest of the labs to allow you to generate text that's indistinguishable from human prose. Especially since nobody would pay $1,000/mo just to generate text - but would do so for tasks like coding.
- rtrgrd 3mo agoOh man all of his Chinese humour is lost in translation T_T
- amai 3mo agoShouldn't there be a Kaggle contest for this?