6 ms·
Ask HN: How do systems (or people) detect when a text is written by an LLM
Hello guys, just curious about how can people or systems (computers) detect when a text was written by an LLM. My question is mainly focused to if there is some API or similar to detect if a text was written by an LLM. Thanks!!!
- dipb 6mo agoHumans detect them mostly through pattern matching. However, for systems, my guess is that a ML model is trained on AI genres texts to detect AI generated texts.
- moonu 6mo agoPangram is probably the best known example of a detector with low false positives, they have a research paper here: https://arxiv.org/pdf/2402.14873 https://arxiv.org/pdf/2402.14873. They do have an API but not sure if you need to request access for it. For humans I think it just comes down to interacting with LLMs enough to realize their quirks, but that's not really fool-proof.
- spindump8930 6mo agoPangram has time after time been shown as the only detector that mostly works. And that paper is pretty old now! There are recent papers from academics independently bench-marking and studying detectors e.g. https://arxiv.org/abs/2501.15654 https://arxiv.org/abs/2501.15654
- Someone1234 6mo agoThey cannot. Unfortunately many believe they can, and it is impossible to disprove. So now real people need to write avoiding certain styles, because a lot of other people have decided those are "LLM clues." Bullets, EM Dash, certain common English phases or words (e.g. Delve, Vibrant, Additionally, etc)[0]. Basicaly you need to sprinkle subtle mistakes, or lower the quality of your written communications to avoid accusations that will side-track whatever youre writing into a "you're a witch" argument. Ironically LLM accusations are now a sign of the high quality written word. [0] https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
- loloquwowndueo 6mo agoThe key insight is to avoid – em dashes. You’re absolutely right. It’s not the content, it’s the style.
- LoganDark 6mo agoThat's an en-dash.
- loloquwowndueo 6mo agoSorry! Is this ok? —
- singpolyma3 6mo agoYou're absolutely right. That is an em dash
- LoganDark 6mo agoYou're absolutely right. They are absolutely right
- sumeno 6mo agoYou're absolutely right! I unintentionally used an en-dash instead of an em-dash. Here is the em-dash you requested: –
- sanex 6mo agoIronically one of the big tells for me is the "It's not this. It's that." Your comment uses a comma though so you're probably a real person :)
- rcxdude 6mo agoI assume they were aping those terms ironically (especially given the 'you're absolutely right')
- haarlemist 6mo agoI
- PufPufPuf 6mo agoI "detect" them through overuse of some patterns, like "It's not X. It's Y." This is an artifact of the default LLM writing style, cross-poisoned through training on outputs -- not an "universal" property.
- deleted 6mo ago[deleted]
- mjlee 6mo agoPeople Look For: Specific language tells, such as: unusual punctuation, including em–dashes and semicolons; hedged, safe statements, but not always; and text that showcases certain words such as “delve”. Here’s the kicker. If you happen to include any of these words or symbols in your post they’ll stop reading and simply comment “AI slop”. This adds even less to the conversation than the parent, who may well be using an LLM to correct their second or third language and have a valid point to make.
- booleandilemma 6mo agoI'm not going to tell you. I don't want that information going into the dark forest :)
- m_w_ 6mo agoI don’t think there’s a reliable system or API for doing so, unclear that arms race will ever favor the side of the detectors. As far as how I / other people do it, there are some obvious styles that reek of LLMs, I think it’s chatgpt. There’s a very common structure of “nice post, the X to Y is real. miscellaneous praise — blah blah blah. Also curious about how you asjkldfljaksd?" From today: This comment is almost certainly AI-generated: https://news.ycombinator.com/item?id=47658796 https://news.ycombinator.com/item?id=47658796 And I'm suspicious of this one too - https://news.ycombinator.com/item?id=47660070 https://news.ycombinator.com/item?id=47660070 - reads just a bit too glazebot-9000 to believe it's written by a person.
- elC0mpa 6mo agoThanks a lot for the detailed answer, will take a look at the examples
- huflungdung 6mo ago[dead]
- volume_tech 6mo ago[dead]
- neurworlds 6mo ago[dead]
- dezgeg 6mo agoFor HN comments, the LLMs seem to really like 2 or 3 paragraphs long responses. It's pretty obvious when you click a profile's comments and see every comment being that exact same structure.
- RestartKernel 6mo agoPeople look for tells, systems detect word distributions. Though neither is as reliable as active fingerprinting using an encoded watermark.
- deleted 6mo ago[deleted]
- rcxdude 6mo agoThere are some systems which can use the LLMs themselves to detect writing (basically, if the text matches what the LLM would predict too well, it's probably LLM generated), but they are far from infallible (with both false positives and false negatives). There's also certain tropes and quirks which LLMs tend to over-use which can be fairly obvious tells but they can be suppressed and they do represent how some people actually write.
- block_dagger 6mo agoEm dashes, “it’s x, not y”, excessive emojis and arrows.
- mghackerlady 6mo agoEspecially where the emoji serves practically no purpose other than to get your attention. If it is especially abstract what the emoji is there to represent, I start looking for other signs
- blanched 6mo agoI don't think there's any reliable way to tell. To me, it often feels like the text version of the uncanny valley. But again, that's just "feels", I don't have proof or anything.
- mghackerlady 6mo agoOveruse of "it's not X, it's Y" kind of writing, strange shifts in writing or thinking patterns, and excessive formatting (or, when I'm on wikipedia especially, ineffective formatting (such as using MD where it isn't supported))
- rwc 6mo agoContrastive negation continues to be a dead giveaway.
- tatrions 6mo ago[flagged]
- leumon 6mo agoYou can try to use an ai detector, here is a leaderboard of the best ones according to this benchmark: https://raid-bench.xyz/leaderboard https://raid-bench.xyz/leaderboard Results should of course always be taken with a grain of salt, but in most cases detectors are quite good in my opinion.
- gwbas1c 6mo agoI don't think you can 100% detect AI content, because at some point someone will just prompt the AI to not sound like AI. I think the better question to ask is: What are your goals? Is it to prevent AI SPAM, or to discourage people copy-pasting AI? Those are two very different problems: in the case of AI SPAM you look for patterns of usage, (IE, unusually high interaction from a single IP, timing patterns around when things are read and the response comes in,) and in the other case it all comes down to cultural norms.
- Havoc 6mo agoYou don't really. There are a couple of tells like em dashes and similar patterns but you should be able to suppress that with even a simple prompt.
- noufalibrahim 6mo agoIt's a lot easier to detect when you mostly interact with non English speakers. I asked an LLM to rewrite this to make it nicer and got the following. I'd flag the first because I don't usually hear "majority of your interactions" in conversation but I might miss it. The second will probably get by me. As for the third, I never say "considerably easier" unless I'm trying to sound artificially posh. 1. It becomes much more noticeable when the majority of your interactions are with non-native English speakers. 2.It tends to stand out more when most of the people you interact with speak English as a second language. 3. It's considerably easier to identify when most of your interactions involve people whose primary language isn't English.
- sigotirandolas 6mo agoI don't look at whether the text is written by an LLM but at whether it has substance and whether the writer understands what they are doing and is respecting my time. If the text is full of punchy three word phrases or nonsense GenAI images then that's an obvious sign. But so is if the other person has some revolutionary project with great results but they can't really explain why their solution works where presumably many failed in the past (or it's a word salad, or some lengthy writing that doesn't show any signs of getting you to an "aha, that's some great insight" moment). A good sign is also if the author had something interesting going before 2022, and they didn't fall into the earliest low quality LLM waves. Unfortunately some genuinely talented people have started using LLMs to turbocharge their output while leaving some quality on the table nowadays, so I don't really know. I'm becoming a lot more sceptical of the Internet, to be honest.
- fwip 6mo agoYou can smell it.
- vednig 6mo agohttps://detect.ai/ https://detect.ai/
- johnwhitman 6mo ago[dead]
- DavideNL 6mo agoI was wondering the same today, when i got this search result from Kagi: https://linuxvox.com/blog/conntrack-linux/ https://linuxvox.com/blog/conntrack-linux/ Note that on Kagi, you can click "Report this page as AI-generated" [1]. Unfortunately though, my last report from January is still "under review" :/ [1] https://help.kagi.com/kagi/features/slopstop.html https://help.kagi.com/kagi/features/slopstop.html
- SyntaxErrorist 6mo agoInstead of trying to detect AI in the final string of text the industry seems to be moving towards proof of process. version history, drafts protocol or even UI level logging of keystrokes. If i can not prove i spent three hours in a doc via a series of incremental diffs, the humanness of my prose becomes irrelevant in a high stakes environment. Detection is a lagging indicator the only leading indicator is the audit trail of the labor itself.
- raw_anon_1111 6mo agoI bet you a paycheck that of anyone read one of the “97 things” books, they would think the essays are AI generated even though they came out way before LLMs. https://github.com/97-things/97-things-every-programmer-should-know/blob/master/en/SUMMARY.md https://github.com/97-things/97-things-every-programmer-shou...
- aimadetools 6mo ago[flagged]
- cadamsdotcom 6mo agoThis is a handy resource - for humans and for bots: https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing