5 ms·
How we measured AI writing across arXiv, and where the measurement breaks
- dopamine_daddy 2mo agoI scored the full text of 12,750 arXiv papers from 2021 through 2026 to find out how many of these get flagged as machine written and how much it increased since the release of chatGPT. I purposely tuned the detector to avoid false positives. My detection rate pre chatGPT is around .4% for that reason. The biggest results: in Jan of 2026 about 39% of papers got flagged as AI written. In computer science speicifcally the peak was at 65%. Mathematics barely moved away from 0.7%, though the proof heavy math texts might just not get picked up by the detector properly. All this is a detector estimate of a statistical signal and not a proof any given author used AI. Machine written can also mean heavy AI-assisted editing.
- economistbob 2mo agoThank you for this work.
- phreeza 2mo agoThis is a pretty stunning result. The time series looks really convincing. Is the way the detector itself is trained orthogonal to this or could there be some "leakage" in that the pre-chatgpt text is in the (positive) training data?
- dopamine_daddy 2mo agoI tried my best to avoid leakage. If you're curious about how I trained the detector I have a writeup on it: https://unslop.run/blog/how-our-ai-text-detector-works https://unslop.run/blog/how-our-ai-text-detector-works FYI this is all relatively new so there might be lots of issues and iterations coming.
- Paracompact 2mo agoMy honest first impressions, since I think the project is well intended: This writeup is itself AI, and I would venture to call it slop. The Calibration section is borderline uninterpretable, and I challenge any non-author who claims to understand it to answer some basic peer review questions about it.
- simonreiff 2mo agoI have to concur, unfortunately. This isn't reliable, reproducible, or interpretable.
- andycasey 2mo agothis is neat! is the model available somewhere for local execution? or even a lookup table with your results for all arXiv pre-print codes? I want to run it on lots of pre-prints and I don't want to kill your server
- dopamine_daddy 2mo agoThank you, I plan to release the arxiv preprint codes. Also don't worry about killing my server, let me know if you succeed :D
- exe34 2mo agoWouldn't that provide an excellent signal to tune models to be less like AI? I suppose that's a good thing ultimately.
- paxys 2mo agoWhat detector are you using? How can you be sure of its accuracy given that every commercial AI detector has been debunked?
- JamesBarney 2mo agoAlmost every. Pangram is pretty accurate on longer texts. I haven't seen any glaring examples of false positives.
- adamgordonbell 2mo agoThis! Pangram is very good. They claim 1 in 10,000 fp rate. I had to change my mind on AI detectors after playing around with it. It would be interesting to hear how this detector compares. It also seems to be aiming for low fp rate.
- malshe 2mo agoI’ve heard good things about Pangram but it’s trivially easy to fool it. I generated text using Opus 4.8 and then humanized it using Grammarly. Pangram determined it 100% human-written. Tried it numerous times with the same result. GPTZero claimed with 80% confidence that it was human-written but AI polished. Nothing in that text was human-written.
- JamesBarney 2mo agoHow long was the text? And mind sharing?
- deleted 2mo ago[deleted]
- jmcqk6 2mo agoExtend it to scan papers from pre-2020. That should give you a better baseline accuracy for your detection system.
- 0x000xca0xfe 2mo agoCould you share some pre LLM false positives? Would be interesting to see what is tripping the detection. Did a tiny fraction of authors write like LLMs, before LLMs?
- cgio 2mo agoA signal, I put an AI generated of mine (with heavy human guidance on aesthetics mostly and some minor human editing) and got 5% only.
- drewcrawford 2mo agoCouple of methodological notes * pre-chatGPT is not an effective control because language evolves. In the arxiv corpus in particular there are "fashions" in research depending on what gets funded lately, not to mention many new words and topics not invented before a given paper. * In general, detecting AI from content seems difficult as humans write like they read. To the extent there are unique factors recognizable as AI and to the extent humans read them, they will eventually incorporate them into their writing style. Accordingly, you'd need to model a rolling window of "AI tells" that decay at some rate.
- RugnirViking 2mo agoHave you considered trying the same sort of analysis between two time periods pre-ai? What I mean is, are there statistical differences in how people write that would give similar looking results between say 1990-2000 and 2000-2008? Are we just detecting the natural progression of language here?
- jrm4 2mo agoThe important question is: So what? Genuinely. I get that there may be some visceral reaction against this, but when I break it down, I mostly fail to see the problem. Seems like what is actually important is: Compared to before, when a human reads it, do they -- or society -- get something good out of it? Is it worth it to add this to the "pantheon?" If that's not what's happening enough, and if this doesn't describe the process -- then the problem lies elsewhere, no?
- economistbob 2mo agoThe problem is that humans start with credulity. AI hallucinates and makes up stuff some percentage of the time. Humans are not "default deny" when given information. Unleashing that was a disservice to mankind and has created a future filled with lies and people who are confident in them.
- jrm4 2mo agoOh, to be technically correct: AI hallucinates and makes up stuff 100% percent of the time. Never been a fan of that word for this. Again, I fail to see the problem here that isn't solved by careful reading WHICH IS WHAT PEOPLE SHOULD BE DOING ANYWAY. I would like to see room for AI disclosure, maybe a statement of "this is how much AI I used." But this blanket X% of this looks like AI? Again, so what?
- JSR_FDED 2mo agoI find the fact that 39% of papers use a style that signals lack of effort somewhat worrisome. Even if that’s 100% wrong, and they’re all high effort papers, the fact that they give off the same aura as low effort papers is a problem for the authors.
- jsrozner 2mo agoIf as a primary school student you had needed to reverify every fact /experiment that was presented in your textbooks, you would not have gotten very far. Trust is very important to human progress.
- willquack 2mo ago> If a tool marks 40% of new papers as machine-written but also marks 20% of papers written before ChatGPT existed, the real story is the 20% nobody mentioned. When 65% of the papers you read have the characteristics of being AI written, whether or not you use AI to write, your writing will be influenced by the AI style. I imagine this must be particularly the case for newbie researchers who are still developing their writing style
- serial_dev 2mo agoI might be missing something but what’s the real story of the 20%. To me it sounds like 1. Either your tool is just not that good and reliable as you thought, 2. AI is trained on human written articles, so some of that human written content informed the now established “AI slop”. There are people who shipped “slop” before AI.
- cisophrene 2mo ago> There are people who shipped “slop” before AI. The funny thing is that "slop" was defined by the writing habits of AI model, which we have learned to pick upon and recognize. The "It's not X, it's Y", the rhetorical questions and other patterns would have been the tools of a skilled writer, and those people writing "like AI" before AI most likely would have been recognized as such.
- aionwikipedia 2mo agoeven if your writing is influenced by "the AI style," it would likely only show up in patches, a word or rhetorical flourish maybe, but AI-generated text doesn't "sound like AI" in patches, it does so consistently across the entirety of the output. it's very hard for a person to maintain the syntactical patterns of AI writing over an extended period of time unless they know exactly what those are -- probably more knowledge than any person currently has, really -- and are applying a level of detail to each clause comparable to forging a painting. there's also the fact that "the AI style" has changed over time. for instance, the word "delve" is notorious as an "AI sign," which it was up until mid-2024, at which point it dropped off sharply and has now basically disappeared from LLM output. so if someone happened to pick that up due to reading it everywhere, their writing is now less characteristic of AI, not more.
- warumdarum 2mo agoWeimar moment of science
- sph 2mo agoMore like the death of the public Internet before our eyes
- WhyIsItAlwaysHN 2mo agoAwesome work, the detector seems accurate on a bunch of texts I tried, surprisingly even human/ai hybrid texts
- dopamine_daddy 2mo agoThat's honestly so good to hear, thank you.
- ianm218 2mo agoSomething I've thought about a lot is that there is having someone with some domain knowledge or reason to care a lot about a particular issue spend a bunch of tokens and cycles on it until something useful comes out the other end. The most obvious ones are the math problems that have been coming out and help push the frontier of various areas of math. Another example is taking all of the public NYC open data ecosystem and crunching it to get some value which I have spent a lot of time and tokens on but not found a great medium to share. The question is just how to organize these outputs and conclusions in a way that is consistently reproducible and also how to correct errors or remove LLM nonsense where it refuses to take a position on something. Before it made sense to do this in papers but it feels like we need something like a paper format.. that is fully reproducible ideally and optimized for aggregating knowledge in a better way. I.e. before a person spent months on one of these and there was just more filtering, and the output itself was a clear signal of time spent and effort that no longer exists.
- dopamine_daddy 2mo agoYes, I’ve thought about this too. The strength of these models is that there is a lot more knowledge encoded in them than the average scientist has in mind at any given time. That means they can explore many more possible combinations of concepts. If we imagine a set of all human ideas that these models have access to, then the set of possible discoveries would be something like the superset of all possible combinations of those ideas. I think all LLM discoveries are bounded by that space. Looking at the recent OpenAI math discoveries, that seems to be pretty much what happened. Existing ideas were used as building blocks, the model found a valuable combination, and the result was something new that had real value.
- ianm218 2mo agoYeah and I think even what might be the most useful is extracting how Codex got to a particular solution into a skill or into even more customized software so it can be applied in lots more places. I know people are working on these things I just haven’t seen the right way yet. Like in manufacturing right now people are trying to encode what skilled machinists do into software and scale it up, we need to go further on that for Math/ data analysis etc.
- cute_boi 2mo agoNo wonder reading all these paper is tiring these days. I don't want to read slop generated by AI. AI written articles are generally low effort.
- deleted 2mo ago[deleted]
- Kuinox 2mo agoI just generated some docs for a lib I'm writing and it says: > 9 % machine likely human-written
- cat-whisperer 2mo agothe funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. it makes sense because it's part of their training data. but this post makes me wonder, if more papers' are written with AI, or the shape of knowledge of converging?
- n_e 2mo ago> the funniest part of these AI detectors is that if I were to upload any of einstin's paper's they will all be flagged as AI-written. Have you tried doing that or even read the article? The article says that their detector flags 0.4% of pre-AI papers as AI-written. If I paste the first page from this paper (https://www.fourmilab.ch/etexts/einstein/specrel/specrel.pdf https://www.fourmilab.ch/etexts/einstein/specrel/specrel.pdf) in https://unslop.run/app https://unslop.run/app, I get a 0% chance that it was AI-written.
- consp 2mo agoEven the translated version of the field equations paper has a 1% match only. Which I expected to be a tiny bit higher but still near 0.
- pkage 2mo agoAs with all text-only AI detection schemes, I am concerned about the accuracy of the detection. I'm skeptical of the methodology, specifically the final join of the three detector scores---how can you be sure that that final step does not introduce any biases? There's no source available, so it's difficult to tell exactly how this works or reproduce the research. I've also uploaded text samples from my own (unreleased) research from pre-LLM era, and it's seemingly scoring pretty high on the LLM-detection scores. On other papers, nearly every sentence is highlighted as red "machine-leaning," but that does not impact the score? Additionally, there are dramatic differences between the scores for identical text with and without LaTeX formatting, despite the fact that it should not matter. The takeaway from this should be "it is difficult to detect generated text and we should be careful about accepting results simply because they confirm a hypothesis." -- Relatedly, the text above scores as highly machine-written, despite the fact that I just wrote it with my human hands, I promise :)
- make3 2mo agoThe most concerning is the use of these tools to detect cheaters in schools.
- lingeringpine 2mo agoI am not a native English speaker. This is not surprising to me. I think most of the papers we write would be flagged by AI detectors. It is not because we ask LLM to write us a paper about X. It is because we are bad at writing in a scientific style, and american editors expect us to do it. With LLMs, we can write in basic sentences and tell the LLM the idea and it converts that to nice writing. If you write each paragraph and have an LLM make that paragraph more scientific, it is entirely your paper, but it is flagged as LLM generated. If you have an LLM write the paper but speak good enough English, you can make it look human even though it is not human.
- bjourne 2mo agoThere is a fine line between using an LLM to clean up grammar and spelling and using it for text generation.
- bananaflag 2mo ago[flagged]
- JadeNB 2mo agoWhich part is a sin? Using LLMs to deal with a lack of English-language fluency? I am a scientist (actually a mathematician, if it matters), and, if that's the way to deal with the practical hegemony of English in the scientific literature, then I have no problem with it. Rather that than people with important ideas can't get them before the scientific community. As long as the authors personally check and stand behind the scientific content of their papers, what do I care how the exposition was produced, especially if the role of LLMs is properly disclosed?
- slopinthebag 2mo ago[flagged]
- never_inline 2mo agohttps://lists.isocpp.org/std-proposals/att-0486/Reply_to_Zero-overhead_deterministic_exceptions:_Throwing.pdf https://lists.isocpp.org/std-proposals/att-0486/Reply_to_Zer... Here's a paper by a non native English speaker. The lack of the formal style doesn't cause any issue. Preciseness is what matters.
- bjourne 2mo ago[flagged]
- bilsbie 2mo agoJust to play devils advocate. These kind of papers are very verbose and boilerplate. I can imagine using AI to write 90% but then the actual novel content and explaining what’s important could be handwritten. Perhaps that’s what’s happening.
- NoImmatureAdHom 2mo agoYes, this is the thing. In the scientific enterprise, the writing is mostly wasted time. Why would you write it up if machines can do it competently? Sure, some people are artists - but most aren't.
- breezybottom 2mo agoIf 90% of it is boilerplate, then you really need to question if you have something worth publishing as an academic paper.
- IncreasePosts 2mo agoIt's generally always better for an academic's career to publish something, rather than publish nothing.
- pbui 2mo agoI'm not sure what to think... I uploaded a PyHPC workshop paper I wrote in 2011 and it said 27% machine. I also uploaded my PhD dissertation from 2012 and got back 40% machine, which is just barely below the 42% threshold. I don't publish anymore... but does this mean I wrote like a LLM or did LLMs learn from me? :p Update: I also uploaded a IEEE CLUSTERS paper I wrote in 2015 and it came back 74% machine written :|
- cansofgrease 2mo agoIt's been trained on a work and then distilled, the false positive noise has to be absurd.
- LearnYouALisp 2mo ago"You made this? "... "I made this."
- NitpickLawyer 2mo agoWhen "detectors" first started popping up all over the place, all of them rated the declaration of independence as 100% AI written, so... Yeah, these things just don't work. And what's even more dangerous is that people that don't understand how any of it works use these tools, and accuse people of using AI, sometimes with grave consequences. Students have been through this, at all levels of education.
- yorwba 2mo agoThey use a "threshold calibrated so pre-ChatGPT papers flag at 0.4%", so these things do work most of the time. It also means that there are known false positives, so for any given paper, scoring above the threshold isn't irrefutable proof of AI usage. But for things like estimating the overall proportion of AI writing, you only need to be correct on average, so individual false positives don't matter much.
- jasonfarnon 2mo ago
- tstactplsignore 2mo agoOne challenge with this approach is: could it be possible that the detector is simply learning to recognize words and jargon used more in the literature post 2022 as 'AI'? For example, LLMs love to talk about LLMs (and the people who write with LLMs love to write about LLMs). Could "large language model" itself therefore be flagged as an AI-like phrase by this approach? It didn't exist much in the literature before 2022, does now, and certainly does more in AI-generated text: but, it is not actually a great way to distinguish modern AI generated text from human written text. A helpful control would be to show that on some cohort of papers that can be declared reasonably clean of LLM generated text post 2023 there are very low rates compared to the arxiv. For example, while papers in the journals Nature and Science are unlikely to be entirely LLM free at this point, if those were tested through 2026, we should see a line significantly lower than the arxiv's growth.
- dopamine_daddy 2mo agoYou raise a valid point. Just by intuition I'd say if this were true, it would probably just be a small fraction of the actual flagged articles. I will still look into how I can mitigate this when I update the detector. The difficulty with this is then: How do you get a clean post 2023 dataset? I have no straightforward idea for this. You can't use other AI detectors to build it because then you'd never outperform them.
- guywithahat 2mo agoPeople are saying this is a bad thing but is it really a problem? The compelling aspect of research is the data and/or description of work, not the writing. Papers probably should be written by AI so that they're clear and well presented, while the researchers should focus on generating good data. If there is no data or work behind the paper, we should question whether the research group needs funding.
- epq22 2mo agobut the assumption you are making is that the underlying ideas are compelling. In practice, people often decide to publish incremental and/or mediocre work for various reasons, and then dress these ideas up to seem as compelling as possible to get past peer review or make a press release. at the least, this is problematic for peer-review because the absolute number of submissions outpaces the time availability of a finite number of expert reviewers. We cannot quickly generate expert human reviewers, and so the community might converge towards half-baked solutions (AI-generated reviews or rejection systems, vastly expanded referee pools, etc.) that tend to erode trust and and make scientific communities more adversarial.
- JamesBarney 2mo agoHonestly academic writing is the only place where I think AI slop might be an improvement over the status quo writing style.
- lqr 2mo agoAs an academic who is forced to review AI-generated conference/journal submissions, I can assure you this is not the case. LLM-heavy papers are terrible. Even if they manage to avoid straight-up factual incorrectness, the writing is full of ambiguities and vagueness when you look closely. Meanwhile, space is wasted repeating the same claim in multiple ways, or explaining something simple. They also seem unable to resist the hype/advertising tone, overselling the contribution while exaggerating the limitations of related work. I would much rather read grammatically incorrect or awkward sentences. While academic writing does have a few pointless historical conventions, the huge majority of "status quo writing style" is a logical consequence of 1) minimizing ambiguity, 2) organizing ideas coherently, 3) distinguishing opinion/interpretation from fact, and 4) providing enough detail to be reproducible.
- arjunvrofficial 2mo agoGood contents should always thrive, with or without AI
- arjunvrofficial 2mo agoGood contents will always thrive, with or without AI
- linolevan 2mo agoI’m very skeptical of these results. I run a pretty large group paper website on top of arXiv and we run pangram on papers. The numbers are not nearly this high. One thing I see a lot is papers flagged as AI because they include llm rollouts in the paper as examples.
- hereme888 2mo agoI run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think). The article doesn't seem to mention consideration of AI for polishing human work.
- lelanthran 2mo ago> I run most professional statements and articles through LLMs before submission. This helps correct grammar and improves accessibility through better syntax (because I write exactly how I think). > The article doesn't seem to mention consideration of AI for polishing human work. Because it isn't a consideration. You are what they are looking for.
- SoftTalker 2mo agoAI doesn't polish human work. It is more like an extruder that squeezes anything you put into it into a generic, formulaic shape, indistinguishable from writing that was lazily prompted because the writer couldn't be bothered to put forth any more effort than that.
- hereme888 2mo ago"AI does not polish human work. It acts more like an extruder, forcing anything fed into it into the same generic, formulaic shape—indistinguishable from writing produced by a lazy prompt from someone unwilling to put in any more effort." There. AI-polished sentence.
- speedstyle 2mo agoYes, this is noticeably worse than the original. The first sentence is rhythmically stilted and carries less impact. "is" to "acts" is not stylistic, we're no longer discussing the nature of the object but its behaviour. 'squeezing into a generic shape' is a transformation, each piece of writing stamped into conformity, whereas 'forcing into the same generic shape' must be pulled from a grammatical reading where all writing is being packed into one box. I think "writing produced by a lazy prompt from someone" is a wash, depends if they want to further detach the writer's agency from the result (and it still isn't produced by the prompt); I do prefer "unwilling" for the original tone; but it's certainly blander writing overall. If I came across it I'd even discount some of the metaphors and choices that are still present – since it's clearly written by AI, or a marketing copywriter, I'd guess they were "polish" rather than attempting to convey any particular connotation or nuance.
- ryandvm 2mo agoThere some real game theory mechanics at play when it comes to LLM usage in corporations. Devs are cranking out superficially superior code and documentation by just aiming the Claude Code fire hose at everything they can. Leadership encourages this because from what they can tell, there is no downside. It's hard to say if this code is structurally better or worse than before, but it's certainly voluminous and as far as leadership could ever tell, with their flawed metrics, that is all that matters. It will be years before we figure out if this is a good idea and worth the cognitive atrophy. Anyone not using LLMs all day is just not going to be as prolific. I can't imagine that the same factors aren't at play in the scientific research community where it's all about how much you can publish.
- MetaWhirledPeas 2mo ago> Leadership encourages this because from what they can tell, there is no downside. This was my big fear before we saw price increases. Now I'm pinning all my hopes on AI being too expensive to justify further big corporate pushes. (Sigh.) I love having new tools, but I hate being pushed to use ______ tool to meet some managerial metric.
- jknoepfler 2mo ago> "superficially superior code" Can you unpack that a bit? It produces measurably, meaningfully inferior code everywhere I see it in use. > "leadership encourages this because from what they can tell, there is no downside" As said 'leadership' I find this a bit puzzling. I'm seeing strong, quantifiable evidence of increasing churn, increasing incident count, and length of downtime from the date of our biggest push into GenAI, and I'm organizing efforts on my teams to mitigate those issues and actively reduce GenAI adoption. If you mean my c-suite, you're mostly correct although they are already rumbling about seeing zero or negative ROI on GenAI investments. > Anyone not using LLMs all day is just not going to be as prolific Agreed, but prolific != productive.
- OleksandrC 2mo agoThe code written by agents generally seems to reflect what you (the human driver) asked for. If you discuss the approach and architecture first, asking the right questions in the process, and then let it implement - the result is quite on point with the frontier models. Might need some minor touch-ups if agent missed some common conventions or guessed the expectations wrong (but again, salvageable with follow-up prompts). And when it comes to line-by-line logic within functions, I would argue that today's models are LESS likely to make a mistake in there than humans - especially if you cross-review with another model (e.g. "write with Claude, review with GPT"). Humans write slop too, you know. Just saying.
- IshKebab 2mo ago> and an honest account of the limitations. Gotta be trolling :-D
- themeiguoren 2mo agoI ran two papers and three blog posts of mine through here, and all but one (correctly) flagged as 0-1% machine written. The other (human written) blog post was 20%. Pretty good afaict!
- epq22 2mo agoTo corroborate this - I found pretty similar rates of AI-flagged papers over time from running pangram on ArXiv (abstracts) in a physics subfield, which is currently at about 25% as of April (write up here: https://peterse.github.io/2026/06/15/The-rising-tide-pt1.html https://peterse.github.io/2026/06/15/The-rising-tide-pt1.htm...). Its not clear from your writeup what threshold needs to be reached to be classified as "machine written". A preprint where half the text is human and half is 100% AI should be a different category than a preprint where 100% of the text is AI-assisted. Also its cool that you're making the detector available. When you say "cheap to run", do you know how this compares to pricing for a commercial detector pangram or GPTZero?
- alexpotato 2mo agoI work with several researchers (primarily working on consensus protocols etc). We were discussing research in general and I asked them: "Do you prefer the writing of the papers or the research?" They, almost unanimously, agreed that they preferred the research. This makes sense as if they preferred writing they probably would have chosen another profession. I say this b/c having LLMs available to turn research diagrams, code etc into a paper (or at least the starting point of a paper) will probably lead to MORE quality research papers. This is b/c I'm sure there was some friction in a researcher's mind of "I would love to do the research on this but don't want the trouble of writing the paper". Put another way: on a 2D plot with one axis being the skills as a researcher and the other being hatred of writing, LLMs may "unlock" the people high on both axes to get more papers out. Post Script: I agree that this could also lead to more BAD papers but the net may turn out to be positive in the long run.
- ergl 2mo agoOn my knees begging that the developers of AI detectors use them on their own writing and that at least pretended to mask the obvious LLM cliches in their blog posts.
- throwaway0123_5 2mo agoI get 0% (accurately) on my latest paper. Not super surprised, as I intentionally avoid some LLM-isms that I used to use because I don't want reviewers to have even the slightest indication that text is LLM-generated (even if in principle I'm not opposed to polishing or even wholesale generating academic text if it can convey the original research well, especially for non-native speakers). I don't think the problem is as bad as a naive reading of this article suggests. I'm highly skeptical that anywhere near 65% of recent CS papers that I've read (mostly systems papers) are substantially AI-written. I threw some recent papers I've read into the system and they come back as 0-7%.
- luciana1u 2mo ago[flagged]
- kimonsodu 2mo ago[flagged]
- jamesriso 2mo ago[dead]
- Eextra953 2mo agoI've looked at this same problem in an academic environment and have come to conclude that there is no way to reliably detect AI writing using only text. The reason for this is that no detector can take two identical inputs and classify one as synthetic (LLM) and one as organic (Human) and this situation can easily happen at the sentence or even paragraph level. Posed as a question, if a human writes a paragraph that happens to be exactly the same as a paragraph written by an LLM how do you classify that paragraph? There are a finite number of words and a finite way of combining them within an academic setting/field which means that we can't build a perfect classifier with just text. In practice, there are obvious tells and LLM-isms but these also change with time and each model has different tells so that even if a paper is full of LLM like writing we have no way to disambiguate between organic text that appears synthetic or synthetic text that appears organic. A more interesting question, to me, is looking at a corpus of essays and analyzing how writing has changed with the introduction of LLMs. We can look at changes in vocabulary, linguistic features, style embeddings, regular embeddings, typos, errors, and references over time. When looked at in this way it is clear that academic writing has changed at the population level but what has led the change is harder to track down.
- Calavar 2mo ago> There are a finite number of words and a finite way of combining them within an academic setting/field which means that we can't build a perfect classifier with just text. Maybe for very short phrases, but otherwise I disagree. Phrasing very quickly runs into a combinatorial explosion. In the words of Noam Chomsky, "Virtually every sentence that a person utters or understands is a brand-new combination of words, appearing for the first time in the history of the universe." In my opinion, the difficulty in LLM/human text discrimination isn't that a person might coincidentally write exactly the same text as an LLM would, but rather that 1) LLMs aren't hard locked to a single phrasing (so this is a tougher problem than matching to a single static document, e.g. plagiarism detection) and 2) text has relatively low information density, so you need quite a bit of it to gather enough data to run a statistical test with a reasonably narrow confidence interval.
- stereolambda 2mo agoWhen some closed models are "retired", interestingly we might not even be able to extract all their tells. We'd be limited to people writing about them on forums (where they don't often specify version, and there's a lot of memes and confirmation bias) and I don't know, doing some statistical Bayesian guessing that given the likelihood it's FooGPT 10.5, maybe we should update our beliefs about its characteristics etc.
- leawi 2mo agoIs it just me, or does this article itself read as AI generated?.. Ironically, I tried running the text of this article through their own classifier, and pretty much all of it was highlighted in red. I'm not sure what to make of this. Is this some kind of intentional irony?
- brokenodo 2mo agoIt obviously is! Human pattern-matching abilities are rather amazing, and unfortunately I've seen so much AI text that it instantly flagged in my brain within 3 seconds of opening the page. Pangram agrees: https://www.pangram.com/history/3de33376-94e3-404d-bbb0-751a0891a62b https://www.pangram.com/history/3de33376-94e3-404d-bbb0-751a...
- edot 2mo agoYes, 100%. "where the measurement breaks" and "honest account of the limitations" are all I needed to see.
- wxw 2mo agoNot convinced that this slop measurement is useful. I asked Codex to generate an article with a high score and then asked Codex to (reverse?) hill climb that score. The original generation scored 97% and then the optimized one scored 1%. Both are pretty bad and read like slop. https://gist.github.com/wbew/8a2bd6686bf875210f2244ac8ea65bf5 https://gist.github.com/wbew/8a2bd6686bf875210f2244ac8ea65bf...
- digitalPhonix 2mo ago> Here is the method, the results, and an honest account of the limitations. Pot meet kettle?
- replatformradar 2mo ago[dead]
- nullc 2mo agoWhat's the point of looking when no one even cares? consider the fraudster that went around suing people on the basis of his absurd claims of being bitcoin's creator. He's now transitioned to using AI to gather graduate degrees and is obtaining masters and doctoral degrees at a regular place and writing multiple 'papers' per day that are all quite obviously AI slop. People report his cheating and publications and simply no one cares... and this is someone court adjudicated to have fabricated evidence in court on a massive scale, including through the use of AI. But when it comes to the degrees and publication everyone involved that wanted paid got paid, and apparently that's all that matters.
- rpm91 2mo agoI tried the detector on several of my own Stack Overflow answers, and the fourth or fifth one I tried was flagged as "85 % machine / likely machine-written" with the default settings, well above the default sensitivity threshold of 42%. When I turn off the option to strip LaTeX commands (the source is Markdown), that jumps to 96% machine. Admittedly, a Stack Overflow answer is somewhat outside the realm of scientific writing, so it's still possible that the detector may be accurate within that domain. That said, it's a cautionary tale on the hazards of applying classifiers like this outside of the domain that they were trained on.
- efitz 2mo ago> Here is the method, the results, and an honest account of the limitations. This paper was written using AI, to be honest
- localhoster 2mo agoI once read a comment on hn that says that ai detection in writing is so hard because text doesn't hold enough information. Now I'm no researcher,but that intuitivly makes sense to me. I mean, apart from the obvious cases, how can you detect ai written text?
- edifierxuhao 2mo ago[flagged]
- Caixu 2mo ago[flagged]
- ilamont 2mo ago> We took papers submitted in 2021 and 2022, before ChatGPT, treated them as ground-truth human How is the author so sure that people or companies weren’t already using early tools at that time?
- rekpero 2mo ago[flagged]
- wulfkaal 2mo ago[flagged]