9 ms·
I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine) The paper was https://openreview.net/forum?id=0ZnXGzLcOg https://ope
by j2kun 9mo ago
I spot-checked one of the flagged papers (from Google, co-authored by a colleague of mine)
The paper was https://openreview.net/forum?id=0ZnXGzLcOg https://openreview.net/forum?id=0ZnXGzLcOg and the problem flagged was "Two authors are omitted and one (Kyle Richardson) is added. This paper was published at ICLR 2024." I.e., for one cited paper, the author list was off and the venue was wrong. And this citation was mentioned in the background section of the paper, and not fundamental to the validity of the paper. So the citation was not fabricated, but it was incorrectly attributed (perhaps via use of an AI autocomplete).
I think there are some egregious papers in their dataset, and this error does make me pause to wonder how much of the rest of the paper used AI assistance. That said, the "single error" papers in the dataset seem similar to the one I checked: relatively harmless and minor errors (which would be immediately caught by a DOI checker), and so I have to assume some of these were included in the dataset mainly to amplify the author's product pitch. It succeeded.
- davidguetta 9mo agoYeah even the entire "Jane Doe / Jame Smith" my first thought is that it could have been a latex default value There was dumb stuff like this before the GPT era, it's far from convincing
- ls612 9mo agoThere are people who just want to punish academics for the sake of punishing academics. Look at all the people downthread salivating over blacklisting or even criminally charging people who make errors like this with felony fraud. Its the perfect brew of anti AI and anti academia sentiment. Also, in my field (economics), by far the biggest source of finding old papers invalid (or less valid, most papers state multiple results) is good old fashioned coding bugs. I'd like to see the software engineers on this site say with a straight face that writing bugs should lead to jail time.
- worik 9mo ago> I'd like to see the software engineers on this site say with a straight face that writing bugs should lead to jail time. My hand is up. I do not believe in gaol, but I do agree with the sentiment.
- ls612 9mo agoLet he who is without sin cast the first stone…
- girvo 9mo agoIf there were real consequences, we wouldn't be forced to churn out buggy nonsense by our employers. So we'd be able to take the time to do the right thing. Bug free software is possible, the world just says its not worth it today.
- ls612 9mo ago>Bug free software is possible, ... Mr. Turing and his halting problem would like to politely disagree with this assertion.
- ted_dunning 9mo agoYou misread the comment and DR Turing's paper. Getting all possible software correct is impossible, clearly. Getting all the software you release is more possible because you can choose not to release the software that it is too hard to prove correct. Not that the suggestion is practical or likely, but your assertion that it is impossible is incorrect.
- ls612 9mo agoIf you want to be pedantic I’m pretty sure every single general purpose OS (and thus also the programs running under it) falls into the category of not provably correct so it’s a distinction without a difference in real life.
- miki123211 9mo agoAnd research codebases (in AI and otherwise) are usually of extremely bad quality. It's usually a bunch of extremely poorly-written scripts, with no indication which order to run them in, how inputs and outputs should flow between them, and which specific files the scripts were run on to calculate the statistics presented in the paper.
- davidguetta 9mo agoCodebase can bé of high quality but still you have no idea how they got the paper result
- nativeit 9mo ago> Between 2020 and 2025, submissions to NeurIPS increased more than 220% from 9,467 to 21,575. In response, organizers have had to recruit ever greater numbers of reviewers, resulting in issues of oversight, expertise alignment, negligence, and even fraud. I don’t think the point being made is “errors didn’t happen pre-GPT”, rather the tasks of detecting errors have become increasingly difficult because of the associated effects of GPT.
- ctoth 9mo ago> rather the tasks of detecting errors have become increasingly difficult because of the associated effects of GPT. Did the increase to submissions to NeurIPS from 2020 to 2025 happen because ChatGPT came out in November of 2022? Or was AI getting hotter and hotter during this period, thereby naturally increasing submissions to ... an AI conference?
- amitav1 9mo agoI guess the way one would verify that this is more general trend in academia would be to run this on accepted papers to a non-AI conference?
- mturmon 9mo agoI was an area chair on the NeurIPS program committee in 1997. I just looked and it seems that we had 1280 submissions. At that time, we were ultimately capped by the book size that MIT Press was willing to put out - 150 8-page articles. Back in 1997 we were all pretty sure we were on to something big. I'm sure people made mistakes on their bibliographies at that time as well! And did we all really dig up and read Metropolis, Rosenbluth, Rosenbluth, Teller, and Teller (1953)? Edited to add: Someone made a chart! Here: https://papercopilot.com/statistics/neurips-statistics/ https://papercopilot.com/statistics/neurips-statistics/ You can see the big bump after the book-length restriction was lifted, and the exponential rise starting ~2016.
- dekhn 9mo agoI cited Watson and Crick '53 in my PhD thesis and I did go dig it up and read it. I had to go to the basement of the library, use some sort of weird rotating knob to move a heavy stack of journals over, find some large bound book of the year's journals, and navigate to the paper. When I got the page, it had been cut out by somebody previous and replaced with a photocopied verison. (I also invested a HUGE amount of my time into my bibliography in every paper I've written as first author, curating a database and writing scripts to format in the various journal formats. This involved multiple independent checks from several sources, repeated several times.
- bjourne 9mo agoStill a citation to a work you clearly have not read...
- nativeit 9mo agoI see your point, but I don’t see where the author makes any claims about the specifics of the hallucinations, or their impact on the papers’ broader validity. Indeed, I would have found the removal of supposed “innocuous” examples to be far more deceptive than simply calling a spade a spade, and allowing the data to speak for itself.
- gowld 9mo agoThe point is that they should focus on the meaningful errors, not the automiation of meaningless errors.
- reliabilityguy 9mo agoWhy these are meaningless? How do I know now that the whole paper is not a slop?
- j2kun 9mo agoThe author calls the mistakes "confirmed hallucinations" without proof (just more or less evidence). The data never "speak for itself." The author curates the data and crafts a story about it. This story presented here is very suggestive (even using the term "hallucination" is suggestive). But calling it "100 suspected hallucinations", or "25 very likely hallucinations" does less for the author's end goal: selling their service.
- gold23 9mo agoObviously a post on a startup's blog will be more editorialized than an academic paper. Still, this seems like an important discussion to have.
- m-schuetz 9mo agoBibtex are often also incorrectly generated. E.g., google scholar sometimes puts the names of the editors instead of the authors into the bibtex entry.
- worik 9mo ago> Bibtex are often also incorrectly generated ...and including the erroneous entry is squarely the author's fault. Papers should be carefully crafted, not churned out. I guess that makes me sweetly naive
- daveFNbuck 9mo agoYou want the content of the paper to be carefully crafted. Bibtex entries are the sort of thing you want people to copy and paste from a trusted source, as they can be difficult to do consistently correctly.
- tuckerman 9mo agoI don't think the original comment was saying this isn't a problem but that flagging it as a hallucination from an LLM is a much more serious allegation. In this case, it also seems like it was done to market a paid product which makes the collateral damage less tolerable in my opinion. > Papers should be carefully crafted, not churned out. I think you can say the same thing for code and yet, even with code review, bugs slip by. People aren't perfect and problems happen. Trying to prevent 100% of problems is usually a bad cost/benefit trade-off.
- m-schuetz 9mo agoThat's not happening for a similar reason people do not bug-check every single line of every single third-party library in their code. It's a chore that costs valuable time that you can instead spend on getting the actual stuff done. What's really important is that the scientific contribution is 100% correct and solid. For the references, the "good enough" paradigm applies. They mustn't be complete bogus, like the referenced work not existing at all which would indicate that the authors didnt even look at the reference. But minor issues like typos or rare issues with wrong authors can happen.
- i_am_proteus 9mo ago>this error does make me pause to wonder how much of the rest of the paper used AI assistance And this is what's operative here. The error spotted, the entire class of error spotted, is easily checked/verified by a non-domain expert. These are the errors we can confirm readily, with obvious and unmistakable signature of hallucination. If these are the only errors, we are not troubled. However: we do not know if these are the only errors, they are merely a signature that the paper was submitted without being thoroughly checked for hallucinations. They are a signature that some LLM was used to generate parts of the paper and the responsible authors used this LLM without care. Checking the rest of the paper requires domain expertise, perhaps requires an attempt at reproducing the authors' results. That the rest of the paper is now in doubt, and that this problem is so widespread, threatens the validity of the fundamental activity these papers represent: research.
- ls612 9mo agoGoogle scholar and the vagaries of copy/paste errors has mangled bibitex ever since it became a thing, a single citation with these sorts of errors may not even be AI, just “normal” mistakes.
- fn-mote 9mo agoThis seems like finding spelling errors and using them to cast the entire paper into doubt. I am unconvinced that the particular error mentioned above is a hallucination, and even less convinced that it is a sign of some kind of rampant use of AI. I hope to find better examples later in the comment section.
- j2kun 9mo agoI actually believe it was an AI hallucination, but I agree with you that it seems the problem is far more concentrated to a few select papers (e.g., one paper made up more than 10% of the detected errors).
- gold23 9mo agoWhy don't you look at the actual article? There are several more egregious examples, e.g., the authors being cited as "John Smith and Jane Doe"
- nearbuy 9mo agoThe rate here (about 1% of papers) just doesn't seem that bad, especially if many of the errors are minor and don't affect the validity of the results. In other fields, over half of high-impact studies don't replicate.
- currymj 9mo agothe earlier list of ICLR papers had way more egregious examples. Those were taken from the list of submissions not accepted papers however.
- nazgul17 9mo agoThe thing is, when you copy paste a bibliography entry from the publisher or from Google Scholar, the authors won't be wrong. In this case, it is. If I were to write a paper with AI, I would at least manage the bibliography by hand, conscious of hallucinations. The fact that the hallucination is in the bibliography is a pretty strong indicator that the paper was written entirely with AI.
- arjvik 9mo agoI'm not sure I agree... while I don't ever see myself writing papers with AI, I hate wrangling a bibtex bibliography. I wouldn't trust today's GPT-5-with-web-search to do turn a bullet point list of papers into proper citations without checking myself, but maybe I will trust GPT-X-plus-agent to do this.
- joshvm 9mo agoReference managers have existed for decades now and they work deterministically. I paid for one when writing my doctoral thesis because it would have been horrific to do by hand. Any of the major tools like Zotero or Mendeley (I used Papers) will export a bibtex file for you, and they will accept a RIS or similar format that most journals export.
- storystarling 9mo agoThis seems solvable today if you treat it as an architecture problem rather than relying on the model's weights. I'm using LangGraph to force function calls to Crossref or OpenAlex for a similar workflow. As long as you keep the flow rigid and only use the LLM for orchestration and formatting, the hallucinations pretty much disappear.
- jmmcd 9mo agoGoogle Scholar provides imperfect citations - very often wrong article type (eg article versus conference paper), but up to and including missing authors, in my experience.
- 9mo ago
- David_Osipov 9mo agoGreat job! I've tried to test their tool as well, but was totally paywalled.
- fmbb 9mo ago> So the citation was not fabricated, but it was incorrectly attributed (perhaps via use of an AI autocomplete). Well the title says ”hallucinations”, not ”fabrications”. What you describe sounds exactly like what AI builders call hallucinations.
- j2kun 9mo agoRead the article. The author uses the word "fabricate" repeatedly to describe the situation where the wrong authors are in the citation.
- _alternator_ 9mo agoThe missing analysis is, of course, a comparison with pre-LLM conferences, like 2022 or 2023 that would show a “false positive” rate for the tool.
- anishrverma 9mo agoAgreed. What I find more interesting is how easy these errors are to introduce and how unlikely they are to be caught. As you point out, a DOI checker would immediately flag this. But citation verification isn’t a first-class part of the submission or review workflow today. We’re still treating citations as narrative text rather than verifiable objects. That implicit trust model worked when volumes were lower, but it doesn’t seem to scale anymore There’s a project I’m working on at Duke University, where we are building a system that tries to address exactly this gap by making references and review labor explicit and machine verifiable at the infrastructure level. There’s a short explainer here that lays out what we mean, if useful context helps: https://liberata.info/ https://liberata.info/
- dexdal 9mo agoCitation checks are a workflow problem, not a model problem. Treat every reference as a dependency that must resolve and be reproducible. If the checker cannot fetch and validate it, it does not ship.
- janalsncm 9mo agoThis is par for the course for GPTZero, which also falsely claims they can detect AI generated text, a fundamentally impossible task to do accurately.
- ainch 9mo agoI'm not going to bat for GPTZero, but I think it's clearly possible to identify some AI-written prose. Scroll through LinkedIn or Twitter replies and there are clear giveaways in tone, phrasing and repeated structures (it's not just X it's Y). Not to say that you could ever feasibly detect all AI-generated text, but if it's possible for people to develop a sense for the tropes of LLM content then there's no reason you couldn't detect it algorithmically.
- janalsncm 9mo ago> there's no reason you couldn't detect it algorithmically For any real world classifier there is a precision/recall tradeoff. Do you care more about false positives or false negatives? If you choose to truly minimize false positives you should simply always predict negative. For your example “it’s not just X it’s Y” I agree it’s a red flag. But the origin of the pattern is from human text which the LLM picked up on. So some people did (and likely still do) use that construction.
- NedF 9mo ago[dead]
- bjourne 9mo agoSorry, but blaming it on "AI autocomplete" is the dumbest excuse ever. Author lists come from BibTeX entries and while they often contains errors since they can come from many sources, they do not contain completely made up authors. I don't share your view that hallucinated citations are less damaging in background section. Background, related works, and introduction is the sections where citations most often show up. These sections are meant to be read and generating them with AI is plain cheating.
- j2kun 9mo agoI'm not blaming anything on anything, because I did not (nor did the authors) confirm the cause of any of these errors. > I don't share your view that hallucinated citations are less damaging in background section. Who exactly is damaged in this particular instance?
- bjourne 9mo agoTrust is damaged. I cannot verify that the evidence is correct only that the conclusions follow from the evidence. I have to rely on the authors to truthfully present their evidence. If they for whatever reason add hallucinated citations to their background that trust is 100% gone.
- j2kun 9mo agoYou are speaking in the abstract. Did you read this paper? I suspect you did not.
- StopDisinfo910 9mo agoThe example you provided doesn't sit right me. If the mistake is one error of author and location in a citation, I find it fairly disingenuous to call that an hallucination. At least, it doesn't meet the threshold for me. I have seen this kind of mistakes done long before LLM were even a thing. We used to call them that: mistakes.
- beowulfey 9mo agoAs with anything, it is about trusting your tools. Who is culpable for such errors? In the days of human authors, the person writing the text is responsible for not making these errors. When AI does the writing, the person whose name is on the paper should still be responsible—but do they know that? Do they realize the responsibility they are shouldering when they use these AI tools? I think many times they do not; we implicitly trust the outputs of these tools, and the dangers of that are not made clear.
- lou1306 9mo ago> relatively harmless and minor errors They are not harmless. These hallucinated references are ingested by Google Scholar, Scopus, etc., and with enough time they will poison those wells. It is also plain academic malpractice, no matter how "minor" the reference is.