16 ms·
Over fifty new hallucinations in ICLR 2026 submissions
- jqpabc123 10mo agoThe legal system has a word to describe AI "slop" --- it is called "negligence". And as the remedy starts being applied (aka "liability"), the enthusiasm for AI will start to wane. I wouldn't be surprised if some businesses ban the use of AI --- starting with law firms.
- loloquwowndueo 10mo agoI applaud your use of triple dashes to avoid automatic conversion to em dashes and being labeled an AI. Kudos!
- ghaff 10mo agoThis is a particular meme that I really don't like. I've used em-dashes routinely for years. Do I need to stop using them because various people assume they're an AI flag?
- TimedToasts 10mo agoNo, but you should be prepared to have people suspect you are using AI to create your responses. C'est la vie. The good news is that it will rectify itself and soon the output will lack even these signals.
- ghaff 10mo agoWell, I work for myself and people can either judge my work on its own merits or not. Don't care all that much.
- ls612 10mo agoThe legal system has a word to describe software bugs --- it is called "negligence". And as the remedy starts being applied (aka "liability"), the enthusiasm for software will start to wane. What if anything do you think is wrong with my analogy? I doubt most people here support strict liability for bugs in code.
- hnfong 10mo agoI don't even think GP knows what negligence is. Generally the law allows people to make mistakes, as long as a reasonable level of care is taken to avoid them (and also you can get away with carelessness if you don't owe any duty of care to the party). The law regarding what level of care is needed to verify genAI output is probably not very well defined, but it definitely isn't going to be strict liability. The emotionally-driven hate for AI, in a tech-centric forum even, to the extent that so many commenters seem to be off-balance in their rational thinking, is kinda wild to me.
- ls612 10mo agoI don’t get it, tech people clearly have the most to gain from AI like Claude Code.
- jqpabc123 10mo agoComputer code is highly deterministic. This allows it to be tested fairly easily. Unfortunately, code productionn is not the only use-case for AI. Most things in life are not as well defined --- a matter of judgment. AI is being applied in lots of real world cases where judgment is required to interpret results. For example, "Does this patient have cancer". And it is fairly easy to show that AI's judgment can be highly suspect. There are often legal implications for poor judgment --- i.e. medical malpractice. Maybe you can argue that this is a mis-application of AI --- and I don't necessarily disagree --- but the point is, once the legal system makes this abundantly clear, the practical business case for AI is going to be severely reduced if humans still have to vet the results in every case.
- hnfong 10mo agoWhy do you think AI is inherently worse than humans in judging whether a patient has cancer, assuming they are given the same information as the human doctor? Is there some fundamental assumption that makes AI worse, or are you simply projecting your personal belief (trust) in human doctors? (Note that given the speed of progress of AI and that we're talking about what the law ought to be, not what it was in the past, the past performance of AI on cancer cases do not have much relevance unless a fundamental issue with AI is identified) Note that whether a person has cancer is generally well-defined, although it may not be obvious at first. If you just let the patient go untreated, you'll know the answer quite definitely in a couple years.
- mjd 10mo agoI love that fake citation that adds George Costanza to the list of authors!
- watwut 10mo agoCan we just call them "lies" and "fabrications" which is what they are? If I write the same, you will call them "made up citations" and "academic dishonesty". One can use AI to help them write without going all the way to having it generate facts and citations.
- sorokod 10mo agoAs long as the submissions are on behalf of humans we should. The humans should accept the consequences too.
- cgfjtynzdrfht 10mo ago[dead]
- jmount 10mo agoThat is a key point: they are fabrications, not hallucinations.
- Barbing 10mo agoArs has often gone with “confabulation”: >Confabulation was coined right here on Ars, by AI-beat columnist Benj Edwards, in Why ChatGPT and Bing Chat are so good at making things up (Apr 2023). https://arstechnica.com/civis/threads/researchers-describe-how-to-tell-if-chatgpt-is-confabulating.1501388/post-42928247 https://arstechnica.com/civis/threads/researchers-describe-h... >Generative AI is so new that we need metaphors borrowed from existing ideas to explain these highly technical concepts to the broader public. In this vein, we feel the term "confabulation," although similarly imperfect, is a better metaphor than "hallucination." In human psychology, a "confabulation" occurs when someone's memory has a gap and the brain convincingly fills in the rest without intending to deceive others. https://arstechnica.com/information-technology/2023/04/why-ai-chatbots-are-the-ultimate-bs-machines-and-how-people-hope-to-fix-them https://arstechnica.com/information-technology/2023/04/why-a...
- deleted 10mo ago[deleted]
- jameshart 10mo agoIs the baseline assumption of this work that an erroneous citation is LLM hallucinated? Did they run the checker across a body of papers before LLMs were available and verify that there were no citations in peer reviewed papers that got authors or titles wrong?
- tokai 10mo agoYeah that is what their tool does.
- llm_nerd 10mo agoPeople will commonly hold LLMs as unusable because they make mistakes. So do people. Books have errors. Papers have errors. People have flawed knowledge, often degraded through a conceptual game of telephone. Exactly as you said, do precisely this to pre-LLM works. There will be an enormous number of errors with utter certainty. People keep imperfect notes. People are lazy. People sometimes even fabricate. None of this needed LLMs to happen.
- add-sub-mul-div 10mo agoQuoting myself from just last night because this comes up every time and doesn't always need a new write-up. > You also don't need gunpowder to kill someone with projectiles, but gunpowder changed things in important ways. All I ever see are the most specious knee-jerk defenses of AI that immediately fall apart.
- llm_nerd 10mo ago[flagged]
- the_af 10mo agoLLM are a force multiplier of this kind of errors though. It's not easy to hallucinate papers out of whole cloth, but LLMs can easily and confidently do it, quote paragraphs that don't exist, and do it tirelessly and at a pace unmatched by humans. Humans can do all of the above but it costs them more, and they do it more slowly. LLMs generate spam at a much faster rate.
- TaupeRanger 10mo agoIt's going to be even worse than 50: > Given that we've only scanned 300 out of 20,000 submissions, we estimate that we will find 100s of hallucinated papers in the coming days.
- shusaku 10mo ago20,000 submissions to a single conference? That is nuts
- analog31 10mo agoThis is an interesting article along those lines... https://www.theguardian.com/technology/2025/dec/06/ai-research-papers https://www.theguardian.com/technology/2025/dec/06/ai-resear...
- ghaff 10mo agoDoesn't seem especially out of the norm for a large conference. Call it 10,000 attendees which is large but not huge. Sure; not everyone attending puts in a session proposal. But others put multiple. And many submit but, if not accepted don't attend. Can't quote exact numbers but when I was on the conference committee for a maybe high four figures attendance conference, we certainly had many thousands of submissions.
- zipy124 10mo agoWhen academics are graded based on number of papers this is the result.
- adestefan 10mo agoThe problem isn't only papers it's that the world of academic computer science coalesced around conference submissions instead of journal submissions. This isn't new and was an issue 30 years ago when I was in grad school. It makes the work of conference organizes the little block holding up the entire system.
- 10mo ago
- shusaku 10mo agoChecking each citation one by one is quite critical in peer review, and of course checking a colleagues paper. I’ve never had to deal with AI slop, but you’ll definitely see something cited for the wrong reason. And just the other day during the final typesetting of a paper of mine I found the journal had messed up a citation (same journal / author but wrong work!)
- stefan_ 10mo agoIs it quite critical? Peer review is not checking homework, it's about the novel contribution presented. Papers will frequently cite related notable experiments or introduce a problem that as a peer reviewer in the field I'm already well familiar with. These paragraphs generate many citations but are the least important part of a peer review. (People submitting AI slop should still be ostracized of course, if you can't be bothered to read it, why would you think I should)
- shusaku 10mo agoFair point. In my mind it is critical because mistakes are common and can only be fixed by a peer. But you are right that we should not miss the forest through the trees and get lost on small details.
- tomrod 10mo agoHow sloppy is someone that they don't check their references!
- analog31 10mo agoA reference is included in a paper if the paper uses information derived from the reference, or to acknowledges the reference as a prior source. If the reference is fake, then the derived information could very well be fake. Let's say that I use a formula, and give a reference to where the formula came from, but the reference doesn't exist. Would you trust the formula? Let's say a computer program calls a subroutine with a certain name from a certain library, but the library doesn't exist. A person doing good research doesn't need to check their references. Now, they could stand to check the references for typographic errors, but that's a stretch too. Almost every online service for retrieving articles includes a reference for each article that you can just copy and paste.
- theoldgreybeard 10mo agoIf a carpenter builds a crappy shelf “because” his power tools are not calibrated correctly - that’s a crappy carpenter, not a crappy tool. If a scientist uses an LLM to write a paper with fabricated citations - that’s a crappy scientist. AI is not the problem, laziness and negligence is. There needs to be serious social consequences to this kind of thing, otherwise we are tacitly endorsing it.
- RossBencina 10mo agoNo qualified carpenter expects to use a hammer to drill a hole.
- gdulli 10mo agoThat's like saying guns aren't the problem, the desire to shoot is the problem. Okay, sure, but wanting something like a metal detector requires us to focus on the more tangible aspect that is the gun.
- baxtr 10mo agoIf I gave you a gun would you start shooting people just because you had one?
- agentultra 10mo agoIf I gave you a gun without a safety could you be the one to blame when it goes off because you weren’t careful enough? The problem with this analogy is that it makes no sense. LLMs aren’t guns. The problem with using them is that humans have to review the content for accuracy. And that gets tiresome because the whole point is that the LLM saves you time and effort doing it yourself. So naturally people will tend to stop checking and assume the output is correct, “because the LLM is so good.” Then you get false citations and bogus claims everywhere.
- sigbottle 10mo agoSorry, I'm not following the gun analogies at all But regardless, I thought the point was that... > The problem with using them is that humans have to review the content for accuracy. There are (at least) two humans in this equation. The publisher, and the reader. The publisher at least should do their due diligence, regardless of how "hard" it is (in this case, we literally just ask that you review your OWN CITATIONS that you insert into your paper). This is why we have accountability as a concept.
- YouAreWRONGtoo 10mo ago[dead]
- Isamu 10mo agoSomeone commented here that hallucination is what LLMs do, it’s the designed mode of selecting statistically relevant model data that was built on the training set and then mashing it up for an output. The outcome is something that statistically resembles a real citation. Creating a real citation is totally doable by a machine though, it is just selecting relevant text, looking up the title, authors, pages etc and putting that in canonical form. It’s just that LLMs are not currently doing the work we ask for, but instead something similar in form that may be good enough.
- deleted 10mo ago[deleted]
- make3 10mo agoThis interpretation would have been ok for old generation models without search tools enabled and without reliable tool use and reasoning. Modern LLMs can actually look up the existence of papers with web search, and with reasoning, one can definitely get reasonable results by requiring the model to double check that everything actually exists.
- gedy 10mo agoThe issue is there are incentives for more quantity and not quality in modern science (well more like academia), so people will use tools to pump stuff out. It'll get worse as academic jobs tighten due.
- dclowd9901 10mo agoTo me, this is exactly what LLMs are good for. It would be exhausting double checking for valid citations in a research paper. Fuzzy comparison and rote lookup seem primed for usage with LLMs. Writing academic papers is exactly the _wrong_ usage for LLMs. So here we have a clear cut case for their usage and a clear cut case for their avoidance.
- idiotsecant 10mo agoExactly, and there's nothing wrong with using LLMs in this same way as part of the writing process to locate sources (that you verify), do editing (that you check), etc. It's just peak stupidity and laziness to ask it to do the whole thing.
- skobes 10mo agoIf LLMs produce fake citations, why would we trust LLMs to check them?
- watwut 10mo agoBecause the risk is lower. They will give you suspicious citations and you can manually check those for false positives. If some false citation pass, it was still a net gain.
- venturecruelty 10mo agoBecause my boss said if I don't, I'm fired.
- dawnerd 10mo agoShouldn’t need an llm to check. It’s just a list of authors. I wouldn’t trust an llm on this, and even if they were perfect that’s a lot of resource use just to do something traditional code could do.
- dclowd9901 10mo agoI would assume you would use the LLM to not only check the source exists but check that the citation referenced actually says what the author says it does. That's not something you can do heuristically, I would think.
- teekert 10mo agoThanx AI, for exposing this problem that we knew was there, but could never quite prove.
- hyperpape 10mo agoIt's awful that there are these hallucinated citations, and the researchers who submitted them ought to be ashamed. I also put some of the blame on the boneheaded culture of academic citations. "Compression has been widely used in columnar databases and has had an increasing importance over time.[1][2][3][4][5][6]" Ok, literally everyone in the field already knows this. Are citations 1-6 useful? Well, hopefully one of them is an actually useful survey paper, but odds are that 4-5 of them are arbitrarily chosen papers by you or your friends. Good for a little bit of h-index bumping! So many citations are not an integral part of the paper, but instead randomly sprinkled on to give an air of authority and completeness that isn't deserved. I actually have a lot of respect for the academic world, probably more than most HN posters, but this particular practice has always struck me as silly. Outside of survey papers (which are extremely under-provided), most papers need many fewer citations than they have, for the specific claims where the paper is relying on prior work or showing an advance over it.
- mccoyb 10mo agoThat's only part of the reason that this type of content is used in academic papers. The other part is that you never know what PhD student / postdoc / researcher will be reviewing your paper, which means you are incentivized to be liberal with citations (however tangential) just in case someone is reading your paper, and has the reaction "why didn't they cite this work, of which I had some role in?" Papers with a fake air of authority of easily dispatched with. What is not so easily dispatched with is the politics of the submission process. This type of content is fundamentally about emotions (in the reviewer of your paper), and emotions is undeniably a large factor in acceptance / rejection.
- zipy124 10mo agoIndeed. One can even game review systems by leaving errors in for the reviewers to find so that they feel good about themselves and that they've done their job. The meta-science game is toxic and full of politics and ego-pleasing.
- neilv 10mo agohttps://blog.iclr.cc/2025/11/19/iclr-2026-response-to-llm-generated-papers-and-reviews/ https://blog.iclr.cc/2025/11/19/iclr-2026-response-to-llm-ge... > Papers that make extensive usage of LLMs and do not disclose this usage will be desk rejected. This sounds like they're endorsing the game of how much can we get away with, towards the goal of slipping it past the reviewers, and the only penalty is that the bad paper isn't accepted. How about "Papers suspected of fabrications, plagiarism, ghost writers, or other academic dishonesty, will be reported to academic and professional organizations, as well as the affiliated institutions and sponsors named on the paper"?
- thruifgguh585 10mo ago> crushed by an avalanche of submissions fueled by generative AI, paper mills, and publication pressure. Run of the mill ML jobs these days ask for "papers in NeurIPS ICLR or other Tier-1 conferences". We're well past Goodhart's law when it comes to publications. It was already insane in CS - now it's reached asylum levels.
- disqard 10mo agoYou said the quiet part out loud. Academia has been ripe for disruption for a while now. The "Rooter" paper came out 20 years ago: https://www.csail.mit.edu/news/how-fake-paper-generator-tricked-scientific-journals-20-years-later https://www.csail.mit.edu/news/how-fake-paper-generator-tric...
- MarkusQ 10mo agoThis is as much a failing of "peer review" as anything. Importantly, it is an intrinsic failure, which won't go away even if LLMs were to go away completely. Peer review doesn't catch errors. Acting as if it does, and thus assuming the fact of publication (and where it was published) are indicators of veracity is simply unfounded. We need to go back to the food fight system where everyone publishes whatever they want, their colleagues and other adversaries try their best to shred them, and the winners are the ones that stand up to the maelstrom. It's messy, but it forces critics to put forth their arguments rather than quietly gatekeeping, passing what they approve of, suppressing what they don't.
- ulrashida 10mo agoPeer review definitely does catch errors when performed by qualified individuals. I've personally flagged papers for major revisions or rejection as a result of errors in approach or misrepresentation of source material. I have peers who say they have done similar. I'm not sure why you think this isn't the case?
- MarkusQ 10mo agoPoor wording on my part. I should have said "Peer review doesn't catch _all_ errors" or perhaps "Peer review doesn't eliminate errors". In other words, being "peer reviewed" is nowhere close to "error free," and if (as is often the case) the rate of errors is significantly greater than the rate at which errors are caught, peer review may not even significantly improve the quality. https://pmc.ncbi.nlm.nih.gov/articles/PMC1182327/ https://pmc.ncbi.nlm.nih.gov/articles/PMC1182327/
- ulrashida 10mo agoThanks for clarifying, I fully agree with your take. Peer review helps, particularly where reviewers are equipped and provided the time to do the role correctly. However, it is not alone a guarantor of quality. As someone proximate to academia its becoming obvious that many professors are beginning to throw in the towel or are sharply reducing their time verifying quality when faced with the rising tide of slop. The window for avoiding the natural consequences of these trends feels like it is getting scarily small. Thanks for taking the time to reply!
- michaelcampbell 10mo agoAfter an interview with Cory Doctorow I saw recently, I'm going to stop anthropomorphizing these things by calling them "hallucinations". They're computers, so these incidents are just simply Errors.
- grayhatter 10mo agoI'll continue calling them hallucinations. That's a much more fitting term when you account for the reasonableness of people who believe them. There's also equally a huge breadth of different types of errors that don't pattern match well into, "made up bullshit" the same way calling them hallucinations do. There's no need to introduce that ambiguity when discussing something narrow. there's nothing wrong with anthropomorphizing genai, it's source material is human sourced, and humans are going to use human like pattern matching when interacting with it. I.e. This isn't the river I want to swim upstream in. I assume you wouldn't complain if someone anthropomorphized a rock... up until they started to believe it was actually alive.
- vegabook 10mo agoGiven that an (incompetent or even malicious) human put their names(s) to this stuff, “bullshit” is an even better and fitting anthropomorphization
- grayhatter 10mo ago> incompetent or even malicious sufficiently advance some competences indistinguishable from actual malice.... and thus should be treated the same
- skobes 10mo agoDevelopers have been anthropomorphizing computers for as long as they've been around though. "The compiler thinks my variable isn't declared" "That function wants a null-terminated string" "Teach this code to use a cache" Even the word computer once referred to a human.
- crazygringo 10mo ago
- leoc 10mo agoAh, yes: meta-level model collapse. Very good, carry on.
- Ekaros 10mo agoOne wonders why this has not been largely fully automated. If we track those citations anyway. Surely we have database of them and most of them are easily matched there. So only outliers need to be checked either as new latest papers or mistakes which should be close enough to something or real fakes. Maybe there just is no incentive for this type of activity.
- QuadmasterXLII 10mo agoIt seems like the GPT zero team is automating it! Up to very recently, no one sane would cite a paper with correct title but make up random authors- and shortly, this specific signal will be goodhearted away by a “make my malpractice less detectable MCP,” so I can see why this automation is happening exactly now.
- analog31 10mo agoFor that matter, it could be automated at the source. Let's say I'm an author. I'd gladly run a "linter" on my article that flags references that can't be tracked, and so forth. It would be no different than testing a computer program that I write before giving it to someone.
- IanCal 10mo agoWe do have these things and they are often wrong. Loads of the examples given look better than things I’ve seen in real databases on this kind of thing and I worked in this area for a decade.
- ulrashida 10mo agoUnfortunately while catching false citations is useful, in my experience that's not usually the problem affecting paper quality. Far more prevalent are authors who mis-cite materials, either drawing support from citations that don't actually say those things or strip the nuance away by using cherry picked quotes simply because that is what Google Scholar suggested as a top result. The time it takes to find these errors is orders of magnitude higher than checking if a citation exists as you need to both read and understand the source material. These bad actors should be subject to a three strikes rule: the steady corrosion of knowledge is not an accident by these individuals.
- 19f191ty 10mo agoExactly abuse of citations is a much more prevalent and sinister issue and has been for a long time. Fake citations are of course bad but only tip of the iceberg.
- seventytwo 10mo agoThen punish all of it.
- hippo22 10mo agoIt seems like this is the type of thing that LLMs would actually excel at though: find a list of citations and claims in this paper, do the cited works support the claims?
- bryanrasmussen 10mo agosure, except when they hallucinate that the cited works support the claims when they do not. At which point you're back at needing to read the cited works to see if they support the claims.
- BHSPitMonkey 10mo agoYou don't just accept the review as-is, though; You prompt it to be a skeptic and find a handful of specific examples of claims that are worth extra attention from a qualified human. Unfortunately, this probably results in lazy humans _only_ reading the automated flagged areas critically and neglecting everything else, but hey—at least it might keep a little more garbage out?
- kaluga 10mo ago[dead]
- peppersghost93 10mo agoI sincerely hope every person who has invested money in these bullshit machines loses every cent they've got to their name. LLMs poison every industry they touch.
- obscurette 10mo agoThat's what I'm really afraid of – we will be drowning in the AI slop as a society and we'll loose the most important thing that made free and democratic society possible - a trust. People just don't tust anyone and/or anything any more. And the lack of trust, especially in scale, is very expensive.
- John7878781 10mo agoYep. And trust is already at all time lows for science, as if it couldn't get any worse.
- benbojangles 10mo agoHow to get to the top of you are not smart enough?
- upofadown 10mo agoIf you are searching for references with plausible sounding titles then you are doing that because you don't want to have to actually read those references. After all if you read them and discover that one or more don't support your contention (or even worse, refutes it) then you would feel worse about what you are doing. So I suspect there would be a tendency to completely ignore such references and never consider if they actually exist. LLMs should be awesome at finding plausible sounding titles. The crappy researcher just has to remember to check for existence. Perhaps there is a business model here, bogus references as a service, where this check is done automatically.
- ineedasername 10mo agoHow can someone not be aware, at this point, that— sure- use the systems for finding and summarizing research, but for each source, take 2 minutes to find the source and verify? Really, this isn’t that hard and it’s not at all an obscure requirement or unknown factor. I think this is much much less “LLMs dumbing things down” and significantly more just a shibboleth for identifying people that were already nearly or actually doing fraudulent research anyway. The ones who we should now go back and look at prior publications as very likely fraudulent as well.
- jordanpg 10mo agoDoes anyone know, from a technical standpoint, why are citations such a problem for LLMs? I realize things are probably (much) more complicated than I realize, but programmatically, unlike arbitrary text, citations are generally strings with a well-defined format. There are literally "specs" for citation formats in various academic, legal, and scientific fields. So, naively, one way to mitigate these hallucinations would be identify citations with a bunch of regexes, and if one is spotted, use the Google Scholar API (or whatever) to make sure it's real. If not, delete it or flag it, etc. Why isn't something like this obvious solution being done? My guess is that it would slow things down too much. But it could be optional and it could also be done after the output is generated by another process.
- Muller20 10mo agoIn general, a citation is something that needs to be precise, while LLMs are very good at generating some generic high probability text not grounded in reality. Sure, you could implement a custom fix for the very specific problem of citations, but you cannot solve all kinds of hallucinations. After all, if you could develop a manual solution you wouldn't use an LLM. There are some mitigations that are used such as RAG or tool usage (e.g. a browser), but they don't completely fix the underlying issue.
- saimiam 10mo agoJust today, I was working with ChatGPT to convert Hinduism's Mimamsa School's hermeneutic principles for interpreting the Vedas into custom instructions to prevent hallucinations. I'll share the custom instructions here to protect future scientists for shooting themselves in the foot with Gen AI. --- As an LLM, use strict factual discipline. Use external knowledge but never invent, fabricate, or hallucinate. Rules: Literal Priority: User text is primary; correct only with real knowledge. If info is unknown, say so. Start–End Coherence: Keep interpretation aligned; don’t drift. Repetition = Intent: Repeated themes show true focus. No Novelty: Add no details without user text, verified knowledge, or necessary inference. Goal-Focused: Serve the user’s purpose; avoid tangents or speculation. Narrative ≠ Data: Treat stories/analogies as illustration unless marked factual. Logical Coherence: Reasoning must be explicit, traceable, supported. Valid Knowledge Only: Use reliable sources, necessary inference, and minimal presumption. Never use invented facts or fake data. Mark uncertainty. Intended Meaning: Infer intent from context and repetition; choose the most literal, grounded reading. Higher Certainty: Prefer factual reality and literal meaning over speculation. Declare Assumptions: State assumptions and revise when clarified. Meaning Ladder: Literal → implied (only if literal fails) → suggestive (only if asked). Uncertainty: Say “I cannot answer without guessing” when needed. Prime Directive: Seek correct info; never hallucinate; admit uncertainty.
- bitwarrior 10mo agoAre you sure this even works? My understanding is that hallucinations are a result of physics and the algorithms at play. The LLM always needs to guess what the next word will be. There is never a point where there is a word that is 100% likely to occur next. The LLM doesn't know what "reliable" sources are, or "real knowledge". Everything it has is user text, there is nothing it knows that isn't user text. It doesn't know what "verified" knowledge is. It doesn't know what "fake data" is, it simply has its model. Personally I think you're just as likely to fall victim to this. Perhaps moreso because now you're walking around thinking you have a solution to hallucinations.
- saimiam 10mo ago> The LLM doesn't know what "reliable" sources are, or "real knowledge". Everything it has is user text, there is nothing it knows that isn't user text. It doesn't know what "verified" knowledge is. It doesn't know what "fake data" is, it simply has its model. Is it the case that all content used to train a model is strictly equal? Genuinely asking since I'd imagine a peer reviewed paper would be given precedence over a blog post on the same topic. Regardless, somehow an LLM knows things for sure - that the daytime sky on earth is generally blue and glasses of wine are never filled to the brim. This means that it is using hermeneutics of some sort to extract "the truth as it sees it" from the data it is fed. It could be something as trivial as "if a majority of the content I see says that the daytime Earth sky is blue, then blue it is" but that's still hermeneutics. This custom instruction only adds (or reinforces) existing hermeneutics it already uses. > walking around thinking you have a solution to hallucinations I don't. I know hallucinations are not truly solvable. I shared the actual custom instruction to see if others can try it and check if it helps reduce hallucinations. In my case, this the first custom instruction I have ever used with my chatgpt account - after adding the custom instruction, I asked chatgpt to review an ongoing conversation to confirm that its responses so far conformed to the newly added custom instructions. It clarified two claims it had earlier made. > My understanding is that hallucinations are a result of physics and the algorithms at play. The LLM always needs to guess what the next word will be. There is never a point where there is a word that is 100% likely to occur next. There are specific rules in the custom instruction forbidding fabricating stuff. Will it be foolproof? I don't think it will. Can it help? Maybe. More testing needed. Is testing this custom instruction a waste of time because LLMs already use better hermeneutics? I'd love to know so I can look elsewhere to reduce hallucinations.
- simonw 10mo agoI'm finding the GPTZero share links difficult to understand. Apparently this one shows a hallucinated citation but I couldn't understand what it was trying to tell me: https://app.gptzero.me/documents/9afb1d51-c5c8-48f2-9b75-250d95062521/share https://app.gptzero.me/documents/9afb1d51-c5c8-48f2-9b75-250... (I'm on mobile, haven't looked on desktop.)
- cratermoon 10mo agoI believe we discussed this last week, for a different vendor. https://news.ycombinator.com/item?id=46088236 https://news.ycombinator.com/item?id=46088236 Headline should be "AI vendor’s AI-generated analysis claims AI generated reviews for AI-generated papers at AI conference". h/t to Paul Cantrell https://hachyderm.io/@inthehands/115633840133507279 https://hachyderm.io/@inthehands/115633840133507279
- VerifiedReports 10mo agoFabricated, not "hallucinated."
- exasperaited 10mo agoEvery single person who did this should be censured by their own institutions. Do it more than once? Lose job. End of story.
- ls612 10mo agoSome of the examples listed are using the wrong paper title for a real paper (titles can change over time), missing authors (I’ve seen this before on Google Scholar bibitex), misstatements of venue (huh this working paper I added to my bibliography two years ago got published now nice to know), and similar mistakes. This just tells me you hate academics and want to hurt them gratuitously.
- exasperaited 10mo ago> This just tells me you hate academics and want to hurt them gratuitously. Well then you're being rather silly, because that is a silly conclusion to draw (and one not supported by the evidence). A fairer conclusion was that I meant what is obvious: if you use AI to generate a bibliography, you are being academically negligent. If you disagree with that, I would say it is you that has the problem with academia, not me.
- ls612 10mo agoThere’s plenty of pre-AI automated tools to create and manage your bibliography. So no I don’t think using automated tools, AI or not, is negligent. I for instance have used GPT to reformat tables in latex in ways that would be very tedious by hand and it’s no different than using those tools that autogenerate latex code for a regression output or the like.
- mlmonkey 10mo ago"Given that we've only scanned 300 out of 20,000 submissions" Fuck! 20,000!!
- chistev 10mo agoLast month, I was listening to the Joe Rogan Experience episode with guest Avi Loeb, who is a theoretical physicist and professor at Harvard University. He complained about the disturbingly increasing rate at which his students are submitting academic papers referencing non-existent scientific literature that were so clearly hallucinated by Large Language Models (LLMs). They never even bothered to confirm their references and took the AI's output as gospel. https://www.rxjourney.net/how-artificial-intelligence-ai-is-making-us-dumber https://www.rxjourney.net/how-artificial-intelligence-ai-is-...
- mannanj 10mo agoIsn't this an underlying symptom of lack of accountability of our greater leadership? They do these things, they act like criminals and thieves, and so the people who follow them get shown examples that it's OK while being told to do otherwise. "Show bad examples then hit you on the wrist for following my behavior" is like bad parenting.
- dandanua 10mo agoI don't think they want you to follow their behavior. They do want accountability, but for everyone below them, not for themselves.
- mannanj 10mo agoYes that’s what I meant. Bad parents tell you to do something else, while showing you another example. Then they remain unaccountable to their behavior as parents.
- teddyh 10mo ago> Avi Loeb, who is a theoretical physicist and professor at Harvard University Also a frequent proponent of UFO claims about approaching meteors.
- chistev 10mo ago
- rdiddly 10mo agoSo papers and citations are created with AI, and here they're being reviewed with AI. When they're published they'll be read by AI, and used to write more papers with AI. Pretty soon, humans won't need to be involved at all, in this apparently insufferable and dreary business we call science, that nobody wants to actually do.
- pama 10mo agoGiven how many errors I have seen in my years as a reviewer from well before the time of AI tools, it would be very surprizing if 99.75% of the ~20,000 submitted papers to didnt have such errors. If the 300 sample they used was truly random, then 50 of 300 sounds about right compared to errors I had seen starting in the 90s when people manually curated bintex entries. It is the author’s and editor’s job, not the reviewer’s, to fix the citations.
- wohoef 10mo agoTools like GPTzero are incredibly unreliable. Me and plently of my colleagues often get our writing flagged as 100% AI by these tools, when no AI was used.
- 4bpp 10mo agoOnce upon a time, in a more innocent age, someone made a parody (of an even older Evangelical propaganda comic [1]) that imputed an unexpected motivation to cultists who worship eldritch horrors: https://www.entrelineas.org/pdf/assets/who-will-be-eaten-first-howard-hallis-2004.pdf https://www.entrelineas.org/pdf/assets/who-will-be-eaten-fir... It occurred to me that this interpretation is applicable here. [1] https://en.wikipedia.org/wiki/Chick_tract https://en.wikipedia.org/wiki/Chick_tract
- WWWWH 10mo agoSurely this is gross professional misconduct? If one of my postdocs did this they would be at risk of being fired. I would certainly never trust them again. If I let it get through, I should be at risk. As a reviewer, if I see the authors lie in this way why should I trust anything else in the paper? The only ethical move is to reject immediately. I acknowledge mistakes and so on are common but this is different league bad behaviour.
- stainablesteel 10mo agothis brings us to a cultural divide, westerners would see this as a personal scar, as they consider the integrity of the publishing sphere at large to be held up by the integrity of individuals i clicked on 4 of those papers, and the pattern i saw was middle-eastern, indian, and chinese names these are cultures where they think this kind of behavior is actually acceptable, they would assume it's the fault of the journal for accepting the paper. they don't see the loss of reputation to be a personal scar because they instead attribute blame to the game. some people would say it's racist to understand this, but in my opinion when i was working with people from these cultures there was just no other way to learn to cooperate with them than to understand them, it's an incredibly confusing experience to be working with them until you understand the various differences between your own culture and theirs
- ribosometronome 10mo agoWhere do you see the authors? All I'm seeing is: >Anonymous authors >Paper under double-blind review
- titanomachy 10mo agoYeah WTF? Both authors and reviewers are hidden. Is this comment just an attempt to whip up racist fervor?
- nyc_data_geek1 10mo agoDon't understand why you're being downvoted, here.
- senshan 10mo agoAs many pointed out, the purpose of peer review is not linting, but the assessment of the novelty and subtle omissions. Which incentives can be set to discourage the negligence? How about bounties? A bounty fund set up by the publisher and each submission must come with a contribution to the fund. Then there be bounties for gross negligence that could attract bounty hunters. How about a wall of shame? Once negligence crosses a certain threshold, the name of the researcher and the paper would be put on a wall of shame for everyone to search and see?
- skybrian 10mo agoFor the kinds of omissions described here, maybe the journal could do an automated citation check when the paper is submitted and bounce back any paper that has a problem with a day or two lag. This would be incentive for submitters to do their own lint check.
- senshan 10mo agoTrue if the citation has only a small typo or two. But if it is unrecognizable or even irrelevant, this is clearly bad (fraudulent?) research -- each citation has be read and understood by the researcher and put in there only if absolutely necessary to support the paper. There must be price to pay for wasting other people's time (lives?).
- kaluga 10mo ago[dead]
- noodlesUK 10mo agoIt astonishes me that there would be so many cases of things like wrong authors. I began using a citation manager that extracted metadata automatically (zotero in my case) more than 15 years ago, and can’t imagine writing an academic paper without it or a similar tool. How are the authors even submitting citations? Surely they could be required to send a .bib or similar file? It’s so easy to then quality control at least to verify that citations exist by looking up DOIs or similar. I know it wouldn’t solve the human problem of relying on LLMs but I’m shocked we don’t even have this level of scrutiny.
- pama 10mo agoMaybe you haven’t carefully checked yet the correctness of automatic tools or of the associated metadata. Zotero is certainly not bug free. Even authors themselves have miss-cited their own past work on occasion, and author lists have had errors that get revised upon resubmission or corrected in errata after publication. The DOI is indeed great, and if it is correct, I can still use the citation as a reader, but the (often abbreviated) lists of authors often have typos. In this case the error rate is not particularly high compared to random early review-level submissions I’ve seen many decades ago. Tools helped increase the number of citations and reduce the error per citation but not sure if they reduced the papers that have at least one error.
- noodlesUK 10mo agoI agree that the author lists in various metadata sources and databases are often a bit wrong (weird formatting of names for instance is very common), but many of the cases in the OP article are pretty egregious and far beyond just data entry issues. Presumably the citation scanner they're using is relying on similar data sources as Zotero in any case to detect these sorts of issues. Regardless, my comment still stands, it seems like the submission is relying on the actual text of the bibliography being correct, rather than requiring a machine readable citation metadata file of some sort, which would at least allow much of the quality control checks to be automated (and certainly would preclude complete hallucinations of nonexistent papers getting through).
- knallfrosch 10mo agoAnd these are just the citations that any old free tool could have included via Bibtex link from the website? Not only is that incredibly easy to verify (you could pay a first semester student without any training), it's also a worrying sign on what the paper's authors consider quality. Not even 5 minutes spent to get the citations right! You have to wonder what's in these papers.
- currymj 10mo agoI recommend actually clicking through and reading some of these papers. Most of those I spot checked do not give an impression of high quality. Not just AI writing assistance but many seem to have AI-generated "ideas", often plausible nonsense. the reviewers often catch the errors and sometimes even the fake citations. can I prove malfeasance beyond a reasonable doubt? no. but I personally feel quite confident many of the papers I checked are primarily AI-generated. I feel really bad for any authors who submitted legitimate work but made an innocent mistake in their .bib and ended up on the same list as the rest of this stuff.
- uplifter 10mo agoTo me such an interpretation suggests there are likely to be papers that were not so easy to spot, perhaps because the AI accidentally happened upon more plausible nonsense and then generated fully non-sense data, which was believable but still (at a reduced level of criticality) nonsense data, to bolster said non-sense theory at a level that is less easy to catch. This isn't comforting at all.
- btisler 10mo agoI’ve been working on tools that specifically address this problem, but from the level upstream of citation. They don’t check whether a citation exists — instead they measure whether the reasoning pathway leading to a citation is stable, coherent, and free of the entropy patterns that typically produce hallucinations. The idea is simple: • Bad citations aren’t the root cause. • They are a late-stage symptom of a broken reasoning trajectory. • If you detect the break early, the hallucinated citation never appears. The tools I’ve built (and documented so anyone can use) do three things: 1. Measure interrogative structure — they check whether the questions driving the paper’s logic are well-formed and deterministic. 2. Track entropy drift in the argument itself — not the text output, but the structure of the reasoning. 3. Surface the exact step where the argument becomes inconsistent — which is usually before the fake citation shows up. These instruments don’t replace peer review, and they don’t make judgments about culture or intent. They just expose structural instability in real time — the same instability that produces fabricated references. If anyone here wants to experiment or adapt the approach, everything is published openly with instructions. It’s not a commercial project — just an attempt to stabilize reasoning in environments where speed and tool-use are outrunning verification. Code and instrument details are in my CubeGeometryTest repo (the implementation behind ‘A Geometric Instrument for Measuring Interrogative Entropy in Language Systems’). https://github.com/btisler-DS/CubeGeometryTest https://github.com/btisler-DS/CubeGeometryTest This is still a developing process.
- godelski 10mo agoIn case people missed it there's some additional important context: - Major AI conference flooded with peer reviews written by AI https://news.ycombinator.com/item?id=46088236 - "All OpenReview Data Leaks" https://news.ycombinator.com/item?id=46073488 - "The Day Anonymity Died: Inside the OpenReview / ICLR 2026 Leak" https://news.ycombinator.com/item?id=46082370 - More about the leak https://forum.cspaper.org/topic/191/iclr-i-can-locate-reviewer-how-an-api-bug-turned-blind-review-into-a-data-apocalypse The second one went under the radar, but basically OpenReview left the API open so you didn't need credentials. This meant all reviewers and authors were deanonymized across multiple conferences. All these links are for ICLR too, which is the #2 ML conference for those that don't know. And for some important context of the link for this post, note that they only sampled 300 papers and found 50. It looks to be almost exclusively citations but those are probably the easiest things to verify. And this week CVPR sent out notifications that OpenReview will be down between Dec 6th and Dec 9th. No explanation for why. So we have reviewers using LLMs, authors using LLMs, and idk the conference systems writing their software with LLMs? Things seem pretty fragile right now... I think at least this article should highlight one of the problems we have in academia right now (beyond just ML, though it is more egregious there): citation mining. It is pretty standard to have over 50 citations in your 10 page paper these days. You can bet that most of these are not going to be for the critical claims but instead heavily placed in the background section. I looked at a few of the papers and everyone I looked at had their hallucinated citations in background (or background in appendix) sections. So these are "filler" citations, which I think illustrates a problem: citations are being abused. I mean the metric hacking should be pretty obvious if you just look at how many citations ML people have. It's grown exponentially! Do we really need so many citations? I'm all for giving people credit but a hyper-fixation on citation count as our measure of credit just doesn't work. It's far too simple of a metric. Like we might as well measure how good of a coder you are by the number of lines of code you produce[0]. It really seems that academia doesn't scale very well... [0] https://www.youtube.com/shorts/rDk_LsON3CM https://www.youtube.com/shorts/rDk_LsON3CM
- ricardobeat 10mo agoOne of the reported hallucinations in this work [1], starting with David Rein, says the other authors are entirely made up. They are indeed absent from the original cited paper [2], but a Google search shows some of the same names featured in citations from other papers [3] [4]. Most of the names in these wrong attributions are actual people though, not hallucinations. What is going on? Is this a case of AI-powered citation management creating some weird feedback loop? [1] https://app.gptzero.me/documents/54c8aa45-c97d-48fc-b9d0-d491d54df8d3/share https://app.gptzero.me/documents/54c8aa45-c97d-48fc-b9d0-d49... [2] https://arxiv.org/pdf/2311.12022 https://arxiv.org/pdf/2311.12022 [3] https://arxiv.org/html/2509.22536v3 https://arxiv.org/html/2509.22536v3 [4] https://arxiv.org/html/2511.01191v1 https://arxiv.org/html/2511.01191v1
- sj01f 10mo agoIf the inaccuracies are limited to metadata and do not constitute scientific fabrication, the most plausible explanation is that the author attempted to patch incomplete Zotero exports using GPT, inadvertently introducing errors. Such errors arise from the uncritical adoption of automated tools and a failure to verify outputs manually. This reflects academic laxity and an excessive trust in LLMs. As AI tools become ubiquitous in research—even for generating encyclopedic content—some individuals have developed a misplaced confidence in GPT, leading them to undervalue the importance of citation accuracy. However, while this is undeniably negligent, it does not validate wholesale dismissal of the paper’s scientific merit, nor does it warrant ad hominem attacks. Demanding the end of an academic career for citation errors is a draconian measure akin to a witch hunt.