3 ms·
Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is
by aliljet 5mo ago
Is there a good benchmark tracking hallucinations? The models are all incredibly good now, even the open ones, and my hope is that the rate of hallucinations is something that's falling off in concert with larger and larger context lengths.
- Sevii 5mo agoI haven't been bothered by hallucinations in premier models since early last year. Still see it in smaller local models though.
- aliljet 5mo agoI'm really running into this deep at the edges of content creation. Take, for example, a need to general some kind of legal work. The cost of painstakingly checking and rechecking each case cited is reducing the value of these frontier models immensely. Coding, however, is solved like magic. Easier to add tests, to be fair.
- throawayonthe 5mo agowell there is https://artificialanalysis.ai/evaluations/omniscience https://artificialanalysis.ai/evaluations/omniscience
- goldenarm 5mo agoIt's a gibberish input detection benchmark, and does not measure output hallucinations.
- yieldcrv 5mo agoif last year's models were the ones people got familiar with in late 2022, hallucinations would be an underrepresented rumor, there would be no articles about it because its so rare. overconfident lawyers wouldn't have messed up dockets in court with fake case law, in other domains that move faster, sources would be only partially outdated with agentic search and mcp servers filling in the gaps AI psychosis would be the problem people talk about more, not just outright agreement but subtle ways of making you feel confident in your ideas. "yes, buy that domain name buy these other ones for defensibility" (the domain name is dumb and completely unmarketable)
- jampekka 5mo agoThe models still hallucinate bad when called via APIs, especially if web search is not enabled. Gemini hallucinates quite frequently even with the app and search enabled. More recent (e.g. ChatGPT 5.x and Deepseek v4) prompts/harnesses search very aggressively, which does greatly mitigate hallucinations.
- schneehertz 5mo agoVictim of LLM hallucinations, poor guy
- majso 5mo agomaybe something like this? https://petergpt.github.io/bullshit-benchmark/viewer/index.v2.html https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
- WarmWash 5mo agoPeople complain about them incessantly, but I can almost never get people to actually post receipts. Every provider allows sharing chats, and anyone can share a prompt that reliably produces hallucinations. More often than not, people are using images in responses that go awry. Which is fair, the models are sold as multi-modal, but image analyses is still at gpt-4.0 text-analyses levels. Also knowledge cutoff issues, where people forget the models exist months to a year or more in the past.
- saberience 5mo agoI see hallucinations ALL the time. It's only obvious when you're prompting about a subject you know well. And when I say all the time, I mean it, and this is for Opus 4.7 Adaptive. I often have to say, please do searches and cite sources, as if it doesn't it will confidently give me wrong or outdated information. If you're often asking questions about a topic that's not in your specialist knowledge you won't notice them.
- droidjj 5mo agoHallucination is also much better controlled in the context of agentic coding because outputs can be validated by running the code (or linters/LSP). I almost never notice hallucinations when I’m coding with AI, but when using AI for legal work (my real job) it hallucinates constantly and perniciously because the hallucinations are subtle—e.g., making up a crucial fact about a real case.
- krupan 5mo agoYes, you can catch many mistakes that LLMs make whike coding, but I wouldn't necessarily call it "controlled." Every now and then the LLM will run into dead ends where it makes a certain mistake, the compiler or unit tests find the mistake, so it tries a different approach that also fails, and then it goes back to the first approach, then tries the second approach again, and gets stuck in an endless loop trying small variations on those two approaches over and over. If you aren't paying attention it can spend a long time (and a lot of tokens) spinning in that loop. Sometimes there might be more than two approaches in the loop, which makes it even harder to see that it's repeating itself in a loop. It's pretty frustrating to see it working away productively (so you think) for 20 minutes or so only to finally notice what's going on
- FergusArgyll 5mo agoAs long as the model uses web search, they almost never hallucinate anymore. The fast models (haiku, gpt-instant, flash) still sometimes have the problem where they don't search before answering so they can hallucinate
- goldenarm 5mo agoI've seen chatGPT and Gemini hallucinate even from web search, it's better is not sufficient
- krupan 5mo agoIt really depends what you are asking it. If the answer is in the training data, then the odds of it lying to you are much lower than if you are asking it for something it has never seen before.
- vlmutolo 5mo ago> While OpenAI originally pioneered Codex (which went on to power GitHub Copilot), Google’s direct answer for dedicated, native code completion and natural-language-to-code generation is CodeGemma. https://g.co/gemini/share/33e7a589a161 https://g.co/gemini/share/33e7a589a161
- deaux 5mo agoNothing about this is a hallucination. The Codex that it talks about is real, existed, and did go on to power the original Copilot. You neither specified that you meant a different Codex, nor did it make anything up. The CodeGemma isn't made up either, as its referenced working link shows.