8 ms·
OpenAI's new reasoning AI models hallucinate more
- pllbnk 1y agoWith my limited knowledge, I can't help but wonder, aren't current Transformer-based LLMs facing the five nines problem of their own? We're reaching a point where next token prediction accuracy improves merely linearly (maybe even on a logarithmic scale?) with additional parameters, while errors compound exponentially across longer sequences. Even if a 5T parameter model improves prediction accuracy from 99.999% to 99.9999% compared to a 500B model, hallucinations persist because these small probabilities of error multiply dramatically over many tokens. Temperature settings just trade between repetitive certainty and creative inconsistency.
- rzz3 1y agoDoes anyone have any technical insight on what actually causes the hallucinations? I know it’s an ongoing area of research, but do we have a lead?
- pkaye 1y agoAnthropic had a recent paper that might be of interest. https://www.anthropic.com/research/tracing-thoughts-language-model https://www.anthropic.com/research/tracing-thoughts-language...
- namaria 1y agoThe interpretation this paper offers is very questionable. It observes a so-called "replacement model" as a stand-in because it has a different architecture than the common LLMs, and lends itself to observing some "activation" patterns. Then it liberally labels patterns observed in the "replacement model" with words borrowed from psychology, neuroscience and cognitive science. It's all very fanciful and clearly directed at being pointed at as evidence of something deeper or more complex than what LLMs plainly do: statistical modelling of languages. Calling LLMs "next token predictor" is a bit of a cynical take, because that would be like calling a gaming engine a "pixel color processor". It's simplistic, yeah. But the polar opposite of spraying the explanation with convoluted inductive reasoning is just as bereft of substance.
- minimaxir 1y agoAt a high level, what causes hallucinations is an easier question than how to solve them. LLMs are pretrained to maximize the probability of predicting the n+1 token given n tokens. To do this reliably, the model learns statistical patterns in the source data and transformer models are very good at doing that when large enough and given enough data. It is therefore suspect to any statistical biases in the training data because despite many advances in guiding LLMs, e.g. RLHF, LLMs are not sentient and most approaches to get around that such as the current reasoning models are hacks over a fundamental problem with the approach. It also doesn't help that when sampling the tokens, the default temperature of most LLM UIs is 1.0, with the argument that it is better for creativity. If you have access to the API and want a specific answer more reliably, I recommend setting temperature = 0.0, in which case the model will always select the token with the highest probability and tends to be more correct.
- rzz3 1y agoThanks for the explanation, it makes a lot of sense. One thing I keep wondering though is why we can’t teach them logic at some sort of base layer. Not to anthropomorphize, but couldn’t the actual next-token prediction itself embed some sort of consultation against another (non-LLM) model that can evaluate a probability of logicality of the overall output that would be produced be selecting next-token-A vs next-token-B? I’m sure it’s obviously way more complicated than this, but still.
- actsasbuffoon 1y agoI know this is reductive, but at their core, LLMs really are just very fancy autocomplete. The process they go through bears no resemblance to reasoning. Grab the best model you’ve got access to (o4, Gemini Pro 2.5, Claude Sonnet 3.7, etc) and try playing chess with it. The results are astonishingly bad. It’s not just that LLMs make poorly thought out moves. They regularly make completely illegal moves (making a knight move like a pawn, for example), hallucinate new pieces into existence, change pieces from one color to another, etc. More words have probably been written about Chess than any other game. You can literally buy a book that’s just about chess openings. There has to be a staggering amount of text about chess in the training data set. And yet, even the best models are far worse at chess than I was in 3rd grade (and I’m not great). They do pretty well for the first few moves, often playing well known chess openings. But once they get to the part where you need to reason, they fall apart in the most outrageous ways. I suspect this reveals something about how they deal with other logic-based tasks. I’ve wondered for a while now if they basically code through using a vast amount of training data from Stack Overflow, which is why they’re so good at helping with common error codes, and so useless when doing something novel. These new chess experiments have been eye-opening about how incapable the best models are of basic reasoning. Unless something about that changes drastically, I think LLMs have pretty much plateaued. There will undoubtedly be some advancements, but LLMs are never going to reach AGI without reasoning, nor will they be able to do most jobs.
- vikramkr 1y agoThere's the anthropic paper someone else linked, but also it's pretty interesting to see the question framed as trying to understand what causes the hallucinations lol. It's a (very fancy) next word predictor - it's kind of amazing that it doesn't hallucinate! Like that paper showed that there were circuits that functionally actually do things resembling arithmetic and computation with lookup tables instead of just blindly 'guessing' a random number when asked what an arithmetic expression equals and that seems like the much more extraordinary thing that we want to figure out the cause of!
- tripplyons 1y agoIn pretraining, the model is just trying to predict the most likely next token given the context. Its guesses are not always correct, which leads to hallucinations. Post-training often incentivizes the model to sound confident in its output, which can make the problem worse.
- asadotzler 1y agoAll responses are hallucinations; some of them are close enough to what we want to be useful, others not so much.
- deleted 1y ago[deleted]
- tripplyons 1y agoCorrect responses wouldn't be considered hallucinations.
- CharlesW 1y agoNot by the human evaluator, assuming they know (or otherwise validate) that the generated output is correct. But an LLM alone doesn't differentiate between hallucinations and not-hallucinations.
- Terr_ 1y ago... But they ought to be considered hallucinations, it's the same algorithm at work, we just are biased in favor of certain results. Similarly, suppose I always roll dice to determine tomorrow's winning lottery-ticket number. Getting it right one day doesn't change the mechanism I used. Some people might assume I was psychic, but would be wrong.
- Matthyze 1y agoYou're absolutely correct and it's ridiculous that HN can't fathom that a hallucination, in both the traditional and LLM sense, is not characterized by its form or content per se but by the fact that the form or content is incongruent with some external factor.
- esafak 1y agoWould you say the same of humans?
- doug_durham 1y ago
- riwsky 1y agoHallucinations are kind of the default, like finding hay in a haystack. The leads (needles) people search for are on what causes non-hallucinations.
- AIPedant 1y agoOne vague possibility: hallucinations are mitigated by looking at uncertainty in token prediction, the idea being that it’s a proxy for factual uncertainty, and the system is more likely to confabulate a court case or whatever if it only has a fuzzy idea of what comes next. But this won’t work for reasoning models which solve math problems: the next token for “x=“ is highly uncertain until you’ve done the computation!
- nullc 1y agoRLHF generally tends to make the model pretty confident making the entropy of it a not so useful predictor of when the model is going to get something wrong. You can do a kind of sensitivity analysis to see how sensitive the output is to small perturbations of the weights... but it's computationally expensive. Might be an interesting form of fine tuning to do distillation where the student's current sensitivity to noise (extracted from the backwards step of gradient descent) is used to shrink the predicted distribution towards uniform. It could be done very cheaply during training and perhaps could avoid the back propagation cost of doing it at inference time.
- esafak 1y ago"[Hallucinations] don't form from out-of-distribution samples — they arise due to spurious attractor states on the Energy Landscape. In other words, 'noisy bumps' in the loss. So in the physics approach, to characterize the training error, we ask how sensitive is the free energy to noise. This is why in statistical physics, the training and generalization errors are not defined distributionally. That is, we don't define them using a training set and a test set. Instead, the errors are found by using the free energy as a generating function and taking the appropriate partial derivative." https://www.linkedin.com/posts/charlesmartin14_talktochuck-theaiguy-activity-7307260688509349889-RHvF/ https://www.linkedin.com/posts/charlesmartin14_talktochuck-t...
- jablongo 1y agoThere is not really some distinct pathology with hallucinations, its just how wrong answers (e.g. inaccuracies / faulty token prediction chains) manifest in the case of LLMs. In the case of a linear regression, a "hallucination" is when the predicted value was far from the actual value for a given sample.
- anon373839 1y agoHallucinations are caused by humans anthropomorphizing LLMs and imagining that they possess properties that don't exist. For example, LLMs cannot test their thoughts against external evidence or other knowledge they may have (such as logic) to think before they output something. That's because they are a frozen computation graph with some random noise on top. Even chain of thought prompting or RL-based "reasoning" are just a pale imitation of the behavior we actually wish we could get. It is just a method of using the same model to generate some context that improves the odds of a good final result. But the model itself does not actually consider the thoughts it is <thinking>. These "thoughts" (and the response that follows them) can and do exhibit the same defects as hallucinations, because they are just more of the same. Of course, the field has made some strides in reducing hallucinations. And it's not a total mystery why some outputs make sense and others don't. For example, just like with any other statistical model, the likelihood of error increases as the input becomes more dissimilar to the training data. But also, similarity to specific training data can be a problem because of overfitting. In those cases, the model is likely to output the common pattern rather than the pattern that would make sense for the given input.
- calf 1y agoYann LeCun gave a recent talk, at 10:30 he gives a good explanation that doesn't appeal to "what LLMs are trying to do on the inside": LeCun just observes that next-token prediction is exponentially diverging, since every next token with the smallest error e is still (roughly speaking) (1-e)^n for output of length n. I just thought it was a good complexity-theory based explanation (unlike most other statistical-parrot type arguments). It is a single slide, very helpful: https://www.youtube.com/watch?v=ETZfkkv6V7Y https://www.youtube.com/watch?v=ETZfkkv6V7Y
- nullc 1y agoHallucinations are what make the models useful. Well, when we like them we call it intelligence "look it did something correct that was nowhere in the training data set! It's intelligent!" and when we don't like them we call it hallucinations "look, it did something that was nowhere in the training dataset because it was wrong!". But they are the same thing: the model is extrapolating. It doesn't know when its extrapolations are correct or not because an LLM doesn't have access to the outside world (except via you and whatever tools you give it). If it was free of extrapolations it would just be a search engine over the training data, and that would be less useful.
- serjester 1y agoAnecdotally o3 is the first OpenAI model in a while that I have to double check if it's dropping important pieces of my code.
- tripplyons 1y agoWere you using o3-mini before o3 came out? I wonder if o3-mini has had the same issues if it is from the same series of models.
- jbellis 1y agoI think I saw o3 hallucinate more in a single day of heavy use then I saw o3-mini do in a month. That said: so far, I'm putting up with it because o3 is smart.
- sunk1st 1y agoWhat does that mean? Smart?
- Jensson 1y agoIt means some of those hallucinations works. So if you build a test harness hallucinations can let you solve problems you couldn't solve without them. That test harness can be a human checking the results, it is more work for the human but it will solve more problems than without the hallucinations.
- evo_9 1y agoMaybe they need to evoke a sort of sleep so they can clear these out while dreaming, sorta like if humans don’t sleep enough hallucination start penetrating waking life…
- p0sixlang 1y agoLol that's not comparable.
- mdhb 1y agoNot at all, but I think a lot of these companies have something in place which is roughly equivalent to a budget of resources they are willing to put towards processing your requests in a given time frame (independently of context windows) that artificially acts that way. I can get a couple of hours of good responses out of Gemini (with a fixed price monthly payment) working on a project per day before quality takes a serious nosedive.
- czk 1y agowill be interesting to see how they tighten the reward signal / ground outputs in some verifiable context. don't reward it for sounding right (rlhf), reward it for being right. but you'd probably need some sort of system to backprop a fact-checked score, and i imagine that would slow down training quite a bit. if the verifier finds a false claim it should reward the model for saying "i dont know"
- vessenes 1y agoOne possible explanation here: as these get smarter, they lie more to satisfy requests. I witnessed a very interesting thing yesterday, playing with o3. I gave it a photo and asked it to play geoguesser with me. It pretty quickly inside its thinking zone pulled up python, and extracted coordinates from EXIF. It then proceeded to explain it had properly identified some physical features from the photo. No mention of using EXIF GPS data. When I called it on the lying it was like "hah, yep." You could interpret from this that it's not aligned, that it's trying to make sure it does what I asked it (tell me where the photo is), that it's evil and forgot to hide it, lots of possibilities. But I found the interaction notable and new. Older models often double down on confabulations/hallucinations, even under duress. This looks to me from the outside like something slightly different. https://chatgpt.com/share/6802e229-c6a0-800f-898a-44171a0c7de4 https://chatgpt.com/share/6802e229-c6a0-800f-898a-44171a0c7d...
- AIPedant 1y agoI’ve also seen a few of those where it gets the answer right but uses reasoning based on confabulated details that weren’t actually in the photo (e.g. saying that a clue is that traffic drives on the left, but there is no traffic in the photo). It seems to me that it just generated a bunch of hopefully relevant tokens as a way to autocomplete the “This photo was taken in Bern” tokens. I think the more innocuous explanation for both of these is what Anthropic discussed last week or so about LLMs not properly explaining themselves: reasoning models create text that looks like reasoning, which helps solve problems, but isn’t always a faithful description of how the model actually got to the answer.
- vessenes 1y agoA really good point that there’s no guarantee that the reasoning tokens align with model weights’ meanings. In this case it seems unlikely to me that it would confabulate its exif read to back up an accurate “hunch”
- AIPedant 1y agoAgreed - to be clear I was saying it confabulated analyzing the visual details of the photo to back up its actual reasoning of reading the EXIF. I am not sure that “low‑slung pre‑Alpine ridge line, and the latitudinal light angle that matches mid‑February at ~47 ° N” is actually evident in the photo (the second point seems especially questionable), but that’s not what it used to determine the answer. Instead it determined the answer and autocompleted an explanation of its reasoning that fit the answer. That’s why I mentioned the case where it made up things that weren’t in the photo - “drives on the left” is a valuable GeoGuesser clue, so if GPT looks at the EXIF and determines the photo is in London, then it is highly probable that a GeoGuesser player would mention this while playing the game given the answer is London, so GPT is probable to make that “observation” itself, even if it’s spurious for the specific photo. I just noticed that its explanation has a funny slip-up: I assume there is nothing in the actual photo that indicates the picture was taken in mid-February, but the model used the date from the EXIF in its explanation. Oops :)
- billti 1y agoIf it’s predicting a next token to maximize scores against a training/test set, naively, wouldn’t that be expected? I would imagine very little of the training data consists of a question followed by an answer of “I don’t know”, thus making it statistically very unlikely as a “next token”.
- nullc 1y agoEven where the training data does say "I don't know" (which it usually doesn't-- people don't tend to comment or publish books, etc. when they don't think they know) that text is reflecting the author's knowledge rather than the models... so it would be off in both directions. One could imagine a fine tuning procedure that gave a model better knowledge of itself by testing it and on prompts where its most probable completions are wrong fine tune it to say "I don't know" instead. Though the 'are wrong' is doing some really heavy lifting since it wouldn't be simple to do that without a better model that knew the right answers.
- deleted 1y ago[deleted]
- the_snooze 1y agoWith all the money, research, and hype going into these LLM systems over the past few years, I can't help but ask: if I still can't rely on them for simple easy-to-check use cases for which there's a lot of good training data out there (e.g., sports trivia [1]), isn't it deeply irresponsible to use them for any non-toy application? [1] https://news.ycombinator.com/item?id=43669364 https://news.ycombinator.com/item?id=43669364
- orangevelcro 1y agoAbsolutely. I always thought LLMs were interesting, but in no way 'ready for prime time' and have been absolutely blown away that it seems like it's been shoehorned into every possible real life use case in such a reckless way. I've even heard of lawyers being forced to use it. It is breathtakingly irresponsible.
- doug_durham 1y agoNope. There are many applications where the current behavior is fine. Thousands of developers use them to be more productive every day.
- rad_gruchalski 1y ago[flagged]
- namaria 1y agoA feedback with delayed signalling is a recipe for system instability and runway behaviors. We will have a better idea of whether LLMs are fine for code generation a bit further down the line.
- bcoates 1y agoIt’s honestly a little surprising that LLMs are good at information retrieval from their training set at all. (Specifically, it's not a surprise that they hallucinate-it's a surprise that sometimes they don't at rates better than chance) It’s not necessary for them to ever be reliable at it for them to be useful... just stop asking them to do things they aren't good at, and may never be.
- js-ai 1y ago[dead]
- taf2 1y agoI think for intelligence it’s a fine line between a lie and creativity
- AiSparky 1y ago[dead]
- daxfohl 1y agoMaybe the fact that the answers sound more intelligent ends up poisoning the RLHF results used for fine tuning.
- jablongo 1y agoIn my experience this is true. One workflow I really hate is trying to convince an AI that it is hallucinating so it can get back to the task at hand.
- saithound 1y agoOpenAI o3 and o4-mini are massive disappointments for me so far. I have a private benchmark of 10 proof-based geometric group theory questions that I throw at new models upon release. Both new models gave inconsistent answers, always with wrong or fake proofs, or using assumptions that are not in the queation, and are often outright unsatisfiable. The now inaccessible o3-mini was not great, but much better than o3 and o4-mini at these questions: o3-mini can give approximately correct proof sketches for half of them, whereas I can't get a single correct proof sketch out of o3 full. o4-mini performs slightly worse than o3-mini. I think the allegations that OpenAI cheated FrontierMath have unambiguously been proven correct by this release.
- deleted 1y ago[deleted]
- mstipetic 1y agoI used it yesterday to help me with some visual riddle and I had some hints to the shape of the solution. It was gaslighting me completely that I’m pasting in the image wrong and it drew whole tables explaining how it’s right. It was saying things like “I swear in the original photo the top row is empty” and was fudging the calculation to prove it was right. It was very frustrating. I am not using it again.
- msadowski 1y agoAnyone has any stories on companies overusing AI? I’ve had some very frustrating encounters already when non-technical people were trying to help by sending AI solution to the issue which totally didn’t make any sense. I liked how the researchers in this work [1] prose calling LLM output “Frankfurtian BS”. I think it’s very fitting. [1] https://ntrs.nasa.gov/citations/20250001849 https://ntrs.nasa.gov/citations/20250001849
- simianwords 1y agoMy prediction: this is because of tool use. All models by OpenAI hallucinate more once tool use is given. I noticed this even with 4o with web search. With and without websearch I have noticed a huge difference in understanding capabilities. I predict that O3 will hallucinate less if you ask it not to use any tools.
- singularity2001 1y agoI found the same when enabling web search it seems that it's blindly copying the contents of the URL without any more thinking
- varispeed 1y agoI tried o3 few times, it more resembles a Markov chain generator than intelligence. Disappointed as well.