7 ms·
"They claim impressive reductions in hallucinations. In my own usage I’ve not spotted a single hallucination yet, but that’s been true for me for Claude 4 and o
by techpression 1y ago
"They claim impressive reductions in hallucinations. In my own usage I’ve not spotted a single hallucination yet, but that’s been true for me for Claude 4 and o3 recently as well—hallucination is so much less of a problem with this year’s models."
This has me so confused, Claude 4 (Sonnet and Opus) hallucinates daily for me, on both simple and hard things. And this is for small isolated questions at that.
- simonw 1y agoWhat kind of hallucinations are you seeing?
- OtherShrezzing 1y agoI rewrote a 4 page document from first to third person a couple of weeks back. I gave Claude Sonnet 4 the document after editing, so it was entirely written in the third person. I asked it to review & highlight places where it was still in the first person. >Looking through the document, I can identify several instances where it's written in the first person: And it went on to show a series of "they/them" statements. I asked it to clarify if "they" is "first person" and it responded >No, "they" is not first person - it's third person. I made an error in my analysis. First person would be: I, we, me, us, our, my. Second person would be: you, your. Third person would be: he, she, it, they, them, their. Looking back at the document more carefully, it appears to be written entirely in third person. Even the good models are still failing at real-world use cases which should be right in their wheelhouse.
- simonw 1y agoThat doesn't quite fit the definition I use for "hallucination" - it's clearly a dumb error, but the model didn't confidently state something that's not true (like naming the wrong team who won the Super Bowl).
- OtherShrezzing 1y ago>"They claim impressive reductions in hallucinations. In my own usage I’ve not spotted a single hallucination yet, but that’s been true for me for Claude 4 and o3 recently as well—hallucination is so much less of a problem with this year’s models." Could you give an estimate of how many "dumb errors" you've encountered, as opposed to hallucinations? I think many of your readers might read "hallucination" and assume you mean "hallucinations and dumb errors".
- jmull 1y agoThat's a good way to put it. As a user, when the model tells me things that are flat out wrong, it doesn't really matter whether it would be categorized as a hallucination or a dumb error. From my perspective, those mean the same thing.
- simonw 1y agoI mention one dumb error in my post itself - the table sorting mistake. I haven't been keeping a formal count of them, but dumb errors from LLMs remain pretty common. I spot them and either correct them myself or nudge the LLM to do it, if that's feasible. I see that as a regular part of working with these systems.
- OtherShrezzing 1y agoThat makes sense, and I think your definition on hallucinations is a technically correct one. Going forward, I think your readers might appreciate you tracking "dumb errors" alongside (but separate from) hallucinations. They're a regular part of working with these systems, but they take up some cognitive load on the part of the user, so it's useful to know if that load will rise, fall, or stay consistent with a new model release.
- godelski 1y agoI think it qualifies as a hallucination. What's your definition? I'm a researcher too and as far as I'm aware the definition has always been pretty broad and applied to many forms of mistakes. (It was always muddy but definitely got more muddy when adopted by NLP) It's hard to know why it made the error but isn't it caused by inaccurate "world" modeling? ("World" being English language) Is it not making some hallucination about the English language while interpreting the prompt or document? I'm having a hard time trying to think of a context where "they" would even be first person. I can't find any search results though Google's AI says it can. It provided two links, the first being a Quora result saying people don't do this but framed it as it's not impossible, just unheard of. Second result just talks about singular you. Both of these I'd consider hallucinations too as the answer isn't supported by the links.
- techpression 1y agoSince I mostly use it for code, made up function names are the most common. And of course just broken code all together, which might not count as a hallucination.
- ewoodrich 1y agoI think the type of AI coding being used also has an effect on a person's perception of the prevalence of "hallucinations" vs other errors. I usually use an agentic workflow and "hallucination" isn't the first word that comes to my mind when a model unloads a pile of error-ridden code slop for me to review. Despite it being entirely possible that hallucinating a non-existent parameter was what originally made it go off the rails and begin the classic loop of breaking things more with each attempt to fix it. Whereas for AI autocomplete/suggestions, an invented method name or argument or whatever else clearly jumps out as a "hallucination" if you are familiar with what you're working on.
- laacz 1y agoI suppose that Simon, being all in with LLMs for quite a while now, has developed a good intuition/feeling for framing questions so that they produce less hallucinations.
- simonw 1y agoYeah I think that's exactly right. I don't ask questions that are likely to product hallucinations (like citations from papers about a topic to an LLM without search access), so I rarely see them.
- godelski 1y agoBut how would you verify? Are you constantly asking questions you already know the answers to? In depth answers? Often the hallucinations I see are subtle, though usually critical. I see it when generating code, doing my testing, or even just writing. There are hallucinations in today's announcements, such as the airfoil example[0]. An example of more obvious hallucinations is I was asking for help improving writing an abstract for a paper. I gave it my draft and it inserted new numbers and metrics that weren't there. I tried again providing my whole paper. I tried again making explicit to not add new numbers. I tried the whole process again in new sessions and in private sessions. Claude did better than GPT 4 and o3 but none would do it without follow-ups and a few iterations. Honestly I'm curious what you use them for where you don't see hallucinations [0] which is a subtle but famous misconception. One that you'll even see in textbooks. Hallucination probably caused by Bernoulli being in the prompt
- simonw 1y agoWhen I'm using them for code these days it is usually in a tool that can execute code in a loop - so I don't tend to even spot the hallucinations because the model self corrects itself. For factual information I only ever use search-enabled models like o3 or GPT-4. Most of my other use cases involve pasting large volumes of text into the model and having it extract information or manipulates that text in some way.
- bluetidepro 1y agoAgreed. All it takes is a simple reply of “you’re wrong.” to Claude/ChatGPT/etc. and it will start to crumble on itself and get into a loop that hallucinates over and over. It won’t fight back, even if it happened to be right to begin with. It has no backbone to be confident it is right.
- cameldrv 1y agoYeah it may be that previous training data, the model was given a strong negative signal when the human trainer told it it was wrong. In more subjective domains this might lead to sycophancy. If the human is always right and the data is always right, but the data can be interpreted multiple ways, like say human psychology, the model just adjusts to the opinion of the human. If the question is about harder facts which the human disagrees with, this may put it into an essentially self-contradictory state, where the locus of possibilitie gets squished from each direction, and so the model is forced to respond with crazy outliers which agree with both the human and the data. The probability of an invented reference being true may be very low, but from the model's perspective, it may still be one of the highest probability outputs among a set of bad choices. What it sounds like they may have done is just have the humans tell it it's wrong when it isn't, and then award it credit for sticking to its guns.
- ashdksnndck 1y agoI put in the ChatGPT system prompt to be not sycophantic, be honest, and tell me if I am wrong. When I try to correct it, it hallucinates more complicated epicycles to explain how it was right the first time.
- diggan 1y ago> All it takes is a simple reply of “you’re wrong.” to Claude/ChatGPT/etc. and it will start to crumble on itself and get into a loop that hallucinates over and over. Yeah, it's seems to be a terrible approach to try to "correct" the context by adding clarifications or telling it what's wrong. Instead, start from 0 with the same initial prompt you used, but improve it so the LLM gets it right in the first response. If it still gets it wrong, begin from 0 again. The context seems to be "poisoned" really quickly, if you're looking for accuracy in the responses. So better to begin from the beginning as soon as it veers off course.
- squeegmeister 1y agoYeah hallucinations are very context dependent. I’m guessing OP is working in very well documented domains
- Oras 1y agoHere you go https://pbs.twimg.com/media/Gxxtiz7WEAAGCQ1?format=jpg&name=4096x4096 https://pbs.twimg.com/media/Gxxtiz7WEAAGCQ1?format=jpg&name=...
- simonw 1y agoHow is that a hallucination?
- madduci 1y agoI believe it depends in inputs. For me, Claude 4 has consistently generated hallucinations, especially was pretty confident in generating invalid JSONs, for instance Grafana Dashboards, which were full of syntactic errors.
- godelski 1y agoThere were also several hallucinations during the announcement. (I also see hallucinations every time I use Claude and GPT, which is several times a week. Paid and free tiers) So not seeing them means either lying or incompetent. I always try to attribute to stupidity rather than malice (Hanlon's razor). The big problem of LLMs is that they optimize human preference. This means they optimize for hidden errors. Personally I'm really cautious about using tools that have stealthy failure modes. They just lead to many problems and lots of wasted hours debugging, even when failure rates are low. It just causes everything to slow down for me as I'm double checking everything and need to be much more meticulous if I know it's hard to see. It's like having a line of Python indented with an inconsistent white space character. Impossible to see. But what if you didn't have the interpreter telling you which line you failed on or being able to search or highlight these different characters. At least in this case you'd know there's an error. It's hard enough dealing with human generated invisible errors, but this just seems to perpetuate the LGTM crowd
- hhh 1y agoYou can just have a different use case that surfaces hallucinations than someone, they don’t have to by evil.
- simonw 1y agoWhat were the hallucinations during the announcement? My incompetence here was that I was careless with my use of the term "hallucination" here. I assumed everyone else shared my exact definition - that a hallucination is when a model confidently states a fact that is entirely unconnected from reality, which is a different issue from a mistake ("how many Bs in blueberry" etc). It's clear that MANY people do not share my definition! I deeply regret including that note in my post.
- davidmurdoch 1y agoI subscribe to the same definition as you. I've actually never heard someone referring to the mistakes as hallucinating until now, but I can see how it's a bit of a grey area.
- simonw 1y agoI updated that section of my post with a clarification about what I meant. Thanks for calling this out, it definitely needed extra context from me.