18 ms·
AI hallucinations: Why LLMs make things up (and how to fix it)
- sollewitt 2y ago"Why LLMs do the one and only thing they do (and how to fix it)"
- mdaniel 2y agothe only comment on the prior submission 3 days ago summarizes the whole thing: https://news.ycombinator.com/item?id=42285149 https://news.ycombinator.com/item?id=42285149 Also, I saw any such blog title as "how to make money in the stock market:" friend, if you knew the answer you wouldn't blog about it you'd be infinitely rich
- rtsil 2y agoThey don't say how to make big money, don't they? Also, the tl;dr is index funds and patience.
- Terr_ 2y agoWhen people talk about stopping an LLM from "seeing hallucinations instead of the truth", that's like stopping an Ouija-board from "channeling the wrong spirits instead of the right spirits." It suggests a qualitative difference between desirable and undesirable operation that isn't really there. They're all hallucinations, we just happen to like some of them more than others.
- ml_more 2y agoThe problem is that LLMs are just convincing enough that people DO trust them which is sort of a problem since AI slop is creeping into everything. What can be done to solve it (while not perfect) is pretty powerful. You can force feed them the facts (RAG) and then verify the result. Which is way better than trusting LLMs while doing neither of those things (which is what a lot of people do today anyway). See the recent 5 cases of lawyers getting in trouble for ChatGPT hallucinating citations of case law. LLMs write better than most college students so if you do those two things (RAG + check) you can get college graduate level writing with accurate facts... and that unlocks a bit of value out in the world. Don't take my word for it look at the proposed valuations of AI companies. Clearly investors think there's something there. The good news is that it hasn't been solved yet so if someone wants to solve it there might be money on the table.
- wk_end 2y agoHow do you check it? Take the example of case law. Would you need to formalize the entirety of case law? Would the AI then need to produce a formal proof of its argument, so that you can ascertain that its citations are valid? How do you know that the formal proof corresponds to whatever longform writing you ask the AI to generate? Is this really something that LLMs are suited for? That the law is suited for?
- latexr 2y ago> and that unlocks a bit of value out in the world. > Don't take my word for it look at the proposed valuations of AI companies. Clearly investors think there's something there. Investors back whatever they think will make them money. They couldn’t give less of a crap if something is valuable to the world, or works well, of is in any way positive to others. All they care is if they can profit from it and they’ll chase every idea in that pursuit. Source: all of modern history. https://www.sydney.edu.au/news-opinion/news/2024/05/02/how-corporations-harm-your-health-through-everyday-products---a-.html https://www.sydney.edu.au/news-opinion/news/2024/05/02/how-c... https://www.decof.com/documents/dangerous-products.pdf https://www.decof.com/documents/dangerous-products.pdf
- Terr_ 2y ago> Investors back whatever they think will make them money. A not-flagrantly-illegal example of this might be casinos, where IMO it is basically impossible to argue the fleeting entertainment they offer offsets the financial ruin inflicted on certain vulnerable types of patron. > All they care is if they can profit from it Notably that isn't the same as the business itself being profitable: Some investors may be hoping they can dump their stake at a higher price onto a Greater Fool [0] and exit before the collapse. [0] https://en.wikipedia.org/wiki/Greater_fool_theory https://en.wikipedia.org/wiki/Greater_fool_theory
- Gormo 2y ago> They couldn’t give less of a crap if something is valuable to the world "The world" is an abstraction: concretely, every bit of value that is generated within that abstraction accrues to someone in particular -- investors in AI projects, for example.
- Nehnehneh 2y agoThat's just not true. The training data is the underlying truth and that's not nothing but a lot. And hallucinations are pathes inside this space which are there for yet unknown reason. We like answers from LLMs which walk through this space reasonable.
- mrguyorama 2y ago>The training data is the underlying truth Correct. What is the training data? Language in the form of sentences and documents and words and "tokens". No human language has any normal or natural encoding of "fact" or "truthiness" which is the entire point. You can only rarely evaluate a string of text for truthiness without external context. An LLM "knows" the structure and look of valid text. That's why they rarely produce grammar mistakes, even when "hallucinating". A lie, a made up reference, a physical impossibility, contradictions, etc are all "valid sentences". That's why you can never prevent an LLM from producing falsehoods, lies, contradictions etc. Truthiness cannot be hacked in after the fact, and I currently believe that LLMs as an architecture are not powerful enough a statistical tool that you even COULD train an LLM that had "truthiness" of the entire corpus labeled somehow, especially since that's on it's own a fairly impossible task.
- TZubiri 2y agoI disagree with this take, Stallman has expressed it recently by linking some "scientific article". While I get that LLMs generate text in some way that does not guarantee correctness. There is a correlation between generated text and correctness, which is why millions of people use it... You can judge the correctness of a sentence generated by an LLM. In the same way you can judge the correctness of a human generated sentence. Now whether the truthness or correlation with reality of an LLM sentence can be judged on its own or whether it requires a human to interpret it is not very relevant, as sentences produced by the LLM are still correct most of the time. Just because it is not perfect doesn't make the correctness in the other cases useless, albeit perhaps less useful of course. This is nothing surprising of a statistical model, it tends to produce true results.
- Terr_ 2y ago> I disagree with this take, Stallman has expressed it recently by linking some "scientific article". I don't know how to parse this. What article did Stallman "link", and what are you saying Stallman "expressed" by linking/using it? > whether the truthness or correlation with reality of an LLM sentence can be judged on its own or whether it requires a human to interpret it is not very relevant It's incredibly relevant. We wouldn't even be having these debates if complex LLM judgements could always be verified without a human checking the logic. > sentences produced by the LLM are still correct most of the time At least half the problem here is that humans are accustomed to using certain cues as an indirect sign of time-investment, attentiveness, intelligence, truth, etc... and now those cues can be cheaply and quickly counterfeited. It breaks all those old correlations faster than we are adapting.
- TZubiri 2y agohttps://stallman.org/chatgpt.html https://stallman.org/chatgpt.html https://link.springer.com/article/10.1007/s10676-024-09775-5 https://link.springer.com/article/10.1007/s10676-024-09775-5
- demaga 2y ago> They're all hallucinations, we just happen to like some of them more than others. I love it! Puts things into perspective.
- mdp2021 2y ago> It suggests a qualitative difference And what is sought is, in a way, a jump to that qualitative difference. (And surely there are «desirable and undesirable operation[s]».) "Add something to the dices so that they can be well predictive".
- Loughla 2y agoI just recently showed a group of college students how and why using AI in school is a bad idea. Telling them it's plagiarism doesn't have an impact, but showing them how it gets even simple things wrong had a HUGE impact. The first problem was a simple numbers problem. It's 2 digit numbers in a series of boxes. You have to add numbers together to make a trail to get from left to right moving only horizontally or vertically. The numbers must add up to 1000 when you get to the exit. For people it takes about 5 minutes to figure out. The AI couldn't get it after all 50 students each spent a full 30 minutes changing the prompt to try to get it done. The AI would just randomly add numbers and either add extra at the end to make 1000, or just say the numbers added to 1000 even if it didn't. The second problem was writing a basic one paragraph essay with one citation. The humans got it done, when with researching for a source, in about 10 minutes. After an additional 30 minutes none of the students could get AI to produce the paragraph without logic or citation errors. It would either make up fake sources, or would just flat out lie about what the sources said. My favorite was a citation related to dairy farming in an essay that was supposed to be about the dangers of smoking tobacco. This isn't necessarily relevant to the article above, but if there are any teachers here, this is something to do with your students to teach them exactly why not to just use AI for their homework.
- ryanmcbride 2y agoMy go-to to show people who don't understand its limitations used to be the old "how many Ms are there in the word 'minimum' or something along those lines, but looks like it's gotten a bit better at that. I just tried it with GPT4o and it gave me the right number, but the wrong placement. In the past it's given it completely wrong: >how many instances of the letter L are in the word parallel The word parallel contains 3 instances of the letter "L": The first "L" appears as the fourth letter. The second "L" appears as the sixth letter. The third "L" appears as the seventh letter.
- lawlessone 2y agoThey probably have a letter counting tool added to it now. that it just knows to call when asked to do this. you ask it the number of letters and it sends those words off to another tool to count instances of L, but they didn't add a placement one so it's still guessing those. edit: corrected some typos and phrasing. Maybe we'll reach a point where the LLM's are just tool calling models and not really giver their own reply.
- threeseed 2y agoMaybe don't make things up in a blog post about LLMs making things up. Because you don't know how to fix it. Only how to mitigate it.
- lolinder 2y ago> While the hallucination problem in LLMs is inevitable [0], they can be significantly reduced... Every article on hallucinations needs to start with this fact until we've hammered that into every "AI Engineer"'s head. Hallucinations are not a bug—they're not a different mode of operation, they're not a logic error. They're not even really a distinct kind of output. What they are is a value judgement we assign to the output of an LLM program. A "hallucination" is just output from an LLM-based workflow that is not fit for purpose. This means that all techniques for managing hallucinations (such as the ones described in TFA, which are good) are better understood as techniques for constraining and validating the probabilistic output of an LLM to ensure fitness for purpose—it's a process of quality control, and it should be approached as such. The trouble is that we software engineers have spent so long working in an artificially deterministic world that we're not used to designing and evaluating probabilistic quality control systems for computer output. [0] They link to this paper: https://arxiv.org/pdf/2401.11817 https://arxiv.org/pdf/2401.11817
- swatcoder 2y ago> The trouble is that we software engineers have spent so long working in an artificially deterministic world that we're not used to designing and evaluating probabilistic quality control systems for computer output. I think that's a mischaracterization and not really accurate. As a trade, we're familiar with probabilistic/non-deterministic components and how to approach them. You were closer when you used quotes around "AI Engineer" -- many of the loudest people involved in generative AI right now have little to no grounding in engineering at all. They aren't used to looking at their work through "fit for purpose" concerns, compromises, efficiency, limits, constraints, etc -- whether that work uses AI or not. The rest of us are variously either working quietly, getting drowned out, or patiently waiting for our respected colleagues-in-engineering to document, demonstrate, and mature these very promising tools for us. Everything else you said is 100% right, though.
- deleted 2y ago[deleted]
- bongodongobob 2y ago
- PLenz 2y agoEverything an LLM returns is an hallucination, it's just that some of those hallucinations line up with reality
- mountainriver 2y agoThe same is true for humans
- mdp2021 2y agoIt's a good thing that, as you state, "humans hold "togetherness" as a "true" value". But in this context, the value is in how much they have pondered to actually see and evaluate what is eventually seen as "true".
- __MatrixMan__ 2y agoThere's room for splitting hairs in there though. Even fiction, for instance, can succeed or fail at being internally consistent, is or is not grammatically correct... Calling everything an AI does a hallucination isn't incorrect, but it reduces the term to meaninglessness. I'm not sure that's most useful thing we can be doing. Atoms are not indivisible, yet we use the term because it works. I anticipate hallucination will be the same.
- Gormo 2y ago> Calling everything an AI does a hallucination isn't incorrect, but it reduces the term to meaninglessness. I don't think it does. In this case, "hallucination" refers to claims generated entirely within a closed system, but which pertain to a reality external to it. That's not meaningless, and makes "hallucinations" distinguishable from claims verified against direct observation of the reality they are meant to represent.
- namaria 2y ago> Atoms are not indivisible They are the smallest unit of a substance that cannot be broken down into smaller units of the same substance. They are, in a sense, indivisible.
- tshadley 2y agoThe article referenced the Oxford semantic entropy study but failed to clarify that the issue greatly simplifies LLM hallucination (making most of the article outdated). When we are not sure of an answer we have two choices: say the first thing that comes to mind (like an LLM), or say "I'm not sure". LLMs aren't easily trained to say "I'm not sure" because that requires additional reasoning and introspection (which is why CoT models do better); hence hallucinations occur when training data is vague. So why not just measure uncertainty in the tokens themselves? Because there are many ways to say the same thing, so a high entropy answer may only reflect uncertainty in synonyms-- many ways to say the same thing. The paper referenced works to eliminate semantic similarity from entropy measurements, leaving much more useful results, proving that hallucination is conceptually a simple problem. https://www.nature.com/articles/s41586-024-07421-0 https://www.nature.com/articles/s41586-024-07421-0
- sfink 2y ago> proving that hallucination is conceptually a simple problem. ...proving that this one particular piece of the hallucination problem may be conceptually simple. FTFY
- tshadley 2y ago> ...proving that this one particular piece of the hallucination problem may be conceptually simple. Everything mentioned in the article boils down to that one particular piece-- non-detected uncertainty. The architecture constraints referenced are all situations that cause uncertainty. Training data gaps of course increase uncertainty. Their solutions are a shotgun blast of heuristics that all focus on reducing uncertainty-- CoT, RAG, fine-tuning, fact-checking -- while somehow avoiding actually measuring uncertainty and using that to eliminate hallucinations!
- tyronehed 2y ago[dead]
- throwawaymaths 2y agoCompletely misses the fact that a big part of the reason why llms hallucinate sp much is because there's a huge innate bias towards producing more tokens over just stopping.
- TZubiri 2y agoThe less tokens produced at inference the lower the quality of the response will be. The process of thinking for an LLM involves the use of words, which is why prompts that ask the LLM to only return the answer will cause lower quality.
- ausbah 2y agodo you know if prompting without regards for length then asking for a summarization of the previous out out works?
- TZubiri 2y agoIt does. I think this was used in a gpt4 version, they called it Chain of Thought.
- throwawaymaths 2y agoWe're not talking about quality, we're talking about accuracy. In general, a model has to learn to positively say "I don't know" instead of "I don't know" being in the negative space of tokens falling into a weak distribution. The softmax selector also normalizes the token logits, so if no options are any good (all next tokens suck) it could pick randomly from a bunch of bad choices, which then locks the model into a continuation based off of that first bad choice.
- TZubiri 2y agoWell I am talking about quality now as it's a tradeoff. You can reduce token output to 0 and achieve 100% accuracy too.
- Mistletoe 2y agoIs there a way to code an LLM to just say "I don't know" when it is uncertain or reaching some sort of edge?
- sfink 2y agoIf it works properly, it would need to say that it doesn't know that it doesn't know, and then where are you? (Short answer is yes, but it only works for a limited set of things, and that set can be expanded with effort but will always remain limited.)
- gerdesj 2y ago"It" does not know when it does not know. A LLM is a funny old beast that basically outputs words one after another based on probabilities. There is no reasoning as we would know it involved. However, I'll tentatively allow that you do get a sort of "emergent behaviour" from them. You do seem to get some form of intelligent output from a prompt but correctness is not built in, nor is any sort of reasoning. The examples around here of how to trip up a LLM are cool. There's: "How many letter "m"s in the word minimum" howler which is probably optimised for by now and hence held up as a counterpoint by a fan. The one about boxes adding up to 1000 will leave a relative of mine for lost for ever but they can still walk and catch a ball, negotiate stairs and recall facts from 50 years ago with clarity. Intelligence is a slippery concept to even define, let alone ask what an artificial one might look like. LLMs are a part of the puzzle and certainly not a solution. You mention the word "edge" and I suppose you might be riffing on how neurons seem to work. LLMs don't have a sort of trigger threshold, they simply output the most likely answers based on their input. If you keep your model tightly ie domain focussed and curate all of the input then you have more chance of avoiding "hallucinations" than if you don't. Trying to cover the entirety of everything is Quixotic nonsense. Garbage in; garbage out.
- TZubiri 2y ago"It" does not know when it does not know. But it does know when it has uncertainty. In the chatgpt api this is logprobs, each generated token has a level of uncertainty, so: "2+2=" The next token is with almost 100% certainty 4. "Today I am feeling" The next token will be very uncertain, it might be "happy", it might be "sad", it might be all sorts of things.
- pfisch 2y agoAnyone who has raised a child knows they hallucinate constantly when they are young because they are just doing probabilistic output of things they have heard people say in similar situations and saying words they don't actually understand. LLMs likely have a similar problem.
- tokioyoyo 2y agoTo my understanding, the reason why companies don't mind the hallucinations is the acceptable error rate for a given system. Let's say something hallucinated 25% of the time, but if that's ok, then it's fine for a certain product. If it only hallucinates 5% of the time, it's good enough for even more products and so on. The market will just choose the LLM appropriately depended on the tolerable error rate.
- cameronh90 2y agoAt scale, you are doing the same thing with humans too. LLMs seem to have an error rate similar to humans for the majority of simple, boring tasks, if not even a bit better since they don't get distracted and start copying and pasting their previous answers. The difference with LLMs is they simply cannot (currently) do the most complex tasks that some humans can, and when they do produce erroneous output, the errors aren't very human-like. We can all understand a cut and paste error so don't hold it against the operator, but making up sources feels like a lie and breeds distrust.
- maeil 2y ago> At scale, you are doing the same thing with humans too. LLMs seem to have an error rate similar to humans for the majority of simple, boring tasks, if not even a bit better since they don't get distracted and start copying and pasting their previous answers. This is the big one missed by the frequent comments on here wondering whether LLMs are a fad, or claiming in their current state they cannot be used to replace humans in non-trivial real-world business workflows. In fact, even 1.5 years ago at the time of GPT 3.5, the technology was already good enough. The yardstick is the peformance of humans in the real world on a specific task. Humans, often tired, having a cold, distracted, going through a divorce. Humans who even when in a great condition make plenty of mistakes. I guess a lot of developers struggle with understanding this because so far when software has replaced humans, it was software that on the face of it (though often not in practice) did not make mistakes if bug-free. But that has been never been necessary for software to replace humans - hence buggy software still succeeding in doing so. Of course, often software even replaces humans when it's worse at a task for cost reasons. They're at the very least competitive, if not better than, doctors at diagnosing illnesses [1]. [1] https://www.nytimes.com/2024/11/17/health/chatgpt-ai-doctors-diagnosis.html https://www.nytimes.com/2024/11/17/health/chatgpt-ai-doctors...
- chefandy 2y agoLots of folks in these conversations fail to distinguish between LLMs as a technology and "AI Chatbots" as commercial question answering services. Whether false information was expected or not matters to LLM product developers, but in the context of a commercial question-answering tool, it's irrelevant. Hallucinations are bugs that creates= time-wasting zero-value output, at best, and downright harmful output at worst. If you're selling people LLM pattern generator output, they should expect a lot of bullshit. If you're selling people answers to questions, they should expect accurate answers to their questions. If paying users are really expected to assume every answer is bullshit and vet it themselves, that should probably move from the little print to the big print because a lot of people clearly don't get it.
- int_19h 2y agoI've been playing with Qwen's QwQ-32b, and watching this thing's chain of thought is really interesting. In particular, it's pretty good at catching its own mistakes, and at the same time, gives off a "feeling" of someone very uncertain about themselves, trying to verify their answer again and again. Which seems to be the main reason why it can correctly solve puzzles that some much larger models fail. You can still see it occasionally hallucinate things in the CoT, but they are usually quickly caught and discarded. The only downsides of this approach is that it requires a lot of tokens before the model can ascertain the correctness of its answer, and also that sometimes it just gives up and concludes that the puzzle is unsolvable (although that second part can be mitigated by adding something like "There is definitely a solution, keep trying until you solve it" to the prompt).
- emil_sorensen 2y agoI find it so interesting that it's possible to develop a "feeling" of a new model.
- prollyjethi 2y agoI am honestly very skeptical of articles like these. Hallucinations are a feature of LLMs. The only ways to "FIX" it is to either stop using LLMs. Or use a super bias some how.
- janalsncm 2y agoYou should be. I don’t know anything about kepa.ai but before even clicking the article I assume they’re trying to sell me something. And “how to fix it” makes me think this is some kind of SEO written for people who think it can be fixed, meaning the article is written for robots and amateurs.
- Der_Einzige 2y agoWow, a whole article that didn't mention the word "sampler" once. There's pretty strong evidence coming out that truncation samplers like min_p and entropix are strictly superior to previous samplers (which everyone uses like top_p) to prevent hallucinations and that LLMs usually "know" when they are "hallucinating" based on their logprobs. https://openreview.net/forum?id=FBkpCyujtS https://openreview.net/forum?id=FBkpCyujtS (min_p sampling, note extremely high review scores) https://github.com/xjdr-alt/entropix https://github.com/xjdr-alt/entropix (Entropix) https://artefact2.github.io/llm-sampling/index.xhtml https://artefact2.github.io/llm-sampling/index.xhtml
- TZubiri 2y agoThe debate around "fixing" hallucinations reminds me of the debate around schizophrenia. https://www.youtube.com/watch?v=nEnklxGAmak https://www.youtube.com/watch?v=nEnklxGAmak It's not a single thing, a specific defect, but rather a failure mode, an absence of cohesive intelligence. Any attempt to fix a non-specific ailment (schizophrenia, death, old age, hallucinations) will run into useless panaceas.
- fsckboy 2y agoit's superficially counterintuitive to people that an AI that will sometimes spit out verbatim copies of written texts, also will just make other things up. It's like "choose one, please". MetaAI makes up stuff reliably. You'd think it would be an ace at baseball stats for example, but "what teams did so-and-so play for", you absolutely must check the results yourself.
- mdp2021 2y ago> "counterintuitive" It is consistent with the topic that the reply would be "Tell them that sequences of words that were verbatim in a past input have high probability, and gaps in sequences compete in probability". Which fixes intuition, as duly. In fact, things are not supposed to reply through intuition, but through vetted intuition (and "vetted mature intuition", in a loop). > you absolutely must check the results yourself So, consistently with the above, things are supposed to reply through a sort of """RAG""" of the vetted (dynamically built through iterations of checks).
- fsckboy 2y agoi said "superficially counterintuitive", you misquoted me and proceeded with a non superficial comment.
- mdp2021 2y ago> misquoted me Why? The reply would have been the same if I quoted the whole «superficially counterintuitive to people that [...]» (instead of just pointing to the original). > proceeded with a non superficial comment Well, hopefully ;) Your post went into the right direction of leading towards the idea that "there is intuition, and there is mature thought further from that: and processors must not stop at intuition, immature thought". (Stochastic output falls in said category of "intuition"... As "bad intuition", since it goes in the wrong direction in the vector "naive to sophisticated".)
- mwkaufma 2y agoHow do we discriminate when a response is correct, vs. when it's "hallucinating" an accurate fact, by coincidence? Are all responses hallucinations, independent of correspondence to ground-truth?
- madiator 2y agoFor the specific form of hallucination, which is called grounded factuality, we have trained a pretty good model that can detect if a claim is supported by a context. This is super useful for RAG. More info at https://bespokelabs.ai/bespoke-minicheck https://bespokelabs.ai/bespoke-minicheck.
- mdaniel 2y agoYour playground pre-populated example isn't doing you any favors, and the "examples" folder linked to on curator's GitHub would be better served by showing areas where your model shines, not "generate a poem" which hardly has any factuality to it. I don't have any earthly idea what camel.py is trying to showcase with respect to your model's capabilities I am open to the fact that maybe the value your service provides is in spitting out a percentage, even if it is - itself - hallucinated. But, hey, it's a metric that can be monitored
- IWeldMelons 2y agoLLM hallucinations in fact has a positive side effect too, if you are using them for learning some subject; makes you verify their claims, and finding errors in them is very rewarding.
- skydhash 2y agoWhy not just read a book where the author is sincerely trying to teach you?
- IWeldMelons 2y agoNot as interactive, not gamified.
- skydhash 2y agoSchool, then. Which is so gamified that it has real stakes. And so much interactions.
- IWeldMelons 2y agoYou are not being serious at this point. OTOH I find that chatting with LLMs helps me in studying a lot, esp. when I fight over subtle hallucinations.
- mdp2021 2y agoWith an interactive LLM you can take a manual and start asking questions (about what you read), also recursively. It is a very efficient way of studying. No, doing it with a professor is not the same - unless you can afford an always available tutor of unthinkable erudition.
- fnordpiglet 2y agoWhen trying to learn a subject I find being able to ask my specific questions and getting a specific answer back is helpful. I find books tend to be laborious and filled with frankly filler, often poorly indexed, and when my question isn’t covered in the book I’m left with no recourse other than googling through SEO wastelands or on topic forum questions with off topic replies. At least with LLMs they always have an answer that’s got enough of the truth in it to give me a direction, or often when I’ve gone into an area with genuinely no known answers or the thing doesn’t exist the answer is easily verified as wrong - but that process, as was pointed out above, teaches me a lot too. I actually prefer the mistakes it makes because it forces me to really learn - even to the point of giving me things to look up in the index of a book. Treating LLMs as a single source of truth and a monolithic resource is as bad an idea as excluding them as a tool in learning.
- uz44100 2y agohallucination problem in LLM, been seeing this. Let me know if someone find a fix please
- imchillyb 2y agoToddlers don't understand truth either, until it's taught. This crayon is red. This crayon is blue. The adult asks: "is this crayon red?" The child responds: "no that crayon is blue." The adult then affirms or corrects the response. This occurs over and over and over until that child understands the difference between red and blue, orange and green, yellow and black etcetera. We then move on to more complex items and comparisons. How could we expect AI to understand these truths without training them to understand?
- mdp2021 2y agoYou probably need to be more clear: the LLM is trained with large amounts of data making statements about facts. It is told repeatedly, "according to this source that crayon is blue".
- dschuetz 2y agoI went straight to the "how to fix" section with popcorn in hand and I wasn't disappointed: just add " doubt" layers for self-correction, beginning at the query itself. And then maybe tell the model "do not hallucinate". Sounds like a pun, but I think an AI model actually would take this seriously, because it can't tell the difference. Context is still a huge problem for AI models, and it's probably still the main reason for hallucinating AIs.
- Sergii001 2y agoThat's all really weird. You can watch how chat gpt gives you advice on which mushrooms are safe. And now it can be just hallucinations
- JimmyWilliams1 2y ago[dead]
- PittleyDunkin 2y agoHumans hallucinate, too. We just have less misleading terms for it. Massive mistake in terms of jargon, IMO—"making shit up" is wildly different from the "delusion of perception" implied by hallucination.
- LetsGetTechnicl 2y agoWhy do LLMs make things up? Because that is all that LLMs do, sometimes what it outputs is correct though.
- jrflowers 2y agoI like that none of the suggestions address probabilistic output generation (aside from the first bullet point of section 3C, which essentially suggests that you just use a search engine instead of a language model). TLDR: Hallucinations are inherent to the whole thing but as humans we can apply bubble gum, bandaids and prayers
- rabid_turtle 2y agoI don't like the output = hallucination I like the output = creative
- deleted 2y ago[deleted]