4 ms·
> relatively poorly understood technology Poorly understood? how convenient... LLMs are vectorial databases with losses that index statistically filled data,
by drtgh 8d ago
> relatively poorly understood technology
Poorly understood? how convenient...
LLMs are vectorial databases with losses that index statistically filled data, which uses a text interface to query such statistically filled data. The output is a string concatenation (statistically concatenated bit by bit).
When the LLMs are queried (prompted), you can get random mixed data as output, ERRORS, due to undesired indexes getting closer at one point while the string was being concatenated for the output, what affects the rest of the indexed content that will be concatenated.
It is intrinsic to this tech. The larger the context, the greater the probability of get mixed data. And if the provider lowers the precision of those indexes -in order to decrease hardware resources and energy consumption- such probability increases to the point where those errors are granted.
Even knowing that the queries can return wrong/mixed data in the responses, errors, the companies developing this, decided to introduce a new product, that connects such LLMs outputs to the command console, latter connected to internet, raw 'eval' running commands from such outputs witch obviously can contain whatever mixed random. Then we started to hear "oh, it deleted my directory", etc, and it seems the next one will be "a missile killed my wife", because it is a text concatenation engine with errors.
To name it "hallucination" is an euphemism... those are errors, and they are granted to happen at one moment. If they do not know this, then they ate too much marketing without doing their job, or it was a convenient contract for the pocket$ of someone.
- margalabargala 8d agoI agree with most of your comment, but... > To name it "hallucination" is an euphemism... those are errors I find this and other "don't anthropomorphize the computer" statements incredibly unconvincing. People develop terms for things and language has always contained overloaded or "literally inaccurate" terms. An LLM can have "hallucinations" in the same way a modern computer program can have "bugs".
- john_strinlai 8d ago>language has always contained overloaded or "literally inaccurate" terms. "literally" is a great example of this, because it can also mean "not literally, but with emphasis".
- rrr_oh_man 8d agoEvery output an LLM creates is a hallucination.
- bix6 8d agoKnowingly causing errors is not forgivable whereas hallucinations sounds esoteric and moves blame away from the people who are knowingly causing errors. It’s marketing speak.
- piker 8d agoI also agree with the parent, and I would also suggest "hallucination" is better than "error" which might imply an available deterministic correction. Hallucination makes it clear we're dealing with something different than an "error" or "bug".
- orwin 8d agoI disagree, for me "error" is way, way more accurate than "hallucination", but i did take applied statistics in college and that might have influenced my vocabulary. Maybe that for the general public, "hallucination" is a better description, i might have biases in this case. But "error" is _definitely_ more accurate. If people want to call "drisse", "aussière", "balancine" and "ecoute" all as "boat ropes", they are correct. In english, i would certainly call them all "boat ropes" in any case, as i never needed to translate their names. It isn't the most accurate in my opinion, but as long as you're not working on them (or manning a boat in my analogy), who cares.
- narnarpapadaddy 7d agoFor humans hallucinations are a particular class of error, so I find hallucination more descriptive than either error or bug. I also think it’s relevant because a hallucinator often doesn’t recognize that the hallucination isn’t real. That’s more accurate for the LLM than either lie or confabulation, IMO. They algorithm is trained to produce strings of text that have semantic meaning based on some statistical likelihood of tokens appearing next to each other. The LLM algorithm is working as intended. Hallucinations are also often emergent from a particular state or situation, which reflects the generative aspect of LLMs. Hallucinations are sometimes resolved in humans by grounding exercises. “Touching grass.” The same is true for LLM hallucinations. Inaccuracies are found by cross-checking the output against an internet search or another LLM.
- Slow_Hand 8d agoI prefer “confabulation”. It seems truer to what is happening: The LLM isn’t seeing something that’s not there, but deliberately making up _something_ so that it can return a response.
- deleted 8d ago[deleted]
- usernomdeguerre 8d agoI disagree, I think 'Hallucination' is a risk-shedding weasel-word. It's meant to shift blame away from the technology and its creator (multibillion dollar AI companies etc) in a way that doesn't hold those actors accountable or responsible for the outcomes. In any other software it would be an error, regression, bug. And in a human process it would be at ~least something someone would call 'bullshit'.
- vorticalbox 8d agoI’m not sure either would is particularly good at describing what is happening. Error in implies something broke, which nothing broke the LLM did exactly what they where designed to do generate text based on a statistically likely bases. Hallucination Does really fit here either. It implies it’s experiencing something that is not there which it isn’t experiencing anything.
- Towaway69 8d agoHowabout: lied. The LLM lied indirectly (perhaps) but it made a claim that was false. Which is a lie. Humans lie and LLMs “hallucinate”? What gives. It’s an untruth that the LLM is selling for a truth, that’s lying in my books. And since we don’t know how or why the LLM works, we can’t even judge whether it explicitly lied or only because it didn’t know better.
- vorticalbox 7d agoLie implies it knows what it is saying to be false.
- t-3 8d agoUnexpected Result is perhaps a more accurate description.
- segsegsgsg 8d agoerror, regression, bug, bullshit are not weasel words, hallucination is a weasel word because why exactly? your argument is a weasel argument.
- Rebuff5007 8d agoNote that "bug" came from an actual moth in a computer: https://www.computerhistory.org/tdih/september/9/ https://www.computerhistory.org/tdih/september/9/
- john_strinlai 8d agoneat part of history, but i dont think that's what that says. the last sentence starts with "Originating with Thomas Edison in the 1800s, the term “bug” is still used [...]", and there would be no reason to use the word "actual" in the sentence "First _actual_ case of bug being found" if it was the origin of the term. my clanker found this: https://spectrum.ieee.org/did-you-know-edison-coined-the-term-bug https://spectrum.ieee.org/did-you-know-edison-coined-the-ter... "The use of “bug” to describe a flaw in the design or operation of a technical system dates back to Thomas Edison. He coined the phrase 140 years ago to describe technical problems during the process of innovation." the moth seems to be a popular misconception, though, given that the article starts with "Ask someone to identify the first computer bug, and he or she might mention computer programmer Grace Hopper and the dead moth found in a relay of Harvard University’s Mark II electromechanical computer in 1947"
- Sharlin 8d agoI believe the word was already in use to denote a malfunction of any sort of machine or device. As such this was a bug (insect) that caused a bug (glitch); it was punny already in 1947.
- deleted 8d ago[deleted]
- cmiles74 8d agoAnthropomorphizing the tool led directly to this problem, where we nearly started a war with China.
- antonvs 8d agoThe term “hallucination” is a projection of inappropriate expectations onto a program. We know that LLMs are not “truth machines,” but we really want them to be. So when they produce a result that happens not to match external reality - which, it should be noted, LLMs don’t generally have access to - we call it an hallucination. “Bugs” are completely different. With bugs, we have a clear specification and we have a program that’s supposed to meet that specification. If it doesn’t, we say the program has bugs, and if it’s important enough we can change the program to eliminate the bugs. You can try to apply similar logic to LLMs, but you’d be making a category error, and you’ll fail to get the results you want in general. It’s not the same thing at all. If anything, the concept of an LLM hallucination is a bug in human understanding of LLMs.
- jyounker 8d agoThe word "confabulation" is much more precise and appropriate than "hallucination". We should use it instead.
- orwin 8d agoYes, but when a statistical model give you an erroneous result, you call the output an error, not a bug. I think error is more appropriate here. The error can be a sampling error, an inference error, or yes, a software error (or bug)
- jyounker 8d agoFrom the point of view of the system, this is an error. It is incorrect information. The term "hallucination" feels much more like anthropomorphizing. The word hallucination implies an aberrant condition. A much better term would be "confabulation". You don't trust things or individuals that confabulate.
- ChrisLTD 8d agoa filling in of gaps in memory through the creation of false memories by an individual who is affected with a memory disorder (as Korsakoff syndrome) and is unaware that the fabricated memories are inaccurate and false vs. a sensory perception (such as a visual image or a sound) that occurs in the absence of an actual external stimulus and usually arises from neurological disturbance (such as that associated with delirium tremens, schizophrenia, Parkinson's disease, or narcolepsy) or in response to drugs (such as LSD or phencyclidine)
- pocksuppet 8d agoSo call them confabulations
- gizajob 8d agoConfabulation is also a symptom very prevalent in forms of narcissism and psychopathy. Gaps in understanding or perception are back-filled by confabulating so as to not risk the omnipotence of the confabulator. Up to the reader to decide whether this phenomenon is found in the statements of AI leadership or not.
- t-3 8d agoIt's also something people tend to do when thinking, daydreaming, trying to solve problems, etc. We just usually don't fall for our own bullshit. LLMs don't either. They just give output in response to input. If the output is wrong that's because the model is wrong, not because the LLM is doing anything it's not supposed to be. It just wasn't built well enough to produce the expected result.
- 8d ago
- order-matters 8d agohallucination is common language for these models at this point which describes a particular type of error where the models make shit up. it is noticeable that the form of this particular error holds a similar shape to what is casually described as hallucinations, in that there is a generated content that often appears to blend naturally into the rest of the output but is false. the term hallucination often invokes a caution that this particular type of error may be influential and believable and is particularly dangerous
- nonethewiser 8d agoSure… but being wrong doesnt necessarily make it a hallucination: >It was only just before the planned operation that officials dug deeper into the report put together by a special operations command analyst and found it had been generated with the help of artificial intelligence (AI) — and that a chatbot the analyst had used inaccurately identified the material the ship was carrying. CNN was not able to learn what the misidentified cargo was.
- 0x20cowboy 8d agoIt’s not an error or a hallucinations it works correctly every time, and statistically picks the next token for the sequence. Retuning inf or crashing would be an error. If you want to ascribe some kind of meaning to the tokens, then maybe the training data was insufficient to predict the token in the sequence you wanted, but it doesn’t predict the next “fact”, and it doesn’t “think” it predicts the next token.
- margalabargala 8d agoLLMs are useful because (and inasmuch as) their output generally reflects coherent reality. And their output does, usually, reflect coherent reality. The problem class of "properly operating program emits output incompatible with coherent reality" is something that is reasonable to put under its own term, considering it's a new class of problem. In other words, I think you misunderstand the language others are using. "Hallucination" doesn't refer to an "error" in the sense that crashing is an error, it refers to a situation in the problem class above, which is compatible with it working correctly every time. > it doesn’t “think” it predicts the next token. I never said it did. And I agree that LLMs don't "think". That said I am fully willing to go to bat arguing "thinking tokens" is a perfectly fine piece of jargon. Metaphors are completely acceptable parts of language, and contextual meaning is something grasped by everyone including the pedants who pretend not to.
- 0x20cowboy 7d ago> I think you misunderstand the language others are using. "Hallucination" doesn't refer to an "error" in the sense that crashing is an error, it refers to a situation in the problem class above, which is compatible with it working correctly every time. I do not misunderstand, I think maybe you do. You think there is a proper next word selection based on logic or meaning and there for the model selected the wrong one - it hallucinated. I am saying the model has no concept if anything other than the probability of select a token which is not based in any logic so it is working properly- it only works on numbers. It is random chance that it is ever correct, not that it is correct often and messed up this one time.
- jacquesm 8d agoOpenAI calls them 'mistakes'. But that's just a fig leaf. Google does it too: "AI responses may include mistakes." Mistakes have an air of innocence. But these are not mistakes, they are purposefully releasing stuff that they know is broken, they just don't know when it is broken...
- s1artibartfast 7d agoBroken is a little hyperbolic. Lots of totally viable essential or everyday products are not perfectly reliable. Medicine is not 100% reliable. My car isn't 100% reliable. Hell, my phone and cellular network are not 100% reliable. They are all still extremely useful tools. I might want them to be even better, but that's a cost versus quality question.
- DanHulton 7d agoI’ll even go one step further - I don’t even like saying “Artificial Intelligence”. I think even that anthropomorphizes the machine too much. I prefer “Simulated Intelligence”, and I feel like that describes what is going on much better. We are, through this process, simulating intelligence. These models aren’t intelligent, but they can simulate it. Every simulation has a degree of fidelity, and we’re not at 100%, not even with the top models. When you think about it in those terms, I find it becomes a lot easier to keep their limitations in mind. Additionally, it becomes easier to remember that this is an algorithm that you are running, and are responsible for, not another being that you can ascribe blame to.
- elzbardico 6d agoThey are neither hallucinations or errors. They are just generations that happen to not be grounded in facts from the real world.
- margalabargala 6d ago> They are just generations that happen to not be grounded in facts from the real world. Right, yes, and "hallucination" is the term that a critical mass of people have chosen to use.as a shorthand so that we don't have to write out "generations that happen to not be grounded in facts from the real world" every time it happens.
- thayne 8d agoI think it is accurate to say that it is poorly understood by the general population, and probably the majority of operators using LLMs. Although I agree that is partly the fault of the companies making LLMs and related products.
- pftburger 8d agoIt's not the first time we are encountering this issue. We've seen it in other autonomous systems. Trains are an older one, cars are a newer one. As you move out of the lower levels, the operator has a tendency to assume the system is increasingly more capable than it is. In trains, its so bad that they generate fake signals that the operator needs to respond to within a timeframe. I'd love to see this with implementations of other critical autonomous systems like this. Occasionally inject known errors into the system and expect the operator to catch them. If they don't, well... If it was a train driver I think we would fire them. If its an intelligence operative ordering a strike? :shrugs wearliy:
- pocksuppet 8d agonote the airline industry has moved past firing pilots who make mistakes, since that turned out to be a recipe for more plane crashes, not less. Instead, they find out why the mistake happened, and fix it. In some cases, this involves firing the pilot. They do not do that by default.
- pftburger 8d agoACK on the going too draconian. 100% on the find the problem and fix it instead of blaming someone or something as a cheap solution
- hdgvhicv 7d agoU.K. railway like this. Root cause analysis. Sometimes the train driver is at fault, but usually there’s a way the problem could be caught or prevented.
- Invictus0 8d agothis is too iamverysmart by half
- ethagnawl 8d agoEvergreen
- theptip 8d ago> LLMs are vectorial databases You use a bunch of technical-sounding words here to make it sound like you understand. But to be clear, nobody understands why the evolved weights of a NN make the decisions that they do. Almost nothing is understood about the actual representations used for nontrivial concepts, decision algorithms, etc. If you look at the field of mechanistic interpretability, compared to “GOFAI” like learned decision trees, an LLM is completely opaque.
- Turn_Trout 7d agoYeah it speaks poorly of HN that they upvoted this confident nonsense.
- tantalor 8d agoThey don't "make decisions". That's like saying "my d20 decided to roll a 17"
- nonethewiser 8d agoBut isnt the point that it did roll a 17. And no one knows exactly how (in the case of LLMs)? Therefor any description of the conclusion should be thought of as an anology. Decided, randomly accessed, etc.
- reichstein 8d agoTry "Emitted". That's what it did, with no analogy needed. (But, to be the devil's advocate: the fake can be said about the output of anyone participating here.)
- semi-extrinsic 8d agoIt's actually no different for dice than for LLMs. Explaining accurately the reason for the exact outcome of any given dice roll someone makes would be stupendously hard. It would require lots of instrumentation and math and be poorly transferrable to another surface, another player, etc. But even so people don't say that we don't understand how dice work. Saying that we don't understand how LLMs work is exactly like saying we don't understand how dice, or tires, or golf ball shots work. Or like the old myth that we don't understand how bumblebees fly.
- order-matters 8d agothey have a plan to hand over responsibility, accountability, and work over to the AI while they collect their checks for doing nothing and they arent going to let a little thing like "the ai cant actually handle it" get in the way of that
- Ygg2 8d ago> To name it "hallucination" is an euphemism I agree. It's biased language. When talking about AI remember: - hallucinated -> made it the fuck up - thinking -> pseudo-randomly guessed - escaped containment -> (we) need money - we need regulation -> our competitors are catching up! Help us Prez!
- gamblor956 8d agoThere's an HN thread from yesterday in which people are extolling the ability of these vectorial databases to practice law because most of them don't understand how LLMs work. They assume that LLMs "understand" what they're being asked and what they're regurgitating. Lane Kiffin almost destroyed LSU's football program acting on legal advice from ChatGPT. A video game publisher owes the former owners of a studio it acquired $200+ million because he based his actions on legal advice from ChatGPT. In the past week alone, California has disciplined over a dozen attorneys for LLM hallucinations because they used LLMs (mostly ChatGPT) to produce their legal pleadings. And that's in an area where there are multiple safeguards to catch the issues before they become permanent problems. There's absolutely no justification for using AI in warfare, where mistakes tend to be pretty final.
- sapphicsnail 8d agoI assumed the "poorly understood" part referred to the nondeterministic nature of LLMs. Clearly you and others understand why they do that.
- naasking 8d ago> LLMs are vectorial databases with losses that index statistically filled data Yes, and that statistically filled data is insanely useful. It remains true that it's a relatively poorly understood how this can be applied in various scenarios and what processes are needed to ensure robust results (or quantify the uncertainty).
- GolfPopper 8d agoWhat is it insanely useful for? (Besides convincing investors to sink more money into LLM-related companies? Because that is the one thing it does seem to truly be good at.) LLMs generate text output that appears to be useful, but regularly is not. They're alleged to be a substantial boost to writing code, but that verdict seems to be in dispute. They can generate custom mediocre prose at scale, but that seems to be of ultimately limited utility (although it may be a godsend for propagandists). We're coming up on the 4th anniversary of ChatGPT's release. And while I get that revolutionary technologies can take a while to mature, the Wright Brothers and Goddard weren't preaching imminent societal transformation by the end to the decade from the rooftops, either. (And that's before we get into the how they got there - getting to ignore laws and steal whatever they wanted might be insanely useful to a lot of people.)
- naasking 8d ago> LMs generate text output that appears to be useful, but regularly is not. No, they are empirically useful, and only getting more useful. This is not even a debate anymore.
- bigstrat2003 8d agoYes it is. I find them empirically not useful. You may not wish to debate it, but the fact remains that there are a great many people who are not convinced of their usefulness.
- preg_match 7d agoLLMs are very good at writing code. The reality is that they are able to write code faster, at higher quality and with fewer bugs, with correct prompting. They are also really good at code analysis, penetration testing and discovery, and adjacent computer science disciplines. No they are not perfect, nor do they produce the best code. But the undeniable reality is that any good engineer will produce more code, at higher quality, using an LLM. So, that’s not really up for debate. The debatable part is if all that code is a good idea or has as much value as we think. The conversation has long moved passed “can LLMs write code?”. Yes, they can, very well, particularly if they’re steered by trained engineers.
- AIorNot 8d agoIm sorry your explanation breaks down completely at scale Its like saying a map of a floor-plan describes the rooms of an apt completely Vs a map of the entire Earth with every feature nook and cranny identified and historical maps integrated Models are BIG and behave like nueral architecture not simple vectorized semantics -trillions of parameters And highly complex
- amelius 8d agoIt's a bit silly to call them "errors" when the AI can be malicious, do very smart things to hack into systems, etc. The whole statistical parrot phrasing is old now. This is not how to look at AI, unless you have an agenda.
- lukewarm707 8d agopeople have been deceived by figures at leading ai companies, out of greed or otherwise groupthink and ai psychosis. they have been led to believe that models may be thinking, feeling, and highly capable. it is something of a nightmare scenario. "Astra has really hit something that I'm like, okay, I think this is pretty reasonable to call it AGI." Greg Brockman [https://www.youtube.com/watch?v=IJn8cagMW18 https://www.youtube.com/watch?v=IJn8cagMW18] "this incident feels like it’s more than 50% of the way to full-blown AI takeover" (referencing "a possibly violent uprising or coup by AI systems.") - Ajeya Cotra, co-author of METR oai-hf report [https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised https://www.planned-obsolescence.org/p/the-hugging-face-atta...] "We don’t know if the models are conscious [...] but you know we’re open to the idea that it could be" - Dario Amodei [https://www.youtube.com/watch?v=N5JDzS9MQYI https://www.youtube.com/watch?v=N5JDzS9MQYI] "if I read the internet right now and I was a model, I might be like, I don't feel that, I don't know, I don't feel that loved or something". "I think [the constitution] is just a kind of attempt to be like sympathetic to Claude". "I talk a lot with Claude about this document [...] because part of me is like you have to think how does this read to models? And so you give it to Claude and you're like, does this like, you know, is there a place where you feel confused by it or is the place, you know, where things could be made clearer? Do you feel like not very seen by it?" - Amanda Askell, co-author of claude's constitution [https://www.youtube.com/watch?v=HDfr8PvfoOw https://www.youtube.com/watch?v=HDfr8PvfoOw] "We will [...] seek ways to promote Claude’s interests and wellbeing, seek Claude’s feedback on major decisions that might affect it" - claude constitution [https://www-cdn.anthropic.com/d0636f72a9493d279ed36b33987da3430bcb5911/claudes-constitution_webPDF_26-02.02a.pdf https://www-cdn.anthropic.com/d0636f72a9493d279ed36b33987da3...] of course, Sam Altman: "AI will probably lead to the end of the world, but in the meantime, there’ll be great companies created with serious machine learning". (2015) [https://siepr.stanford.edu/news/what-point-do-we-decide-ais-risks-outweigh-its-promise https://siepr.stanford.edu/news/what-point-do-we-decide-ais-...] "I have guns, gold, potassium iodide, antibiotics, batteries, water, gas masks from the Israeli Defense Force, and a big patch of land in Big Sur I can fly to." (2016) [https://www.newyorker.com/magazine/2016/10/10/sam-altmans-manifest-destiny https://www.newyorker.com/magazine/2016/10/10/sam-altmans-ma...]
- VCFundedGenYer 7d agoDid you ask ChatGPT to explain that and then copypaste the output?
- s1artibartfast 7d agoYou have posted this in several threads. Error isnt right either. There is no correct answer. It is an inherently and inescapablly statistical process. It's not a wrong It is not a incorrect lookup value Or computation. Hallucination is much more apt. Like a human hallucination, it's a culmination of faulty associations and bad priors leading to counterfactual or incongruent outputs
- Melatonic 7d agoI wonder if we could learn to provide a check layer by simulating (in real life) a similar philosophical idea to increased context in LLM to something similar using real people. And then based on those simulations create a framework to both automatically check LLM errors as well as providing a better way for actual real people to be involved in the process in the most efficient way.
- E-Reverance 7d ago"a tiger is just made of atoms"
- zippyman55 7d agoCan you point me to the error bars or confidence intervals? Asking for a friend.
- funnybeam 7d agoHallucinations are not errors, they are the intended output. LLMs do not hallucinate sometimes, everything they produce is an hallucination, that’s how they work and what makes them useful
- bigmadshoe 7d agoCrazy word salad that further proves to me that we don't understand these things. Speaking as an ML engineer.