6 ms·
Yann LeCun has pointed out [0] that due to the autoregressive nature of LLMs, you'll never be able to stop them hallucinating. I think what you need is an arch
by optimalsolver 3y ago
Yann LeCun has pointed out [0] that due to the autoregressive nature of LLMs, you'll never be able to stop them hallucinating.
I think what you need is an architecture that emits a fully built answer directly from memory after working on it for a while, not one that builds it up token by token.
[0] https://www.hopsworks.ai/dictionary/rlhf-reinforcement-learning-from-human-feedback https://www.hopsworks.ai/dictionary/rlhf-reinforcement-learn...
- cm2187 3y agoBut doesn't the human mind make a lot of things up too, just to compensate for the gaps? Like if I tell you: "I have a box, I place 3 balls in it and remove 1, how many balls are left". You have to assume that the box was empty before, that I want to know how many balls are left in the box, not on the side, that the box isn't leaking, etc. This is you making up facts that weren't present in the question.
- moffkalast 3y agoYep, we're the exactly same in this regard. This is why expecting people to do anything consistently right from memory is a recipe for catastrophe. Aerospace knows this very well, so pre-flight checklists are standard.
- jgtrosh 3y agoI feel there's an important difference : we humans can, to some extent, analyse what we just said (or were about to say) (or what we just typed and didn't press send yet!) and conceptualise it as a whole to evaluate its veracity. Such an ability would really elevate current LLMs to a higher level of “intelligence” imo. It is, however, interesting that the “basic” word-spewing ability we now share with LLMs seems vital to the intelligence process. I have no idea how that reckoning phase would be implemented.
- mavhc 3y agoask it to consider what it just wrote
- jebarker 3y agoRight. I think we'll end up seeing models that iterate on larger parts of autoregressive generations before presenting them to the user. It's just more compute, i.e. the bitter lesson.
- TeMPOraL 3y agoMy take for almost a year now is: LLMs are like your inner voice. The part where we can "analyse what we just said (or were about to say) (or what we just typed and didn't press send yet!) and conceptualise it as a whole to evaluate its veracity" involves going back-and-forth in yourself. A loop. LLMs are the part inside the loop, and so if you want to get better results, you have to feed the output back to LLM. This is, arguably, part of what the "conversation mode" does, by always including the growing message log in each request - which is why e.g. GPT-4 is good at correcting a mistake once you say, without pointing it out, that there exist a mistake.
- Jensson 3y ago> which is why e.g. GPT-4 is good at correcting a mistake once you say, without pointing it out, that there exist a mistake. It is also good at adding a mistake, so this is completely unrelated. It just rolls to get another random. What you should add is that ChatGPT does better when it writes out every step, so what you said previously is still true just that your example doesn't properly support it.
- TeMPOraL 3y agoThe "chain of thought" thing where you tell it to "think step by step" is a slightly different thing. That's a single forward pass, where it can only reason based on what it output earlier. But once you reply to it, and the interface feeds the whole conversation back to the LLM, it now can look at the entirety of its previous answer.
- Jensson 3y agoWhen it does it step by step it can look at every step when it writes down its final answer after the steps. There is no extra magic that happens just because you write something in between. Edit: To expand on that, a humans inner thought would have caught the mistake at that point and corrected it by writing more instead of ending the text. The LLM basically never does that unless you tell it to fix the answer.
- cout 3y agoWhat would happen if an existing LLM were trained with the ability to emit a backspace token? (Would that even be possible?)
- jgtrosh 3y agoNo idea, but this makes me wonder how we'd even go about getting training data. And it reminds me one time I built a poc editor[0] which saves deleted content so you can get a similar view to manuscripts in which the author strikes out text. Maybe that would be an interesting starting point… [0]: https://github.com/trosh/nohide https://github.com/trosh/nohide
- vidarh 3y agoNot only that, but we will frequently make up a lot of things directly contradicted by what we're told. Look at any online discussion and you'll usually find people answering different questions to the ones they were asked, for example. I don't think you can have functional intelligence without a willingness to fill in details you believe are missing, and that will not always work well. Consider how much effort we spend reinforcing in children how to recognize and suppress what is fantasy vs. reality, and to make them favour telling the truth. It shouldn't be surprising at all that we get "hallucinations" unless/until we put a massive effort into reinforcing that distinction with LLMs too.
- TerrifiedMouse 3y ago> you'll usually find people answering different questions to the ones they were asked, for example. Most of those people are doing it intentionally though - e.g. politicians when you ask them a question they don't want to answer. > I don't think you can have functional intelligence without a willingness to fill in details you believe are missing I disagree. A big part of intelligence is realising/admitting you don't know. That you are missing information. And you don't just fill it up with BS. You going find that information. > Consider how much effort we spend reinforcing in children how to recognize and suppress what is fantasy vs. reality, and to make them favour telling the truth Do we? I don't think children have problems telling the difference between fantasy and reality. If a child isn't telling the truth, most of the time it's not because they can't tell the difference between truth and falsehood, but because they are deliberately lying to you for one reason or another - e.g. they don't want to get punished.
- TeMPOraL 3y ago> Most of those people are doing it intentionally though - e.g. politicians when you ask them a question they don't want to answer. That's distinct from the failure of communication. I believe GP meant something closer to what we'd call a "brain fart", where you read the text correctly, but "understood" something different. > I don't think children have problems telling the difference between fantasy and reality. If a child isn't telling the truth, most of the time it's not because they can't tell the difference between truth and falsehood Do you have children? I've been saying that GPT-3.5 and GPT-4 failure modes are disturbingly similar to the failure modes of my own kid, who happens to be 4.5 right now. I mean it. It's a realization that hit me when I started noticing her "context window" growing - when she'd go into story mode, telling some fantasy stuff like having a sister giraffe who flies a helicopter or whatnot, I could tell that, if she doesn't mention something again within 30 seconds, it would never be mentioned again - a forgotten detail. That window of time kept growing over subsequent months, and is too large to notice now, but hey, the entire way she'd construct her stories was very similar to what you get when you prompt the LLM for a story and just let it keep writing. Anyway, on the parenting stuff: > it's not because they can't tell the difference between truth and falsehood, but because they are deliberately lying to you for one reason or another - e.g. they don't want to get punished. That's true at a later age. Early on, even at 3 y.o., they get genuinely confused about reality.
- RandomLensman 3y agoYes, but why would we want machines to inherit our "issues"? Machines should at least be able to have ground truth and then build on it. For certain social interactions dialling in some blank filling can then be totally fine. Would you want a calculator that gets regularly gets addition wrong?
- j1elo 3y agoI'm not sure if, in the case of humans, that example falls on making up facts for the logical analysis (as you imply) or in the assumed common context for language processing (the "tool" we socially use to make language efficient: by not having to explain absolutely everything). If you told me that thing about the box, I'd assume you meant a box that doesn't leak, not because of making that fact up in my brain, but because you're verbally describing a system to me, and it's implied you'll tell me the most relevant features of the system in order to transmit the idea of it from your mind to mine. Otherwise I'd be right to complain that you are maliciously hiding important information in your description. Problem is, of course, we don't have a common preestablished language context with an LLM, at least not any more than what is the prevalent one in its training materials ("Ask" vs. "Guess" cultures, etc).
- vidarh 3y agoSo you are making things up. Including making up rules about which things it's ok for you to make up. They might be reasonable things to make up, but you're still inferring them to plug holes in the description based on what seems to be the most probably interpretation. And that is pretty much exactly what an LLM does: It predicts the most likely item of data to fill in. That their training is as of yet insufficient to make them predict reasonable things in all situations should be far less surprising than how quickly we've gotten to a point where we can communicate well enough with them that these conversations are meaningful at all. One of the fascinating things about these discussions to me is that they reveal an astonishing amount of assumptions about human thought processes that appear to be just marginally above LLM hallucinations in rigor.
- j1elo 3y ago> So you are making things up. Including making up rules about which things it's ok for you to make up. I guess yeah, that's correct :) But it's also not an artifact, but within the very definition of human interaction. We need to be able to plug holes. Otherwise conversation evolves into a list of guarding caveats after each and every message. OTOH Twitter is a great example of how different cultures develop different assumptions about what is reasonable to fill the gaps in with. It's notorious for how people just take whatever wrong and unintended meaning they want from controversial tweets. Some times on purpose, granted, but lots of other times not really. I concur it's a very interesting topic!
- tveita 3y agoThat the box starts out empty is an assumption based on shared cultural context, not a hallucination. If you give someone that question and then say "wrong, there was already a ball in the box!" they're not going to say "haha, silly me, I hallucinated", they're going to stop hanging out with you. Lots of separate phenomenon can cause our minds to have "wrong" or incomplete information, calling them all "hallucinations" is reductive and just serves to trivialize the vast differences in operation between human minds and LLMs.
- TeMPOraL 3y ago> That the box starts out empty is an assumption based on shared cultural context, not a hallucination. Sure, in the exact same way LLM operates within context of everything it learned. It so happens that it's "cultural context" is derived from ours, by means of the training data, which is why it's so similar. > If you give someone that question and then say "wrong, there was already a ball in the box!" they're not going to say "haha, silly me, I hallucinated" And yet people are doing that to LLMs. "Haha, look at that LLM hallucinating", where half the time the prompt is misleading or wrong. > they're going to stop hanging out with you. That's the social equivalent of RLHF.
- maximus-decimus 3y ago> And yet people are doing that to LLMs. "Haha, look at that LLM hallucinating", where half the time the prompt is misleading or wrong. I mean... I ask it to give me Python code and it gives me code that doesn't work and makes up libraries that don't actually exist. That's not a prompt problem. A human would understand there's an implicit "code that actually works and isn't pulled out of your ass".
- harimau777 3y agoCould that be similar to how most people have a few bugs in the first draft of their code; particularly when writing from memory (e.g. during a whiteboard interview)? For example, I always have to lookup whether the function is Array.every or Array.some in JavaScript. That doesn't seem too different than making up a library to me.
- rightbyte 3y agoIt is fair to assume a riddle is solvable, which means two balls. "It is unsolvable" is the correct answer but it is a gotcha riddle.
- xnzakg 3y agoFeels like this is where having the LLM "think out loud" helps. I've seen examples where a model makes a mistake and then corrects itself.
- akrymski 3y agoPretty certain you can never stop humans from hallucinating either. Any human answer will have a probability/confidence level attached to it, it's never zero, because any statement relies on assumptions of the world encoded by the same humans. Unless we are actually doing symbolic reasoning like a theorem proof? But even then are we ever completely certain the proof is correct? This doesn't come naturally to us.
- Jensson 3y ago> Any human answer will have a probability/confidence level attached to it, it's never zero I have been to a grocery store at some point in my life. That is 100%. I trust you also have statements you can be certain of. Humans are 100% certain about many things.
- vidarh 3y agoYour claim relies on an assumption that your memory is correct and that your sensory input is trustworthy. While those are highly likely, and so it's reasonable to "round up" and claim to be 100% sure, the probability of it being true is most certainly not 100%.
- RugnirViking 3y agohere's the thing though, we aren't talking about an objective fact. Maybe humans don't exist at all, maybe we're boltzmann brains floating in a nebula. We're talking about self reported probabilities humans are aware of. And yes, I am actually 100% sure of some facts about myself. (I have a sister, I live in an apartment). I am aware, abstractly, that such things may not be true in a sort of plato's cave shadows etc, but I will for all intents and purposes act and believe that they are certain. If such things were proven to not be true, I would be shocked to my very core; I am shocked because it broke a notion I had, if I was aware in my actual core of the possibility that it may not be true, I wouldn't be shocked. To survive, our biological machinery must make assumptions and round probabilities. To do otherwise would be to be paralysed with indecision. (what if there is some small chance an elaborate deathtrap has been sprung that will kill me if I move even slightly at this very moment? what about the next moment?)
- pfd1986 3y agoI may be wrong but isn't that what chain of thought, verification etc claim to solve? As I understand them, you essentially feed back the llm back as a prompt and ask to verify if it's true and fix mistakes until "convergence"