5 ms·
In your opinion, how does the "hallucination" issue differ from the same behaviour we see in humans? I don't claim or believe that any LLM is actually intellig
by InvertedRhodium 3y ago
In your opinion, how does the "hallucination" issue differ from the same behaviour we see in humans?
I don't claim or believe that any LLM is actually intelligent. It just seems that we (at least on an individual basis) can also meet the criteria outlined above. I know plenty of people who are confidently incorrect and appear unwilling to learn or accept their own limitations, myself included.
In my opinion, even if we did have AGI it would still exhibit a lot of our foibles given that we'd be the only ones teaching it.
- dmarchand90 3y agoYeah I've never gotten this argument at all. "Humans aren't actually intelligent they're just machines designed to optimize their probability of reproducing "
- f4f4f4f43f 3y agoYes, you may be. But you still have an internal world model - through conditioning or otherwise that you're playing off against. An LLM doesn't have that. It's very impressive parlour trick (and of course a lot more), but it's use is hence limited (albeit massive) to that. Chaining and context assists resolving that to some extent, but it's a limited extent. That's the argument anyway, that doesn't mean it's not incredibly impressive, but comparing it to human self-awareness, however small, isn't a fair comparison. It's next token prediction, which is why it does classification so well.
- samuellevy 3y agoSo humans have a level of knowledge, understanding, and reasoning ability that LLMs simply don't have. I'm writing a response to you right now, and I "know" a certain amount of information about the world. That knowledge has limits, and I can expand it, I can forget it, all sorts of things... "Hallucination" is a term that works well for actual intelligence - when you "know" something that isn't true, and has no path of reasoning, you might have hallucinated the base "knowledge". But that doesn't really work for LLMs, because there's no knowledge at all. All they're doing is picking the next most likely token based on the probabilities. If you interrogate something that the training data covers thoroughly, you'll get something that is "correct", and that's to be expected because there's a lot of probabilities pointing to the "next token" being the right one... but as you get to the edge of the training data, the "next token" is less likely to be correct. As a thought experiment, imagine that you're given a book with every possible or likely sequence of coloured circles, triangles, and squares. None of them have meaning to you, they're just colours and shapes that are in random seeming sequences, but there's a frequency to them. "Red circle, blue square, gren triangle" is a much more common sequence than "red circle, blue square, black triangle", so if someone hands you a piece of paper with "red circle, blue square", you can reasonably guess that what they want back is a green triangle. Expand the model a bit more, and you notice that "rc bs gt" is pretty common, but if there's a yellow square a few symbols before with anything in between, then the triangle is usually black. Thus the response to the sequence "red circle, blue square" is usually "green triangle", but "black circle, yellow square, grey circle, red circle, blue square" is modified by the yellow square, and the response is "black triangle"... but you still don't know what any of these things _mean_. When you get to a sequence that isn't covered directly by the training data, you just follow the process with the information that you _do_ have. You get "red triangle, blue square" and while you've not encountered that sequence before, "green" _usually_ comes after "red, blue", and "circle" is _usually_ grouped with "triangle, square", so a reasonable response is "green circle"... but we don't know, we're just guessing based on what we've seen. That's the thing... the process is exactly the same whether the sequence has been seen before or not. You're not _hallucinating_ the green circle, you're just picking based on probabilities. LLMs are doing effectively this, but at massive scale with an unthinkably large dataset as training data. Because there's so much data of _humans talking to other humans_, ChatGPT has a lot of probabilities that make human-sounding responses... It's not an easy concept to get across, but there's a fundamental difference between "knowing a thing and being able to discuss it" and "picking the next token based on the probabilities gleaned from inspecting terabytes of text, without understanding what any single token means"
- deleted 3y ago[deleted]
- hashhar 3y agoWhat you're describing is very close to what the thought experiment of Chinese Room (https://en.wikipedia.org/wiki/Chinese_room https://en.wikipedia.org/wiki/Chinese_room). But yes, it's unfortunate that when the next tokens are joined token and laid out in the form of a sentence it appears "intelligent" to people. However if you instead lay out the individual probabilities of each token instead then it'll be more obvious what ChatGPT/LLMs actually do.
- samuellevy 3y agoYep, the "chinese room" is the classic thought experiment, but I feel like it fails to get the point across because the characters still represent language, so you could conceivably "learn" the language. I prefer the idea of symbols that aren't inherently language, as it really nails in the idea that it doesn't matter how long you spend, there's not something that you can ever learn to "speak" fluently.
- hackinthebochs 3y agoWhat do you think your brain does when deciding the next word to speak? It is scoring words based on the appropriateness considering context and all the relevant known facts, as well as your communicative intent. But it is not obvious that there is nothing like communicative intent in LLMs. When you prompt it, you are engaging some subset of the network relevant to the prompt that induces a generative state disposed to produce a contextually appropriate response. But the properties of this "disposition to contextually appropriate responses" is sensitive to the context. In a Q&A context, the disposition is to produce an acceptable answer, in a therapeutic context, the disposition is to produce a helpful or sensitive response. The point is that communicative intent is within the solution space of text prediction when the training data was produced with communicative intent. We should expect communicative intent to improve the quality of text prediction, and so we cannot rule out that LLMs have recovered something in the ballpark of communicative intent.
- joshuahedlund 3y ago> how does the "hallucination" issue differ from the same behaviour we see in humans? In humans “hallucination” means observing false inputs. In GTP it means creating false outputs. Completely different with massively different connotations.
- samuellevy 3y agoThat's kind of the point, but also kind of not. GPT isn't making true or false outputs. It's just making outputs. The truthiness or falseness of any output is irrelevant because it has no concept of true or false. We're assigning those values to the outputs ourselves, but like... it doesn't know the difference. It's like blaming a die for a high or a low roll - it's just doing rolls. It has no knowledge of a good or a bad roll. GPT is like a Rube Goldberg machine for rolling dice that's _more likely_ to roll the number that you want, but really it's just rolling dice.
- hnfong 3y agoYou could apply the same logic to humans. Whenever a human speaks, it's just vibrations of wave molecules, triggered by the mouth and throat, which in turn are controlled by electric signals in the human's neural network. Those neurons, they just make muscles move. They don't have any concept of true of false. At least nobody has found a "true of false" neuron in the brain.
- parthianshotgun 3y agoall of it coheres to consciousness, we know what it's like to be a human, but I think it'd be hubris to think we've cracked the code and made a blueprint of anything other than a word calculator
- hnfong 3y agoHubris goes both ways. It is also hubris to assume our intelligence is special, instead of a boring neural network with sufficient number of neurons that exhibit emergent properties.
- Quarrelsome 3y ago> In your opinion, how does the "hallucination" issue differ from the same behaviour we see in humans? I feel like if you have any belief in philosophy then LLMs can only be interpreted as a parlour trick (on steroids). Perhaps we are fanciful in believing we are something greater than LLMs but there is the idea that we respond using rhetoric based on trying to find reason within in what we have learned and observed. From my primitive understanding, LLMs rhetoric and reasoning is entirely implied based on an effectively (compared to the limitations of human capacity to store information) infinite amount of knowledge they've consumed. I think if LLMs were equivalent to human thinking then we'd all be a hell of a lot stupider, given our lack of "infinite" knowledge compared to LLMs.
- krageon 3y ago> if you have any belief in philosophy [...] You're going to have to explain which part of philosophy you mean, because what came after this doesn't follow from that premise at all. It's like saying a Chinese Room is fundamentally different from a "real" solution even though nobody can tell the difference. That's not a "belief in philosophy", that's human exceptionalism and perhaps a belief in the soul.
- Quarrelsome 3y agoThe belief that your thoughts are constructed based on an understanding of principles such as logic, rationality, ethics. That your interactions are built from a solid understanding of these ideas. As opposed to every train of thought just being glued together from pertinent fragments you can recall from your knowledge in response to a prompt provided by the circumstances of reality. > that's human exceptionalism and perhaps a belief in the soul. I would also argue that LLMs are not proven to be equivalent to what's going on in our minds. Is it really "human exceptionalism" to state that LLMs are not yet and perhaps never will be what we are? I feel like from their construction it is somewhat evident that there are differences, since we don't raise humans the same way we raise LLMs. In terms of CPU years babies require significantly less time to train.