7 ms·
I see remarks like this a lot, and I don't know what to do with them. Talking about "real understanding" without fleshing out what we mean by the phrase doesn'
by caturopath 3y ago
I see remarks like this a lot, and I don't know what to do with them.
Talking about "real understanding" without fleshing out what we mean by the phrase doesn't strike me as at all insightful in this context. In day-to-day use, the term is imprecise but usually easy enough to talk about, but I'd struggle to apply it to dogs, let alone to LLMs or AGIs or whatever.
You say that _of course_ they cannot self-correct because of this. If Deepmind's tests showed the opposite result, would you say that the system they were testing did have "real understanding"? It seems likely to me that future systems will demonstrate self-correction using Huang et al.'s methodology...will we think those systems are different in some way involving understanding?
> The LLM I was using gave me code that worked, but did not do what I wanted. I pointed this out, so I then got code that syntactically did what I wanted, but could never work. This went back and forth twice, before I gave up.
> The LLM was incapable of recognizing the falsity of its answers. It certainly wasn't capable of suggesting a different approach that would work (and yes, there was one, I found after a bit more research).
I really don't know what to do about these sorts of anecdotes. Often with slightly different random values or temperature or phrasing, the same LLM will give different results for the same problem. I've repeatedly seen people make much more narrow claims about what LLMs get wrong or what their biases are, where simply regenerating will give a dramatically different answer.
Your pain here sounds like something I see in humans all the time. If we trained a human harder never to ask follow-up questions, I imagine it would be even worse. There are clearly important differences between an LLM and a human being, but I don't think it's always easy to describe what they are in useful ways.
- tnecniv 3y agoTypically, when people say it lacks real understanding, they seem to mean it lacks an understanding of any object’s semantic meaning. In the above example, the person asked for an example of X API solving problem Y, and X lacked the functionality needed to solve Y. If the AI understood the formal specifications of the API, it would be able to know that functionality doesn’t exist and reply as such. However, the AI is just a next token predictor here, it only knows certain words tend to follow others. Thus if you never told it that X API can’t do Y in training, it’s unlikely to reply as such. Instead it might give you a fictional function call in the style of X API to solve your query.
- RandomLensman 3y agoAn LLMs can explain a concept but is unable to apply said concept to a given problem - for me that shows a lack of "real understanding".
- all2 3y ago> I see remarks like this a lot, and I don't know what to do with them. There's no "reasoning loop" built into LLMs yet. Keyword, yet. For now we're left with single-shot answers from "memory" rather than a reasoning loop akin to what a human would do, which is read the docs, try some stuff out, discover that what you're asking for is impossible, and then telling you that it isn't possible.
- jaidhyani 3y agoAlternatively, the prior on "this is not possible" is very low because RLHF & Friends have targeted metrics that, inadvertently or not, discourage that outcome.
- robertlagrant 3y agoI think that's the right answer - human trainers prefer an answer, even a made up one, to "I don't know".
- Jensson 3y agoDataset as well. In a forum if you don't know the answer you simply don't post. Only people who think they know will post an answer. In a dialogue you see a lot more "I don't know" since there they are expected to respond, but there isn't a lot of dialogue data to be found on the internet compared to open forum data.
- SAI_Peregrinus 3y agoAmazon product Q&A has a lot of "I don't know" answers. Unlike just about everywhere else on the internet.
- abm53 3y agoThese models produce correct answers to many problems that require “reasoning” (for any sensible meaning of the word), that are not in their training set.
- YeGoblynQueenne 3y ago>> I've repeatedly seen people make much more narrow claims about what LLMs get wrong or what their biases are, where simply regenerating will give a dramatically different answer. It will, but what's the point of a random answer generator? What's useful is a system that gives correct answers, and can discriminate correct answers from incorrect answers. Otherwise, you can only know that an answer is correct when you already know it, and that's not very useful at all.
- LASR 3y agoAssuming “correctness” as being a hard criterion will greatly limit your ability to get powerful and very useful output from LLMs. A person doing some knowledge work will fail on this criterion. But they still are employable and generally regarded as being useful.
- caturopath 3y agoI think usefulness is a different question than capabilities. "LLMs can't do X" is an interesting claim, and it takes some level of systematic work to make it. If the capability exists at all, this is interesting, even if it's not useful. That being said, it doesn't mean such an existing capability will not prove useful. If something is fallible, it might be above a threshold worth using, it might be improved by doing the same techniques that are already being used, or it might be improved by adding on some sort of system to filter or refine responses. > you can only know that an answer is correct when you already know it, and that's not very useful at all. This statement makes sense on the surface, but I don't think it's true. All the time I get told things that I don't know the answer to beforehand, but judge pretty well whether it's right or not.
- layer8 3y agoYou are right that “understanding” requires further explanation. In the anecdote above, however, it seems to be clear that the LLM didn’t realize a number of things, such as: – It might not have sufficient information to answer the question. – It’s just making a guess of what a correct answer might plausibly look like. – The fact that this is different from actual correctness. – That by asking questions back, it might actually be able to come up with a working answer, with the help of its collocutor. Even after being informed multiple times that it made an error, it doesn’t seem to realize any of the above. This apparent lack of self-reflection, of addressing and working with the conversational situation — i.e., what a human would typically do — is what people usually mean by LLMs lacking an understanding of what they output.
- godelski 3y agoTo add to this list of things LLMs don't understand: - Numbers (LLMs still can't do math and the ones that "can" are very brute forced and don't generalize) - Commutative properties (The (idk why anyone was surprised) "reversal curse": https://arxiv.org/abs/2309.12288 https://arxiv.org/abs/2309.12288) - Associative properties - World models (i.e. can understand some basic physics and causality. Not mathematically, but the same way a child understands that letting go of something in their hand falls to the ground and not up) - How to say "I don't know" - Any form of meaningful generalization (GPT 3.5 and most LLMs still have difficulties with "Which weighs more, a pound of feathers or a kilogram of bricks". GPT4 gets it, but I'm not convinced it isn't because I've asked it that question too many times and they specifically trained on it.) And a whole load of things. With basically all these things being quite relevant to the collective interpretation of "understanding." They aren't hard to tease out either. You ever have a problem that isn't easy to google because the search terms are too close to a different problem? The LLM will give you very similar results, even if you probe with corrections and specify that your problem is different. Like layer8, many many times I've corrected (many different) LLMs and cannot tease out the correct answer. Because the correct answer is too low likelihood and often the RLHF has decreased the distribution in that region to get to it. Because this is exactly what RLHF does, on purpose.
- famouswaffles 3y ago
- godelski 3y agoLet me give an example[0] > Before LLMs, we had a very crisp test for having, or not having a world model: The former could answer certain questions that the latter could not, as shown in #Bookofwhy. LLMs made it harder to test, for the latter could fake having a model by simply citing texts from authors who had world models, see https://ucla.in/3L91Yvt https://ucla.in/3L91Yvt The question before us is: Should we care? Or, can LLMs fake having a world model so well that it wouldn't show up in performance? If not, we need a new mini-Turing test to distinguish having vs. not-having a world model. The thing here is that LLMs are trained on most of the internet. Some people are surprised by some results but don't seem to know what kind of content is on the internet or in the training set (to be fair, we don't know what all the training sets are). There's lots of people that believe LLMs have world models (there have even been papers written about this!) but it's actually pretty likely that the LLMs were trained on similar examples, and this also helps explain why it is easy to break the world model (it isn't a world model if it is very brittle). We can see similar things with our tests like LSAT and GRE subject tests. Well guess what, there are whole subreddits and stack exchanges dedicated to these. Reddit is is most of these datasets and if you're testing on training data you're spoiled (there's a lot of spoiling that happens these days, but like Judea said, does it matter?) The problem with hyping up LLMs/ML/AI up too much is that we can no longer discuss how to improve the system. If you are convinced that they have everything solved then there's nothing to do. But no system is perfectly solved. Never confuse someone criticizing a tool with someone saying a tool is useless. I'm pretty critical of LLMs and am happy to talk about their limitations. That doesn't mean I'm also not wildly impressed and use them frequently. There's too much reaction to criticism as if people are throwing the thing in the dumpster rather than just discussing limitations. FWIW, I wouldn't change my mind if DM's test showed the opposite result. You can check my comment history. I'd probably dig in and rather comment why DM's tests were bullshit. I've even made comments about how chain of thought is frequently a type of spoiling. The reason being not because I don't think we can't create AGI (I very much do), but because I have a deep experience with these models and ML in general and nothing in my experience and understanding leads me to even believe half the things people claim. You'll see me make many comments ranting about the difficulties of metrics and how there's this absurd evaluation that people do (not just in ML) of using a proxy test set and using some proxy metric and saying performance on it is enough to claim a model is better. It is a ridiculous notion. [0] https://twitter.com/yudapearl/status/1710038543912104050 https://twitter.com/yudapearl/status/1710038543912104050
- klyrs 3y ago> Often with slightly different random values or temperature or phrasing, the same LLM will give different results for the same problem. You're fundamentally describing a stochastic parrot here, and not an intelligence. When you say that humans can also make the same mistake, you're ignoring the fact that every human testing these systems and finding it lacking are also comparing their experience to a lifetime of interacting with human intelligence. Every anecdote of this sort is an example, typically based on repeatedly trying an assortment of prompts (eliminating two of your variables -- random values & varying prompts), on a large variety of tasks, which is curated down to a single example for the sake of brevity. To say that "oh maybe it would have gotten the right answer if you got lucky or tried harder or twiddled a knob that you don't have access to" is simply nowhere near the extraordinary evidence that is required to prove the extraordinary claim of "real understanding." The reason that you don't know what to do about these anecdotes is that you lack the evidence to properly rebut them. You can only wave your hands at ill-defined properties and gaslight about the user holding it wrong. Perhaps you should ask your parrot buddy what to do. But if you were really serious about this, you wouldn't be going after the strawfolk down in the comments, you'd rebut the paper itself.
- yawnxyz 3y agoThis is super interesting because you’ll end up with different people with different opinions on “what is the best answer” if they all have different experiences. Most of us when using LLMs feel it should only produce the one correct answer; sometimes that’s the right way to think about it, but other times we might want an opinionated answer, when there is no true answer. So I think if we get better learning and self correcting machines, we also get more opinionated and sometimes incorrect machines. Kind of like people. And sometimes they’re stubborn or confidently incorrect, but that I think actually points to more intelligence (in people at least), since it points to people deciding based on their experiences.
- lossolo 3y ago> This is super interesting because you’ll end up with different people with different opinions on “what is the best answer” if they all have different experiences. In context of grandparent comment it's code, so while we can have different solutions to the same problem, we can verify if the solution is correct because it's not an opinion on some subject.
- jfim 3y agoYou can see it in various tasks where the answer for a text generator would be incorrect, yet simple reasoning or understanding yields the correct answer. For example, the everything fits in the boat variant of the farmer across the river problem: https://chat.openai.com/share/7d2de1c4-1cd6-4d7d-97e4-645830c0647a https://chat.openai.com/share/7d2de1c4-1cd6-4d7d-97e4-645830...
- throwaway4233 3y agoI do not think that answer would correctly solve the original riddle. The LLM blindly just gave an answer without raising the question of which pair of objects together would be a problem if left alone.
- jfim 3y agoGood catch, it actually leaves the goat with the cabbage in the answer above.
- fragmede 3y agoIt figures it out if you point out the difference much like a human would when faced with a riddle they didn't get. https://chat.openai.com/share/34df3eb0-c41c-4c5b-80d2-48f5d00d200a https://chat.openai.com/share/34df3eb0-c41c-4c5b-80d2-48f5d0... I'll be honest, the only reason I knew to read your variation of the farmer-boat problem carefully is because you set me up for success - I knew it wasn't going to be the original. Curiously, I haven't been able to get it to one-shot the question, even by prefixing it with instructions.
- godelski 3y agoExcept the problem here is that you spoiled the solution. It is quite difficult to provide hints that don't spoil. Or rather, don't provide hints but rather tell it that it is wrong. Here, I'll demonstrate it. First I'll prod it with vague responses about it just generally being wrong. Trying to leak no information to the model. We'll try to slowly add a bit more and then use your followups. Notice that the model cannot escape the overfit regime without your strong hints. You told the model specifically what to consider. You told it that it is a trick question of a trick question. This is not how a human would handle the situation. To my followups they wouldn't spit out the same answers. They'd actually likely followup with questions if they were confused. Which is a behavior I've never seen from an LLM: asking clarifying questions. https://chat.openai.com/share/57ab9bca-326d-45cb-9257-7fb8c2333a73 https://chat.openai.com/share/57ab9bca-326d-45cb-9257-7fb8c2...
- graeme 3y agoOne example. I asked ChatGPT to take 7 letters and find all the word combos. There were about 80 possible. It maxed out at 19 or so. Now clearly this is a hard task. Humans can’t do it well and none of the text corpus would offer training on this sort of thing. If I asked a human to do it they would say “here’s what I got, I don’t think that is all of them” ChatGPT would apologize when corrected and make a new, equally wrong list with full confidence. That’s how I interpret understanding. Self awareness and doubt. Oddly enough the code interpreter seems much stronger on this point. Now that I think about it I don’t think I was using it for the letter combo task.
- fragmede 3y agoChatGPT will happily tell you it doesn't know anything after January 2022, and it'll say "jdnrirmd-gurlksjd" is not a recognizable term, instead of trying to bullshit a definition of that, if that's your bar for "understanding",
- deleted 3y ago[deleted]
- jiggawatts 3y agoCurrent LLMs use an optimisation of processing data one word at a time. This makes them bad at solving puzzles involving individual letters or digits. This is a well known, well understood limitation that can be overcome (Facebook published a hierarchical model that can parse individual characters), but this technique isn’t used by ChatGPT. Whenever I see a criticism like this, it says more to me about the hubris of humans and the falsity of their assumed superior intelligence. You’re criticising the intelligence of a thing you don’t understand yourself! You didn’t “read the manual”. You didn’t go find out why the AI is failing. You just spouted an angry comment and gave up without gaining any understanding. PS: the “stupid AI” understands all of this. Just ask it to explain it to you: https://chat.openai.com/share/7bc1c1cb-d888-4ccb-904a-79e9ae22b950 https://chat.openai.com/share/7bc1c1cb-d888-4ccb-904a-79e9ae...
- somewhereoutth 3y agoYou have it backwards - there might be some similarities between LLMs and humans. I don't think any such almost certainly superficial similarity is meaningful in any way.