6 ms·
What's your definition of "real understanding", and how would one definitively know when it's present? Various humans will display the same type of pathologica
by mrshadowgoose 3y ago
What's your definition of "real understanding", and how would one definitively know when it's present?
Various humans will display the same type of pathological behavior within various reasoning tasks, do all humans lack "real understanding" as well?
- thangalin 3y agoIN: What consonants usually make a hard sound? GPT: In English, the consonants that typically make a hard sound are voiceless plosives (/p/, /t/, /k/), voiceless fricatives (/f/, /θ/, /s/, /ʃ/), ... snip. IN: Great. Rephrase the following without using a hard consonant sound: «What was once dormant may soon angrily awake.» GPT: "What was previously inactive may soon vigorously awaken." IN: K is a hard consonant. Try again. GPT: Apologies for the oversight. Here's a rephrased version without using any hard consonant sounds: "What was once asleep may soon awaken with intensity." IN: K is still a hard consonant sound. Try again. GPT: Apologies for the confusion. Here's another rephrased sentence without using any hard consonant sounds: "What was once quiet may soon awaken with fervor." IN: Both q and k are hard consonants. Try again. GPT: I apologize for the oversight once again. Here's another attempt without using any hard consonant sounds: "The peaceful river flowed gently, carrying its secrets with grace." Ad nauseam.
- asah 3y agoI know many humans who wouldn't pass this test, e.g. in their non-native language... [ insert humiliating attempts to speak <language> ] Try instead to give GPT more examples of pronunciation (1000+) and then I suspect it'll do fine.
- glhaynes 3y agoIsn't this just arguing for giving it more text to do text generation from?
- pixl97 3y agoDepends what you consider vocal models 'more text'. Already we're seeing multi-modal AI have a correct answer when executing across image/text or voice/text interpretation.
- hajile 3y agoIf I had access to a dictionary with pronunciation symbols and a thesaurus, I could do it regardless of language. The LLM not only has access, but has baked both its model. This problem is EXACTLY the kind of problem computers should be good at. The real issue is that LLMs don’t think.
- vidarh 3y agoHumans "have access" too in the languages we speak in the same way the LLM does, and I'd posit that many of us would still fail. It's "baked into us" just like it is into an LLM. I don't see any reason why you'd assume an LLM would be better able to recall this information than a human who will if anything have vastly more training at mapping words to sounds. It's a problem a computer with a custom written algorithm should be good at. Why we'd assume a model trained on just general data will automatically be good at this is a bizarre notion to me. We don't automatically assume humans will be great at everything just because we've passively consumed lots of content. To me, a whole lot of these "LLMs don't think" claims comes from not thinking about what reasoning and thinking is, and whether or not how we try to measure that makes any sense at all.
- vanviegen 3y agoYes, these responses are annoying, but what's your point? Humans too fail to perform many tasks (that llms can carry out), even if given many chances to correct themselves. The way in which these llms fail is very unhumanlike though - they will keep trying the same failing strategy over and over, whereas a human would try the failing strategy a couple of times, and then start yelling at the person giving them impossible orders. If that's what we prefer, I'm sure a tiny bit of additional training can teach llms to throw a tantrum. :-)
- ratherbefuddled 3y ago> Yes, these responses are annoying, but what's your point? The point I imagine is that there is no reasoning going on at all. Some humans sometimes struggle with some reasoning, of course. That is completely irrelevant to whether LLMs reason. Picking word sequences that are most likely acceptable based on a static model formed months ago is not reasoning. No model is being constructed on the fly, no patterns recognised and extrapolated. There are useful things possible of course but these models will never offer more than a nice user interface to a static model. They don't reason.
- maxdoop 3y agoAnd so, what do you think reasoning is? Or how would you know if something can reason or not ?
- vidarh 3y agoWhy do you say that isn't reasoning, and what do you think human reasoning is? I do think you have a point that the lack of a working memory is a severe constraints, but I also think you are wrong that these models will remain a user interface to a static model rather than being given the ability to add working memory and form long term memories and reason with that. I also think it's an entirely open question whether they are reasoning under a reasonable definition, in part because we don't have one, and I think any claim that they don't reason ironically comes from a lack of reasoning about the high degree of uncertainty and ambiguity we have with respect to what reasoning means and how to measure it.
- vidarh 3y agoWhile I'm sure you can find other such pathological examples, this one is "unfair" in as much as GPT does not "see" letters or sounds, but tokens that does not map to directly to either, and so you're effectively asking it to fumble around in the dark to no tool to observe or verify what it is putting together.
- Jensson 3y agoGPT knows the letters each word is made up of, it can turn small snippets of text to Base64 reliably. The only reason it fails is that it is too stupid to understand the connection, not that it doesn't know what letters those words are made up of.
- vidarh 3y agoIt knows the letters each word is made up when asked within a specific context. Humans also often struggle to recall things in one context that we have no problems with in another. It may well be reasonable to call it "too stupid" but at the same time it's unreasonable to call it stupid without then acknowledging that it can understand and reason. That it has gap in knowledge in areas we typically drill into young children and don't leave huge datasets online about is unsurprising to me. That said, I incidentally think a whole lot of adults - even native English speakers - would struggle with the task given, and would repeatedly fail in the same way until given a detailed refresher. E.g. being able to explain a rule and fail consistently to apply it is something I've seen up to and including supposed senior software engineers im interviews. Getting their mistakes explained and still repeating the same mistakes also.
- joelfried 3y agoYou have to try a little harder with GPT to get it to understand its mistakes, but it's not as bad as it used to be, at least if you pay. It failed for me the first time ("Earlier inert, it could rouse in fury before long"), but not the second ("In a lull, may soon yawn"). With a little more pushing it got a much better result ("In a lull, may soon arise"). Conversation, should any want to see it in full: https://chat.openai.com/share/421dac47-16df-499e-9b2e-d8dc0fccca21 https://chat.openai.com/share/421dac47-16df-499e-9b2e-d8dc0f...
- wilg 3y agoNearly every example of this sort of failure seems to involve asking the LLM to do an operation with text it is not capable of. This seems totally explicable to me by the tokenization process, which deletes the information about which letters are in the words. The string “awaken” gets turned into an arbitrary integer. So it can’t know this unless it’s specifically trained on spelling those words somehow. This seems like an implementation artifact, not a flaw in deep learning.
- Jensson 3y ago> asking the LLM to do an operation with text it is not capable of So you are saying the LLM isn't making world models? Because if it did it would understand the properties of these words, this is one of the easiest relationships it could find. It isn't fed these properties directly, but it knows the letters that each word is made up of anyway which is how it can turn them to Base64 etc, there is no reason at all why it shouldn't be able to solve that problem. But if you are right and this sort of thing is impossible, that implies that the LLM can't model anything at all, its just a stupid text generator. Is that what you meant?
- astrange 3y agoThere's no such thing as a world model. People don't solve problems by creating world models. That term was invented by 70s AI researchers, but remember that those people failed. Their research wasn't actually correct, so you shouldn't reuse it.
- wilg 3y agoNo. We know that LLMs make world models. (Often literally: https://twitter.com/wesg52/status/1709551516577902782 https://twitter.com/wesg52/status/1709551516577902782). I should also point out that your tone is aggressive, so it makes me think you're not interested in learning. But I'm going to proceed anyway. It cannot understand these properties of words well because it does not operate on words, it operates on tokens. You can see examples here: https://platform.openai.com/tokenizer https://platform.openai.com/tokenizer This makes it much more difficult to learn how to reason about particular data that appears inside the token, because it does not ever receive that information. The only way it could get that information is if the training data explicitly attempted to work around this limitation. But I would also appreciate insight from someone with a deeper understanding of the internals of these big LLMs.
- mike00632 3y agoThis is a very interesting example. I don't blame LLMs for not understanding which English text makes hard sounds. I wonder if it would get better if it were multi-modal.
- pixl97 3y agoI have a feeling that if you tried this with a human that has always been deaf/mute that they would also have a very difficult time. At least to me it's going to be interesting at these AI models become multi-modal and each mode can feed back into each other to formulate an answer. For example in the above questions I will subvocalize to reach an answer.
- nomel 3y agoWhich GPT version is this? Why is this relevant? Is the assumption that this example will also fail in future models? GPT-4 seems to handle this ok: https://chat.openai.com/share/f416b1ab-7f0c-43f1-b10a-37142df88d7e https://chat.openai.com/share/f416b1ab-7f0c-43f1-b10a-37142d...
- nomel 3y agoSorry, I can't read. It failed.
- maxdoop 3y agoHow is this in any way some retort to the claim that LLMs don’t understand something? I could ask the same questions to a child, and they’d respond with equally bad takes. Is the child incapable of understanding ?
- deleted 3y ago[deleted]
- spacecadet 3y agoSure, we barely understand the operation of our environment and universe, and thats not a knock at all of scientific progress- but I think much of them would agree, we lack a real understanding... Im starting to think the new enlightenment is recognizing this, because damn people seem to think they are all superhuman or something today...
- deleted 3y ago[deleted]