6 ms·
I'm not sure what you mean by the system learning "perfectly well" as it literally knows nothing about what those words mean. This is why LLMs can give statist
by gibsonf1 3y ago
I'm not sure what you mean by the system learning "perfectly well" as it literally knows nothing about what those words mean. This is why LLMs can give statistical word results that can be true of false or in between and those systems have no way of knowing which status the result has as the system has no understanding - that is, it does not model what the words mean as we humans do to understand if something is true or false etc.
- yorwba 3y agoThe words in the article look like #aaaaaaabbbbbbb#, they do not mean anything.
- naasking 3y ago> as it literally knows nothing about what those words mean Speculation. We don't know what it means, mechanistically, to know what something means. Knowing what something means could consist of having a model of how a word relates to other things, in which case the system does indeed know something about what those words mean. It doesn't have precisely the same meaning as humans, because humans also relate words to other sense data, but there's still meaning.
- gibsonf1 3y agoSure we do. When you read a sentence, and are able to model that in your mind and understand the space-time being referred to, you understand it. LLM is like a calculator in the sense that given an input it gives an output. There is no modeling of space-time in the mind to know if the sentence makes sense etc. That's why if you read a language you don't understand, you get nothing as you can't link those words with the concepts in your mind to understand it. For the LLM, all languages are like that as it understands nothing.
- naasking 3y ago> For the LLM, all languages are like that as it understands nothing. Again, that's too far. An LLM will not assert that "married men are bachelors" because it's been trained that "bachelor" is not associated with "married". Even if it doesn't understand what a bachelor is or marriage is, it understands that these two concepts are not positively associated with each other, and in fact, are negatively correlated. Logical associations like this are semantic content, which thus exhibits some degree of understanding. That's why LLMs can make sense, even when they're lacking the full understanding of a human.
- gibsonf1 3y agoI'm not arguing that they are not fantastic statistics machines, but there is nothing like semantics or "two concepts" for LLM, its just word strings that statistically occur together, etc. This is precisely why they are so convincingly wrong, wright, or in between, and the technology has no way to ever "know" the difference.
- naasking 3y agoAgain, we don't have a precise understanding of what distinguishes syntax and semantics, or whether there's any difference at all. You're just asserting that logical relationships are not semantics, but that's not a proof, nor are fallacious thought experiments like Searle's Chinese Room legit proofs that this distinction is meaningful. At the end of the day, semantics are about making meaningful distinctions. If all I gave you was a formula that related entities X and Y, but you didn't specifically know what X and Y represented, it is simply not correct to say that you don't know anything about X and Y despite not knowing what they are.
- EchoChamberMan 3y agoThese topics are awesome- they go to the heart of what it means to be human. With that, I'm not so sure your premise of "semantics are meaningful distinctions." In my mental model of an apple, I envision an apple against the backdrop of a void. The only existing relationship is "apple" to "not apple." This is a meaningful property, but I don't think it means I have any clue about what an apple is. I wonder about asking chatgpt "what is two plus two" a billion times or whatever. Given errors in the training data, and the negative correlation to answers that are not "four" is only a probability, will it tell me two plus two equals five? My gut says yes. And yes, humans could misspeak, but we can also self correct. "Apple-not apple" also makes me doubt adding space-time to the mix is enough for actual understanding.
- naasking 3y ago> With that, I'm not so sure your premise of "semantics are meaningful distinctions." That's a simplistic definition for introductory purposes. I think a more precise definition is that syntax consists of local properties of X, and semantics are broader, possibly global properties of X, eg. how X is related to Y, Z, A, B, etc. For any real world object, semantic properties of an apple pull in all sorts of sophisticated relationships depending on the perspective, eg. redness, taste, crunch if you're hungry, genes if you're a biologist, particles if you're a physicist, etc. But in the end, we're all just locally acting particles and fields (syntax) lacking any semantic properties, and so semantics must be emerge from syntax. It's the only thing that makes sense in a mechanistic universe. It seems clear that LLMs do build a model relating X, Y, Z, A, B, etc., it's simply lacking some relations, namely, the sensory relations that humans have. That's a lot of missing data, which is one of the reasons why LLMs make mistakes. > My gut says yes. And yes, humans could misspeak, but we can also self correct. Yes, LLMs don't typically have review-loops, where they check what they were about to say before spitting out an answer. But if you explicitly ask it to review its answer for any errors, it's accuracy goes up noticeably. There are plenty of follow-up papers that add this sort of review automatically and show considerable benefits. There's still probably still some form of generalization that's missing to make best use of training data, but I still think it's wrong to say that LLMs lack any understanding. I think this is a good analysis of the situation: https://reddit.com/r/naturalism/comments/1236vzf/on_large_language_models_and_understanding/ https://reddit.com/r/naturalism/comments/1236vzf/on_large_la...
- eli_gottlieb 3y agoYou appear to know literally nothing about what's in the article.
- chaxor 3y agoIf you could, would you mind pointing us to some foundational theory on this that actually proves (preferably with analytical expressions or some other form of mathematics) anything around these statements? It sounds like you read some of Emily Bender, et al's works and then simply regurgitated them. Unfortunately, everyone in this camp simply has large walls of text for their arguments, and basically no data, proofs, or any evidence whatsoever. The 'evidence' around "system X doesn't understand, like I do" amount to 'I don't like this' and 'trust me bro'. In formal lingiistics, Chomsky has a point against ML for LMs, which is based around the lack of interpretability. His works did push our understanding forward - because they included formal mathematics in proving his statements. However, they tended to be more specific about what limits exist around expressivity, rather than your issues laid out here. There are strong arguments to say that we do learn and represent language in a similar way to these NNs, at least at some level. 1) the distributional hypothesis and 'words defined by company they keep' is intuitive for how we learn language, and has provided enormous advances in NLP starting with w2v. 2) Dozens of academic papers from around the globe in different labs all reproducing the result in various ways that the distributional hypothesis word vectors (either contextual as in Bert/Gpt, or not as in gensim) map to brain activations by fMRI. This result also tracks in more resolute experiments in visual modalities for more invasive experiments in non-human primates. I could go on providing data and experiments, but it would be nice if the opposing side to this would provide any at all that were convincing.
- foobarqux 3y agoPutting aside the fact that everything you said is wrong or unsubstantiated, LLMs don’t work like humans because they can learn languages/grammar that humans cannot.
- chaxor 3y agoYou haven't provided any counter argument, and there are many articles with data backing this up. Unfortunately I am not aware of any articles that can show convincing data that "LLMs don't learn like humans". I'm not really even sure what that means precisely. Your statement could be understood to claim that language models cannot predict brain activations or vice versa? Predictive power is the foundation of science, so it's one of the better tests we have for these types of problems. To that point, there is substantial evidence of similarity and predictive power, as noted below. Perhaps you mean something else? Surely you're point is not that GPUs simply aren't human. Of course - there are differences at various levels. So at some level, everyone recognizes they're not identical. Perhaps it's that attention architectures are not reproduced within neural microarchitectures in tbe neocortex? I haven't seen any studies on this within the relevant cortical areas to address this, though microarchitecture matching may not be required for output distribution matching. The interesting endeavour for science to show right now is how these systems are similar to our processing mechanisms. The ways in which they're not similar typically tend to be uninteresting, or obvious. The fact that we can predict brain activity from NN model activations raises many more opportunities in science, and the predictive power has gotten far better with these advances,. Here are a small set of articles from MIT, deepmind, etc. for some of the points I made: - Tuckute, Greta, Aalok Sathe, Shashank Srikant, Maya Taliaferro, Mingye Wang, Martin Schrimpf, Kendrick Kay, and Evelina Fedorenko. "Driving and suppressing the human language network using large language models." Nature Human Behaviour (2024): 1-18. - Schrimpf, Martin, Jonas Kubilius, Ha Hong, Najib J. Majaj, Rishi Rajalingham, Elias B. Issa, Kohitij Kar et al. "Brain-score: Which artificial neural network for object recognition is most brain-like?." BioRxiv (2018): 407007. - Pereira, Francisco, Bin Lou, Brianna Pritchett, Samuel Ritter, Samuel J. Gershman, Nancy Kanwisher, Matthew Botvinick, and Evelina Fedorenko. "Toward a universal decoder of linguistic meaning from brain activation." Nature communications 9, no. 1 (2018): 963. - Arend, Luke, Yena Han, Martin Schrimpf, Pouya Bashivan, Kohitij Kar, Tomaso Poggio, James J. DiCarlo, and Xavier Boix. Single units in a deep neural network functionally correspond with neurons in the brain: preliminary results. Center for Brains, Minds and Machines (CBMM), 2018.