3 ms·
I'm not arguing that they are not fantastic statistics machines, but there is nothing like semantics or "two concepts" for LLM, its just word strings that stati
by gibsonf1 3y ago
I'm not arguing that they are not fantastic statistics machines, but there is nothing like semantics or "two concepts" for LLM, its just word strings that statistically occur together, etc. This is precisely why they are so convincingly wrong, wright, or in between, and the technology has no way to ever "know" the difference.
- naasking 3y agoAgain, we don't have a precise understanding of what distinguishes syntax and semantics, or whether there's any difference at all. You're just asserting that logical relationships are not semantics, but that's not a proof, nor are fallacious thought experiments like Searle's Chinese Room legit proofs that this distinction is meaningful. At the end of the day, semantics are about making meaningful distinctions. If all I gave you was a formula that related entities X and Y, but you didn't specifically know what X and Y represented, it is simply not correct to say that you don't know anything about X and Y despite not knowing what they are.
- EchoChamberMan 3y agoThese topics are awesome- they go to the heart of what it means to be human. With that, I'm not so sure your premise of "semantics are meaningful distinctions." In my mental model of an apple, I envision an apple against the backdrop of a void. The only existing relationship is "apple" to "not apple." This is a meaningful property, but I don't think it means I have any clue about what an apple is. I wonder about asking chatgpt "what is two plus two" a billion times or whatever. Given errors in the training data, and the negative correlation to answers that are not "four" is only a probability, will it tell me two plus two equals five? My gut says yes. And yes, humans could misspeak, but we can also self correct. "Apple-not apple" also makes me doubt adding space-time to the mix is enough for actual understanding.
- naasking 3y ago> With that, I'm not so sure your premise of "semantics are meaningful distinctions." That's a simplistic definition for introductory purposes. I think a more precise definition is that syntax consists of local properties of X, and semantics are broader, possibly global properties of X, eg. how X is related to Y, Z, A, B, etc. For any real world object, semantic properties of an apple pull in all sorts of sophisticated relationships depending on the perspective, eg. redness, taste, crunch if you're hungry, genes if you're a biologist, particles if you're a physicist, etc. But in the end, we're all just locally acting particles and fields (syntax) lacking any semantic properties, and so semantics must be emerge from syntax. It's the only thing that makes sense in a mechanistic universe. It seems clear that LLMs do build a model relating X, Y, Z, A, B, etc., it's simply lacking some relations, namely, the sensory relations that humans have. That's a lot of missing data, which is one of the reasons why LLMs make mistakes. > My gut says yes. And yes, humans could misspeak, but we can also self correct. Yes, LLMs don't typically have review-loops, where they check what they were about to say before spitting out an answer. But if you explicitly ask it to review its answer for any errors, it's accuracy goes up noticeably. There are plenty of follow-up papers that add this sort of review automatically and show considerable benefits. There's still probably still some form of generalization that's missing to make best use of training data, but I still think it's wrong to say that LLMs lack any understanding. I think this is a good analysis of the situation: https://reddit.com/r/naturalism/comments/1236vzf/on_large_language_models_and_understanding/ https://reddit.com/r/naturalism/comments/1236vzf/on_large_la...
- EchoChamberMan 3y agoInteresting points, and thanks for the link!