5 ms·
I don't get it. The paper reads like 10 pages of opinion and casting aspersions on language models. No math. No graphs.
by peachfuzz 6y ago
I don't get it. The paper reads like 10 pages of opinion and casting aspersions on language models. No math. No graphs.
- tsimionescu 6y agoYou don't need math to explain why, for example, a statistical model trained to find the plausible combinations of words or phrases in its corpus that matches a prompt is not a promising approach to NLP and general understanding, and using it to win at tasks designed to assess NLP success essentialy amounts to cheating. You also don't need math to explain that, in very real terms, the meaning of any phrase produced by GPT-3 lies in the mind of the reader and not in GPT-3's output, which doesn't see any meaning in the phrases it produces (unlike, say, the output of a rule-based system, which, while much more rudimentary, generally has reasoning about the topic at hand behind it, and not about word probabilities).
- blackbear_ 6y ago> to explain why, for example, [...] is not a promising approach to NLP and general understanding I am not aware of any definitive proof or at least convincing argument of why this is not possible, though. Timnit's paper takes that as a given fact and illustrates consequent dangers and high-level workarounds.
- tsimionescu 6y agoI didn't say that it's not possible, just that it's not likely. It's clearly not how the human mind works, and since that is the only computer we know for sure can do NLP, any approach that is so obviously alien to it is unlikely to be a good approach. In case this is not clear, the human mind obviously doesn't do stochastic prediction on phrases, it has a model of the world ( based on agents interacting with objects through mechanical laws) and it produces or interprets speech by assessing its interaction with this model of the world. When assessing whether to produce phrase A or phrase B, it doesn't assess the likelihood of this phrase in the corpus of phrases it has seen/heard before, but on its appropriateness to the situation. Most prompts for language use are not language at all, but come from the world itself [0], something which pure LMs can't even in principle do (they they could potentially be combined with other kinds of models to achieve this). [0] for example, 'seeing a fire in a crowded theater' is the prompt for yelling 'Fire!'; the correct response/next step from hearing a shout of 'Fire!' in a crowded theater is not a phrase, it is a desperate attempt to run away
- blackbear_ 6y ago> In case this is not clear, the human mind obviously doesn't do stochastic prediction on phrases. Are you sure? Because this is exactly what you and I are doing right now. A language model is a very appropriate description for how we see each other: you produce some text, I produce some more, and you respond based on that. And I inject some stochasticity by re-typing this paragraph five times trying to make it sound like natural English and a cohesive text, a bit like beam search if you are familiar with that. What I am trying to say is that language models are a good interface (in the programming sense) to describe human interactions on the internet: text in, text out. So while I agree that GPT-3 is not a realistic "human-like" implementation of this interface, I don't see why a priori a neural network cannot eventually incorporate world models with agents and so on and reach actual understanding (whatever that means) of text.
- tsimionescu 6y agoWe may be agreed in fact, to some extent at least. My claim, and I believe the article's as well, is that this model of the world and agents and everything will not probably be learned just by consuming ever larger quantities of text and adding ever more parameters to the model. But I would also claim that our conversation now is unlikely to be predictable without some model of the world. It's not 'you produce some text, I produce some text', it's 'you produce some text, I run it through my model of the world, my inference engine produces some other model of the world, and then that gets translated to some text in reply'.
- ineedasername 6y agoany approach that is so obviously alien to it is unlikely to be a good approach Computers very frequently do not solve problems the same way humans do, so I'm not sure that's a significant point against any particular modeling technique. Otherwise, if you're waiting for language models to understand language he same way that human minds do, you're probably waiting for AGI, not any particular breakthrough in NLP alone.
- ineedasername 6y agoSure, but without something more concrete to support criticisms, it's only a theory, however compelling the reasoning, that this might be a bad approach. Right now it reads like a detailed literature review. Such things are usually the starting point for research, not and end in themselves. I think the authors make a promising start here, but I look forward to further work from them so that if their criticisms are correct the field can turn to more fruitful approaches.
- lupire 6y agoIt's a philosophy paper.