3 ms·
These are tests which require reasoning in humans. A human can't reasonably be expected to pass Quant or Verbal tests without reasoning about the information gi
by pwinnski 4y ago
These are tests which require reasoning in humans. A human can't reasonably be expected to pass Quant or Verbal tests without reasoning about the information given, but--and I can't stress this enough--large language models are not humans.
Here's a sample question from the GRE Verbal Reasoning test:
> Upon visiting the Middle East in 1850, Gustave Flaubert was so [blank] belly dancing that he wrote, in a letter to his mother, that the dancers alone made his trip worthwhile.
> (A) overwhelmed by
> (B) enamored by
> (C) taken aback by
> (D) beseeched by
> (E) flustered by
Whether that specific question was in the training corpus or not, there are enough words in the sentence to suggest a positive association, including, significantly, "worthwhile." That alone possibly serves to narrow the answers down to A or B, with a preference for B, because it's more likely that "worthwhile" and "letter to x mother" are associated with "enamored" in general English-language text.
Look, the whole point of these models is that its not easy or even possible to trace the path any given input takes on its way to output, but we know the principles used in development, so I think it's rather more of a burden to explain how the clearly-explained principles in LLMs result in something other than the obvious. The fact that the results are so overwhelming that we become enamored by them, well, I'm taken aback by the seeming accuracy of some of the responses, but I beseech you to remember the other responses in which these LLMs are dramatically off-base, as if flustered--if LLMs could ever be flustered.
- adamsmith143 4y ago> These are tests which require reasoning in humans. A human can't reasonably be expected to pass Quant or Verbal tests without reasoning about the information given, but--and I can't stress this enough--large language models are not humans. A tale as old as time: "It is illustrated by the success of chess computers. In the 60s, it was said that computers will never beat people at chess, because that requires intelligence and computers aren't capable of intelligent thought. When computers regularly started winning matches in the 80s, it was claimed that playing chess wasn't a test of real intelligence because computers could do it." >Whether that specific question was in the training corpus or not, there are enough words in the sentence to suggest a positive association, including, significantly, "worthwhile." That alone possibly serves to narrow the answers down to A or B, with a preference for B, because it's more likely that "worthwhile" and "letter to x mother" are associated with "enamored" in general English-language text. Yes, this is called reasoning so in other words the LLM is reasoning about language. >Look, the whole point of these models is that its not easy or even possible to trace the path any given input takes on its way to output, but we know the principles used in development, so I think it's rather more of a burden to explain how the clearly-explained principles in LLMs result in something other than the obvious. The inscrutability of the Matrices tells you nothing about it's reasoning ability. Given the correct prompt the LLM will also provide you with a step by step solution to the question it answered. There are also explicit reasoning prompts that these models are able to deal with. I think it's pretty simple, if these questions are not explicitly in the training data it cannot have answered them correctly at such a high success rate with anything other than reasoning. You haven't given any alternative answer to how it does this either.
- pwinnski 4y ago> You haven't given any alternative answer to how it does this either. I have. You refuse to accept it, but I definitely have given an answer that involves tokenization and association, the things we already know LLMs use to construct their responses.
- adamsmith143 4y ago>Whether that specific question was in the training corpus or not, there are enough words in the sentence to suggest a positive association, including, significantly, "worthwhile." That alone possibly serves to narrow the answers down to A or B, with a preference for B, because it's more likely that "worthwhile" and "letter to x mother" are associated with "enamored" in general English-language text. This is reasoning, what you've described is reasoning.
- pwinnski 4y agoIs your claim that all reasoning is ultimately math we do unconsciously? That humans "reason" via word-association and probability? That seems to be what you're suggesting, but I don't want to conclude that without your say-so.
- adamsmith143 4y agoIf I throw a ball at you and you attempt to catch it do you think the brain is NOT doing unconscious math to predict the likely trajectory of the ball so you know where to place your hand to catch it? What processes outside of Physics and Math and Probability is the brain using to interact with the world?
- pwinnski 4y agoIt was a yes/no question! It's like trying to nail jello to a wall, so I'll stop trying to ask what you actually mean.