42 ms·
I did feel genuinely bad for my comment, since it was a few shades nastier than I was going for and I try hard to be positive and not cut people down. It was me
by jointpdf 6y ago
I did feel genuinely bad for my comment, since it was a few shades nastier than I was going for and I try hard to be positive and not cut people down. It was mean spirited and I apologize.
My point was that intelligent humans can and often do make mistakes in logic and computation (arithmetic) in ways that machines typically do not. One reason may be colliding or incomplete representations of certain concepts, and (relatedly) the fact that we are relying on language. I think of neural networks as fuzzy representation composers, so it seems they also fail for similar reasons. Basically, it (GPT) does have some layered representation of the concept of numbers and how they are used in different contexts which gives it some faculty at carrying out common operations, but it doesn’t “add up” to a reliable system of logic (that would allow it to extend addition to say 100-digit numbers, the way even a sharp and/or patient 2nd grader could do, generalizing from the simpler cases).
I think accuracy is sometimes the correct measure, and in this instance it seems fine—at baseline, we should expect ~0% accuracy since it is generating output from essentially the space of all possible text (texts <= 2048 tokens). I agree that it would be interesting to probe the model with better tests, and understanding when/why it fails on certain arithmetic problems or types of reasoning.
I liked what you wrote about finding heuristics, though I disagree with your conclusion that heuristic finding does not qualify as learning—it is just somewhere along the spectrum between a randomized model and an ALU (neither of which can be said to have learned anything) in terms of its ability to perform arithmetic.
Of course, we already have better models for solving proofs and such, so I generally think the way toward more complete AI models is to return to the system design view of AI (meta-learning, integration of different models, etc) rather than trying to evolve one colossal model to rule them all. That is, a meta-model that recognizes what sort of problem it is facing, then selecting a model/program to solve or generate possible solutions to that problem, while revealing or explaining as much of this process as possible to the user.
In any case, I have definitely have more to read on the subject and am mostly musing at this point. Thanks for the references and the conversation.
- YeGoblynQueenne 6y agoYou're welcome, and really, please don't worry about your comment. It's all good :) I guess I can concede that the memorisation explanation is not the only possible one, there's always the possibility of learned heuristics. I still expect very strong evidence before I'm convinced that GPT-3 can learn arithmetic in the general sense and I don't trust the explanation that it's only learning partially- but let's agree to disagree on that. Thank you for the conversation, too.