3 ms·
If that was true, a model trained on one commonsense Q&A dataset would be able to answer questions from another commonsense Q&A dataset without finetuning. But
by euphetar 4y ago
If that was true, a model trained on one commonsense Q&A dataset would be able to answer questions from another commonsense Q&A dataset without finetuning. But they can't. It's the same for every task you can find, but especially evident on commonsense reasoning benchmarks. At least last time I actively researched the question, I haven't found a single task where NNs definitely generalize.
When researchers dig in, they find that the neural network is learning wrong things. Word matching between answer and question, learning to model the annotator who asked a lot of questions because his Q&A answers are predictable, and stuff like this.
There was a great review of the problem, but I can't find it, so I will have to link to this article [1], which gives an overview of issues with current NLP models.
[1] https://thegradient.pub/frontiers-of-generalization-in-natural-language-processing/ https://thegradient.pub/frontiers-of-generalization-in-natur...
- geysersam 4y agoThe fact that NN does not generalize in a particular (arguably very challenging) task trained on a particular data set etc. does not mean NN never generalizes. Don't know exactly what is meant by "generalization" in this context. I'd argue that it's very unclear what the difference between "interpolation" and "extrapolation" even is in a high dimensional and sparsely sampled space. Would a NN trained to recognize "fur" trained on different kinds of dog fur also activate when it encounters wolf, cat or even bear fur? This seems quite plausible, and seems like a kind of generalization.
- euphetar 4y agoIt's indeed a very challenging task. I picked it because it's one of those where the lack of generalization is apparent. In fact, researchers study commonsense QA and similar tasks to find how we can reach generalization. I agree, it's very hard to fidn the difference. Especially for GPT or BERT, where the training set is basically the whole internet. It's a very good question about fur. I would suspect that it would correctly recognise all kinds of dog-like fur and sometimes fail on different furs, like bear furs. But in general NN's are very good at textures, so maybe it will just be good on all kinds of fur. One problem that might arise is that you will show it something fur-like but not really fur, and it will think it's fur. I agree it's some kind of generalization. Here I don't have enough background to draw the line, perhaps a more theorically oriented person could, but I can't. I guess the most important judge is that if you devise a benchmark like commonsense Q&A, a neural network fails it. Or how a Tesla will recognise a truck full of red stop signs as a real stop sign, while a "generalizing" thing like a human would definitely know that a core property of a stop sign is that it should be installed near a road. So there is a real problem.