3 ms·
When they don't generalize well it usually means you have over trained on your training data. When you take classes on this stuff they have whole sections that
by johnwatson11218 7y ago
When they don't generalize well it usually means you have over trained on your training data. When you take classes on this stuff they have whole sections that talk about trying to detect this and what to do about it e.g. regularization, better models, more training data etc.
- Cybiote 7y agoWhat you list remains insufficient to tackle the difficulty of extrapolation. Extrapolation of the kind we're able to do with Physics theories is difficult in the general case for all methods, not just deep learning. With even the relevant variables subject to change, things like distribution shift and non-stationarity are but the tip of the ice-berg. For neural networks, if you take something basic like sorting a list or multiplying two decimal numbers, the further you are from range the models were trained on, the worse they will do (yes, transformers too). Only exception I can think of are carefully trained Neural GPUs, which will quickly struggle to be 100% correct as you depart simple tasks. While consuming a great deal of computational resources. Program synthesis is the general area, with no approach clearly dominant in the same way deep learning has dominated machine learning.
- deleted 7y ago[deleted]
- YeGoblynQueenne 7y agoWhat is meant above by "generalisation" is generalisation on out-of-sample data, i.e. not the sample on which you train and test your model (whether you have access to the test set or not). If out-of-sample data has a different distribution than the sampled dataset, then you're SOL, regularisation or not. As a for instance, here's an interesting paper I found out about from HN that describes how image classifiers trained on one standard dataset (ImageNet etc) do much worse on other standard datasets and how it's even possible to identify the dataset a classifier was trained on: Unbiased look at dataset bias https://people.csail.mit.edu/torralba/publications/datasets_cvpr11.pdf https://people.csail.mit.edu/torralba/publications/datasets_...
- tel 7y agoNot strictly true, regularization in various forms can help have expert knowledge apply to create better extrapolation. Parametric models could be thought of as a kind of regularization even.