3 ms·
Validation on holdout sets. When I was a student in the 1990s, I was taught about hypothesis testing (and all the hassle of p-fishing etc.), and about Bayesian
by datastoat 5y ago
Validation on holdout sets.
When I was a student in the 1990s, I was taught about hypothesis testing (and all the hassle of p-fishing etc.), and about Bayesian inference (which is lovely, until you have to invent priors over the model space -- e.g. a prior over neural network architectures). These are both systems that tie themselves in epistemological knots when trying to answer the simple question "What model shall I use?"
Holdout set validation is such a clean simple idea, and so easy to use (as long as you have big data), and it does away with all the frequentist and Bayesian tangle, which is why it's so widespread in ML nowadays.
It also aligns statistical inference with Popper's idea of scientific falsifiability -- scientists test their models against a new experimental data, data scientists can test their model against qualitatively different holdout sets. (Just make sure you don't get your holdout set by shuffling, since that's not what Popper would call a "genuine risky validation".)
The article mentions Breiman's "alternative view of the foundations of statistics based on prediction rather than modeling". Breiman does make a big deal of evaluation on holdout sets; but his "prediction" idea isn't general enough, since it doesn't accommodate generative modelling (e.g. GPT, GANs). I think it's better to frame ML in terms of "evaluating model fit on a holdout set", since that accommodates both predictive and generative modelling.
- anxrn 5y agoVery much agree with the simplicity and power of separation of training, validation and test sets. Is this really a 'big data' era notion though? This was fairly standard in 90s era language and speech work.
- datastoat 5y agoBig enough data that you can afford not to use some of it for training! Different disciplines hit this threshold at different times -- language and speech much earlier, as you say; clinical trials not there yet. Maybe we could talk about two cardinalities of "big" data. The first is when you can afford not to use all of your data for training. The second is when you can usefully fit highly overparameterized models.
- disgruntledphd2 5y agoTo be fair, there's a psychology paper from the late fifties that suggests this approach. Much like the early days of double descent, this didn't attract the attention it deserved at the time.
- vlmutolo 5y agoIf you don't get the holdout set by shuffling, how do you get it?
- datastoat 5y agoIdeally, there's some field in our data that we can use -- e.g. train on data from cities in several countries, test on data from cities in a different country; train on one corpus, test on another; train on climate data at one level of CO2, test on data at another. If there's no such field, I do a simple 1d dimension reduction, and pick off 20% of the data at the extremes. If we know exactly the domain where our model will be used, and it matches exactly the disposition of the training data, then shuffling is fine. But if we want our model to work well in new domains, then we're like scientists looking for generalizable laws, and the best way to test this is by testing in unseen domains. It's often hard to get hold of data with this domain-diversity, which is why a lot of ML models are fragile.
- monkeybutton 5y agoWalk the walk! If your data has any notion of temporality, train on [t-K, t] and predict t+1.
- PeterisP 5y agoAs you often want to predict generalization to future unseen data, it's important to consider how the unseen data will be different from your current dataset - all the theory operates on the simplifying assumption that your dataset is an IID random sample from the "true" data but, of course, usually that's not really the case and you know and expect that the future data will have a different distribution than what you have - economics and politics trends will have new major events; future textual data will see new names, event names, terms and concepts that did not exist in 2020; etc, etc. So if you want to estimate actual generalization, you have to at least try to isolate some of these aspects. For text, if you take a random sample of shuffled sentences then you'll get a different outcome than if you shuffle whole documents, because it'll be "cheating" as your test sentences will always have some relevant context in your training set, which won't be the case for real unseen data which will sometimes introduce totally new things. And if you train something on data gathered 2015-2020 and evaluate on data from 2021, then you'll likely get worse (but more informative!) measurements than training on a random sample chosen from the whole 2015-2021 range, simply because there are major world events and 'distribution shifts' over time. IMHO a true test for generalization of many approaches would be to train stuff on data before 2020 and test on data from 2020 - to see how well it generalizes given a major event like Covid pandemic that changes all kinds of aspects everywhere in society that generates the new data.