3 ms·
A bit un related but this reminded me of a question I had some months ago: Let's say you train a bunch of ML models to solve a problem, these models are m_1, m
by ccortes 6y ago
A bit un related but this reminded me of a question I had some months ago:
Let's say you train a bunch of ML models to solve a problem, these models are m_1, m_2,..., m_n and are ordered according to their performance on the tests/validation set where m_1 is the best model.
Should we expect to see regression to the mean on their predictions scores once we let them do their thing on fresh sets?
- gwern 6y agoYes, because your validation set performance is imperfectly correlated with performance on more samples (because it's of finite size), so it will tend to overestimate within-distribution. There's also the problem that you probably want to deploy it on data which was not exactly collected the same way. So you have issues of both internal and external validity. In practice, at least for CNN classifiers on ImageNet etc, it seems to be a pretty minor issue overall. Some links: https://arxiv.org/abs/1711.11561 https://arxiv.org/abs/1711.11561 https://arxiv.org/abs/1806.00451 https://arxiv.org/abs/1806.00451 https://arxiv.org/abs/1902.10811 https://arxiv.org/abs/1902.10811 http://gradientscience.org/data_rep_bias/ http://gradientscience.org/data_rep_bias/ http://gradientscience.org/data_rep_bias.pdf http://gradientscience.org/data_rep_bias.pdf https://arxiv.org/abs/1905.10498 https://arxiv.org/abs/1905.10498 https://arxiv.org/abs/2002.02559 https://arxiv.org/abs/2002.02559 https://arxiv.org/abs/2006.07159 https://arxiv.org/abs/2006.07159 https://arxiv.org/abs/1805.08974 https://arxiv.org/abs/1805.08974 https://arxiv.org/abs/1902.10178 https://arxiv.org/abs/1902.10178 https://arxiv.org/abs/1912.11370#google https://arxiv.org/abs/1912.11370#google
- david2ndaccount 6y agoYes, but this is true of anything when you estimate some quality and pick the “best” from a group. See the winner’s curse.
- heyitsguay 6y agoDepends on how you do your validation, basically. It's easier to give an example if we assume m_1 is your worst model and m_n is your best: if you do a train/eval split, train m_1, eval m_1, make some tweaks to m_1 to improve that eval score and call the tweaked model m_2, etc., you'll implicitly be overfitting your better models to the eval data even though you're never actually training on it. In which case, yes, it is quite possible that given some new unrelated data, m_n will do no better than m_1 or m_2 or whatever. If you do a train/eval/test data split instead, train+eval m_1 through m_n like before, and then rank them based on performance on the test data only after model selection is finished, you won't have that implicit overfitting to the test data, and m_n stands a better chance of beating m_1 on new datasets. There's another way things can go wrong - if your train/eval/test data doesn't capture the full range of variability of the data source that will be producing your fresh datasets, then even if you do the train/eval/test split correctly, it's quite possible that m_n will do no better than m_1 when operating on fresh data outside of the train/eval/test distribution. Basically, it is possible to build a sequence of models that actually improve performance on new data, but there are also lots of ways to get it wrong and produce models that look like they're doing better in a controlled environment, but fail to do better in the wild.
- Shorel 6y agoSounds like halfway to a genetic algorithm to train the networks.