4 ms·
The best way to reduce overfitting is with cross-validation. The general way is to set up a hold-out sample (or do n-fold cv if you don't have a lot of data) an
by selectron 10y ago
The best way to reduce overfitting is with cross-validation. The general way is to set up a hold-out sample (or do n-fold cv if you don't have a lot of data) and then use this cross-validation hold-out sample to do feature, parameter, and model selection. With this technique however there is a risk of overfitting to your hold-out sample, so you want to use your domain expertise to consider what features and models to use, especially if you don't have a lot of data.
Overfitting is somewhat of an overloaded term. People often use it to describe the related process of creating models after you have looked at past results (e.g. models which can correctly "predict" the outcomes of all past presidential elections), and also in a more technical sense of fitting a parabola to 3 points. These are technically related, but I think it would be clearer to have two distinct terms for them.
- stdbrouw 10y ago> These are technically related, but I think it would be clearer to have two distinct terms for them. "Fishing" and "researcher degrees of freedom" are two terms I hear a lot in reference to fitting models in a very data-dependent way.
- selectron 10y agoIt is interesting how different fields have different terms for statistics concepts. Statistics really should be taught at the high school level, it is far more useful than for instance calculus. I hadn't heard those terms before. In particle physics we have the "look-elsewhere effect" as a synonym for fishing, and discuss local vs global p-values (which might be similar to researcher degrees of freedom).
- stdbrouw 10y agoThat is interesting! Re: researcher degrees of freedom, it's not really about multiple comparisons but about the fact that as an analyst you can make lots and lots of choices about how to construct your model that, individually, might well be defensible, but that ultimately end up making your model very data-dependent. You see some outliers and you remove them, you see some nonlinearities so you analyze the ranks instead of the raw data, you don't find an overall effect but you do find it in some important subgroups which then becomes the new headline, and so on and so on. At no point was anything you did unreasonable, but the end result is still something that won't generalize. A wonderful article about the phenomenon: http://www.stat.columbia.edu/~gelman/research/unpublished/p_hacking.pdf http://www.stat.columbia.edu/~gelman/research/unpublished/p_...