4 ms·
A lot of what they discuss is in the literature on reference priors, if not in other literature on objective priors as well. It's a little complex for a commen
by ta1929901 9y ago
A lot of what they discuss is in the literature on reference priors, if not in other literature on objective priors as well.
It's a little complex for a comment on HN, but IMHO the best formalization of overfitting is in the literature on minimum description length, and related information-theoretic literature (https://en.wikipedia.org/wiki/Minimum_description_length https://en.wikipedia.org/wiki/Minimum_description_length ; the wikipedia page is a little off on some things but the general points are probably about right).
The relationship between MDL/NML and Bayesian statistics is a little complicated--Barron, Roos, and Watanabe have a nice recent paper about it (https://arxiv.org/abs/1401.7116 https://arxiv.org/abs/1401.7116) -- but the short story is that there's a certain equivalence (or at least very close relationship) between MDL/NML and Bayesian inference with reference priors (i.e., the "capacity-achieving prior" in IT parlance).
So Bayesian inference with reference priors minimizes risk of overfitting in a very technical minimax sense.
The problem is that reference priors are very difficult in general to construct, although that is changing rapidly, and they have been worked out for certain important cases (this paper provides an interesting new approach with a nice overview of recent papers https://arxiv.org/abs/1704.01168 https://arxiv.org/abs/1704.01168). The Jeffreys prior is a form of the reference prior for models meeting certain constraints.
A lot of the issues Gelman et al. touch on have been written about in various places in the MDL/NML/reference prior/information theory literature.
- Xcelerate 9y ago> but IMHO the best formalization of overfitting is in the literature on minimum description length It's always surprised me how few people know about MDL, considering it's pretty darn close to what we might consider a "universal predictor" (granted, the difficulty with MDL is that it's uncomputable in the general sense). Even among most data scientists I know, very few understand what overfitting really is (and thus cross-validation and regularization are merely tools to what seems like the blackbox "art" of model selection). But the concepts of MDL/MML and Kolmogorov complexity are very deep and fundamental—to such an extent that I think the path to true AGI will rely much more heavily upon algorithmic information theory than neural networks in the future.
- ta1929901 9y agoYeah, I have a similar reaction about being surprised MDL and algorithmic statistics isn't better known. As you say, it's very fundamental stuff. It seems so fundamental to me that I just sort of assume without thinking about it or even questioning whether it will eventually become more prominent. My guesses as to why it's been slow to be adopted so far are that (1) it's relatively new, in the grand scheme of things, (2) certain things about it are really challenging to everyone, and (3) it has a certain perspective on inference that can be alien to a lot of people.
- sgt101 9y agoAlso (5) it doesn't really work or make sense. The data isn't necessarily representative of the domain theory, in fact in a lot of domains the data isn't, because the domains are so large that you can't capture the whole of it in a tractable training set, for example : images. Other data doesn't capture the domain because the data is generated in a regime that isn't operant when gathered - for example bull runs vs. bear runs in the markets. Bayesian analysis is attractive, we can include informative priors that capture our knowledge that in circumstances outside of the data other determining behaviours exist. This is also one of the reasons why deep networks can outperform support vector machines; deep networks can learn to prefer domain theories that are not the minimal descriptive one. The other thing that is interesting about MDL is where does the idea that the minimal theory is the right theory come from? Most people say "oh it's Occam's razor" but where did Occam's razor come from - who was Mr (Fr.) Occam?? Well, he was a 13th century philosopher - part of the Cambridge school and part of the tradition of Scotus invented to construct a story that supported the Trinitarian God... and this is why we prefer the idea that "entities will not multiply beyond necessity" because it says that you have a Trinity because The University ABSOLUTELY cannot work without it, and that's why you have three and not two and not four. I am happy with all this but why should we think it's a good way to do machine learning? After all there are lots of examples of theories that were simple but don't work as well as complex alternatives - Gravity is a good one.