4 ms·
Nice article, in the sense that I would assign it to students, although I think it presents sort of a strawman and doesn't really introduce anything that hasn't
by ta1929901 9y ago
Nice article, in the sense that I would assign it to students, although I think it presents sort of a strawman and doesn't really introduce anything that hasn't been discussed at length elsewhere.
It provides a nice introduction to types of objective priors, hinting at their advantages and disadvantages. Also a nice discussion of why priors matter. Incidentally, I agree that "subjective" and "objective" are poor labels for priors--something better would be something like "estimand-predictive" and "inferential-property" priors, respectively.
I'm not really sure why they suggest objective priors violate assumptions about not using information from the data. The model, and thus many objective priors, can be specified given the design, without any knowledge of what the data looks like. This is the strawman part of it.
I get the sense Gelman has been wrestling with objective priors the last few years.
- confounded 9y agoHe (and the Stan team) have been my main source of 'default priors' where I don't have a strong opinion, and am mainly looking for regularization-style properties (most of the time).
- nerdponx 9y agoMany of the Gelman-style default "objective" priors aren't so hard to interpret as "I am skeptical that this parameter is non-zero".
- nonbel 9y agoGelman has been pretty adamant about the idea there probably always is some real correlation between almost any two measurements (that is why he advocates his type M and type S error concepts). Perhaps you are confusing "skeptical the parameter is far from zero" with "skeptical the parameter is non-zero". It makes a big difference.
- sgt101 9y agoOdd because I read the abstract of the paper as a claim for a fundamental advance in statistics; systematic and objective determination of priors in Bayesian analysis. I can't speak to whether the paper is successful in delivering this as I am very dense and it takes me weeks to understand things, but if they have identified mechanisms that allow complex priors to be constructed that minimize overfitting: 1. I will do a little dance 2. I will learn how to construct such priors 3. I will attempt to apply this in practice to see what happens. I don't see overfitting risk as a strawman, I see it as a nasty business that means that production models can't be trusted.
- ta1929901 9y agoA lot of what they discuss is in the literature on reference priors, if not in other literature on objective priors as well. It's a little complex for a comment on HN, but IMHO the best formalization of overfitting is in the literature on minimum description length, and related information-theoretic literature (https://en.wikipedia.org/wiki/Minimum_description_length https://en.wikipedia.org/wiki/Minimum_description_length ; the wikipedia page is a little off on some things but the general points are probably about right). The relationship between MDL/NML and Bayesian statistics is a little complicated--Barron, Roos, and Watanabe have a nice recent paper about it (https://arxiv.org/abs/1401.7116 https://arxiv.org/abs/1401.7116) -- but the short story is that there's a certain equivalence (or at least very close relationship) between MDL/NML and Bayesian inference with reference priors (i.e., the "capacity-achieving prior" in IT parlance). So Bayesian inference with reference priors minimizes risk of overfitting in a very technical minimax sense. The problem is that reference priors are very difficult in general to construct, although that is changing rapidly, and they have been worked out for certain important cases (this paper provides an interesting new approach with a nice overview of recent papers https://arxiv.org/abs/1704.01168 https://arxiv.org/abs/1704.01168). The Jeffreys prior is a form of the reference prior for models meeting certain constraints. A lot of the issues Gelman et al. touch on have been written about in various places in the MDL/NML/reference prior/information theory literature.
- Xcelerate 9y ago> but IMHO the best formalization of overfitting is in the literature on minimum description length It's always surprised me how few people know about MDL, considering it's pretty darn close to what we might consider a "universal predictor" (granted, the difficulty with MDL is that it's uncomputable in the general sense). Even among most data scientists I know, very few understand what overfitting really is (and thus cross-validation and regularization are merely tools to what seems like the blackbox "art" of model selection). But the concepts of MDL/MML and Kolmogorov complexity are very deep and fundamental—to such an extent that I think the path to true AGI will rely much more heavily upon algorithmic information theory than neural networks in the future.