3 ms·
Good to see Bayesian model selection get a mention. Bayesian model averaging is pretty interesting, too, in that it comes, in a sense, with built-in protection
by mjw 13y ago
Good to see Bayesian model selection get a mention. Bayesian model averaging is pretty interesting, too, in that it comes, in a sense, with built-in protection against overfitting.
I still think there is something quite fundamental, though, about validation sets and other related resampling-based methods for estimating generalisation performance (cross-validation, bootstrap, jackknife and so on).
The built-in picture you get about predictive performance from Bayesian methods comes with strong caveats -- "IF you believe in your model and your priors over its parameters, THEN this is what you should expect". Adding extra layers of hyperparameters and doing model selection or averaging over them might sometimes make things less sensitive to your assumptions, but it doesn't make this problem go away; anything the method tells you is dependent on its strong assumptions about the generative mechanism.
Most sensible people don't believe their models are true ("all models are false, some models are useful"), and don't really fully trust a method, fancy Bayesian methods included, until they've seen how well it does on held-out data. So then it comes back to the fundamentals -- non-parametric methods for estimating generalisation performance which make as few assumptions as possible about the data and the model they're evaluating.
Cross-validation isn't the only one of these, and perhaps not the best, but it's certainly one of the simplest. One thing people do forget about it is that it does make at least one basic assumption about your data -- independence -- which is often not true and can be pretty disastrous if you're dealing with (e.g.) time-series data.
- ced 13y agoI agree. As a Bayesian hoping to understand my data, P(X|M1) is useful: it's the probability I have for X under M1's modelling assumptions. Of course M1 is an approximation, but that's how science is done. You get to understand how your model behaves, and you may say "Well, X is a bit higher than it should be, but that's because M1 assumes a linear response, and we know that's not quite true". Bayesian model averaging entails P(X) = P(X|M1)P(M1) + P(X|M2)P(M2). It assumes that either M1 or M2 is true. No conclusions can be derived from that. It might be useful from a purely predictive standpoint (maybe) , but it has no place inside the scientific pipeline. There is a related quantity which is P(M1)/P(M2). That's how much the data favours M1 over M2, and it's a sensible formula, because it doesn't rely on the abominable P(M1) + P(M2) = 1
- mjw 13y agoYeah good perspective -- I guess I was thinking about this more from the perspective of predictive modelling than science. Model averaging can be quite useful when you're averaging over versions of the same model with different hyperparameters, e.g. the number of clusters in a mixture model. You still need a good hyper-prior over the hyperparameters to avoid overfitting in these cases though, as an example IIRC dirichlet process mixture models can often overfit the number of clusters. Agreed that model averaging could be harder to justify as a scientist comparing models which are qualitatively quite different.
- ced 13y agoModel averaging can be quite useful when you're averaging over versions of the same model with different hyperparameters, e.g. the number of clusters in a mixture model. Yeah, but in this case, there's a crucial difference: within the assumptions of a mixture model M, N=1, 2, ... clusters do make an exhaustive partition of the space, whereas if I compute a distribution for models M1 and M2, there is always M3, M4, ... lurking unexpressed and unaccounted for. In other words, P(N=1|M) + P(N=2|M) + ... = 1 but P(M1) + P(M2) << 1 Is the number of clusters even a hyperparameter? Wiki says that hyperparameters are parameters of the prior distribution. What do you think?
- avaku 13y agoGreat explanation. I would like to add to this, that held-out data is often used in Bayesian learning too - for example, in cases when you intentionally over-specify the model (adding more parameters than might be needed) because you don't really know what the best model might be. The inference goes until the likelihood on held-out data keeps increasing. Example, gesture recognition in Kinekt. If someone finds this info useful, I also recommend Coursera course on Probabilistic Graphical Models.
- mailshanx 13y agoWhat are some good resources to understand Bayesian model averaging?
- mjw 13y agoThese slides have a bit on this (although quite dense material): http://www.gatsby.ucl.ac.uk/teaching/courses/ml1-2011/lect5b-handout.pdf http://www.gatsby.ucl.ac.uk/teaching/courses/ml1-2011/lect5b... as part of http://www.gatsby.ucl.ac.uk/teaching/courses/ml1-2011.html http://www.gatsby.ucl.ac.uk/teaching/courses/ml1-2011.html I quite like "Bayesian reasoning and machine learning" too: http://web4.cs.ucl.ac.uk/staff/D.Barber/textbook/090310.pdf http://web4.cs.ucl.ac.uk/staff/D.Barber/textbook/090310.pdf