4 ms·
There has been a lot of rich discussion about the relative merits of the Bayesian and frequentist perspectives on statistical inference. If you are thinking abo
by antics 14y ago
There has been a lot of rich discussion about the relative merits of the Bayesian and frequentist perspectives on statistical inference. If you are thinking about applying inference techniques to some problem, then it is well worth your time to make sure that you really, really understand this debate, because picking the correct tool for your job is likely to make your life a lot easier (and, yes, there are both situations where the Bayesian perspective is clearly better and where it is clearly worse).
Unfortunately this post completely and totally ignores the things that will enable you to make this decision. This post is called "Understand the Math Behind it All", but you will learn nothing about math, or really, Bayesian statistics, at all. You will not learn, for example, how to apply Bayesian inference to problems, what it means to do basic Bayesian tasks like "use an expert" or "condition on evidence", or even what a Bayesian statistic is. In fact, Bayes' theorem is never even mentioned. There's just a hand-wavy collection of statements like "The Bayesian approach is to rely on past knowledge and then adjust accordingly". That is so vague that it is not even clear that they are talking about Bayesian analysis. This is the sort of statement that fools people into believing they understand something that they really don't.
If you really want to understand this material, you should watch[1], a talk by Mike Jordan called "Are you a Bayesian or a Frequentist?". It's a bit much for beginners, but if you are willing to look up some of the math, it is entirely digestible, and it is by far the best comparison of the two communities I have found. I say "by far" because it is (1) a more or less complete representation of both communities, (2) it is pretty much an unbiased representation of both communities' strengths and weaknesses, and (3) it is as direct as it can get, meaning that it is not tied up in a lot of external knowledge, and is intent on delivering this message, rather than delivering it in an off-hand way as a method of getting to something else.
[1] http://videolectures.net/mlss09uk_jordan_bfway/ http://videolectures.net/mlss09uk_jordan_bfway/
- ced 14y agoWhen are Bayesian methods "clearly worse" than frequentist methods, apart from computationally?
- equark 14y agoThere are times when, even as a Bayesian, one is interested in calibration. Model checking without a specified alternative is an example. Frequentist ideas -- sampling from the model and comparing it to the observed distribution -- can be helpful here. I'm thinking of Rubin (1984): http://www.cs.princeton.edu/courses/archive/fall11/cos597C/reading/Rubin1984.pdf http://www.cs.princeton.edu/courses/archive/fall11/cos597C/r...
- antics 14y agoMy Bayesian theory is a bit rusty, but here we go. Say we have data X, and some non-finite dimensional index into the family of functions that describe the the data, called \theta. The Bayesian perspective classically holds \theta constant and optimizes the expected loss, conditioned on the data X. The frequentist perspective, on the other hand, classically optimizes \theta, that is, it picks the best \theta over the data X, unconditionally. This has two impacts. First, all things equal frequentist statistics will tend to be more stable, and more calibrated, but less coherent. It is commonly said that frequentist statistics will "isolate" one from poor decision making, and all things equal, that will be true. Specific, clear wins for frequentists are bootstrapping procedures (e.g., Efron's bootstrap, the b of n bootstrap, Jordan's own scalable "bag of bootstraps" from NIPS 2011), which are methods for building what are called "quantifiers" for "estimators". In short, this means that if you have some estimator (e.g., a classifier, or a mean, or whatever), you want to be able to quantify the certainty of your estimator -- so if you've only seen 5 examples, you want to express that you're less certain about this. This is clearly a frequentist application, not a Bayesian application, and in general, it points to the fact that pure frequentist tools not only have a place in inference, but they fill a niche that Bayesian tools necessarily will not, and in some cases, cannot, fill.
- equark 14y agoIt sounds like you're just saying that if you want to know the frequentist properties of an estimator you have to be frequentist. That's a tautology. The harder question is whether there are any decisions you'd prefer to make using a non-Bayesian procedure. That's basically a tautology in the other direction though.
- epistasis 14y agoThat's a truly excellent talk and really even the first slide, titled "Statistical Inference," should be enough to gain a ton of information, so if you're intimidated by the length just give the first slide a try. Michael Jordan is one of my favorite statisticians/machine learning researchers, and if I see he's speaking somewhere I always try to go. I don't go to many stats talks, but his talks are always some of the most mathy I do see, and he's not afraid to dive in to the mathematical mechanics of methods.
- mjw 14y agoVery much agree. I'd also prefer not to encourage referring to people as "a Bayesian" or "a frequentist", as though it's a hard philosophical preference, a binary choice one has to make. I realise there are historical reasons for this, but really guys, they're both tools. Know the pros and cons and pick the right one for the job; neither is uniformly better. Which framework you use can be a fairly subtle decision sometimes, requiring you to think fairly deeply about what you really want to get out of the analysis, what assumptions you're comfortable making, what kind of interpretation it would be most useful to be able to place on the results. But it's not just an arbitrary choice that you can leave to aesthetic or philosophical preference.
- kylebrown 14y agoThanks. Just yesterday I came across this first paragraph in a paper I now see is by Jordan: "Statistics has both optimistic and pessimistic faces, with the Bayesian perspective often associated with the former and the frequentist perspective with the latter, but with foundational thinkers such as Jim Berger reminding us that statistics is fundamentally a Janus-like creature with two faces." [Janus is a Roman god with two faces] In my current project (stroke and [since the leap] gesture recognition), I'm using the covariance matrix of a training set and the difference a feature vector from the mean to calculate the mahalanobis distance of the input vector from the training set. I plan to use this same covariance matrix in the Gaussian density estimation formula to generate a probability distribution function (and then use the likelihood function instead of mahalanobis distance). Still trying to mentally connect this to the bigger-picture stuff though. http://www.cs.berkeley.edu/~jordan/papers/berger-festschrift.pdf http://www.cs.berkeley.edu/~jordan/papers/berger-festschrift...
- antics 14y agoBerger's actually written a seminal book on the topic called "Statistical Decision Theory and Bayesian Analysis"[1]. If you're interested in that area, consider checking it out of your library. I'm not quite as sharp in the area as I used to be, but feel free to hit me up over email if you have more questions, I can't guarantee that I'll know the answers, but I'm happy to give it a shot. [1] http://www.amazon.com/Statistical-Decision-Bayesian-Analysis-Statistics/dp/0387960988 http://www.amazon.com/Statistical-Decision-Bayesian-Analysis...
- kylebrown 14y agoThanks again! I'll do that if I form a coherent question.
- carbocation 14y agoStatsexchange is actually a great resource as well if you form that question: http://stats.stackexchange.com/ http://stats.stackexchange.com/
- taliesinb 14y agoThat's a cool lecture. Looks like it was part of the Cambridge Machine Learning Summer School 2009 (http://videolectures.net/mlss09uk_cambridge/ http://videolectures.net/mlss09uk_cambridge/), which has a bunch of cool talks, for example Josh Tenenbaum's "Machine Learning and Cognitive Science" (http://videolectures.net/mlss09uk_tenenbaum_mlcs/ http://videolectures.net/mlss09uk_tenenbaum_mlcs/)
- antics 14y agoOne of the cooler things about the field of machine learning is that conferences like ICML, and little mini-summer schools like this are very "rich", so talks like this are all over the Internet. For those who are curious Tenenbaum has recently gotten very famous in the community for Bayesian style work. His students are getting very very prestigious jobs, so a talk like that is worth a look for sure.