3 ms·
When are Bayesian methods "clearly worse" than frequentist methods, apart from computationally?
by ced 14y ago
When are Bayesian methods "clearly worse" than frequentist methods, apart from computationally?
- equark 14y agoThere are times when, even as a Bayesian, one is interested in calibration. Model checking without a specified alternative is an example. Frequentist ideas -- sampling from the model and comparing it to the observed distribution -- can be helpful here. I'm thinking of Rubin (1984): http://www.cs.princeton.edu/courses/archive/fall11/cos597C/reading/Rubin1984.pdf http://www.cs.princeton.edu/courses/archive/fall11/cos597C/r...
- antics 14y agoMy Bayesian theory is a bit rusty, but here we go. Say we have data X, and some non-finite dimensional index into the family of functions that describe the the data, called \theta. The Bayesian perspective classically holds \theta constant and optimizes the expected loss, conditioned on the data X. The frequentist perspective, on the other hand, classically optimizes \theta, that is, it picks the best \theta over the data X, unconditionally. This has two impacts. First, all things equal frequentist statistics will tend to be more stable, and more calibrated, but less coherent. It is commonly said that frequentist statistics will "isolate" one from poor decision making, and all things equal, that will be true. Specific, clear wins for frequentists are bootstrapping procedures (e.g., Efron's bootstrap, the b of n bootstrap, Jordan's own scalable "bag of bootstraps" from NIPS 2011), which are methods for building what are called "quantifiers" for "estimators". In short, this means that if you have some estimator (e.g., a classifier, or a mean, or whatever), you want to be able to quantify the certainty of your estimator -- so if you've only seen 5 examples, you want to express that you're less certain about this. This is clearly a frequentist application, not a Bayesian application, and in general, it points to the fact that pure frequentist tools not only have a place in inference, but they fill a niche that Bayesian tools necessarily will not, and in some cases, cannot, fill.
- equark 14y agoIt sounds like you're just saying that if you want to know the frequentist properties of an estimator you have to be frequentist. That's a tautology. The harder question is whether there are any decisions you'd prefer to make using a non-Bayesian procedure. That's basically a tautology in the other direction though.
- antics 14y agoAs I said, my Bayesian theory is rusty, but there are no "frequentist properties" of an estimator. Frequentist inference is inference -- it doesn't make guarantees about the underlying thing it's approximating, it provides guarantees about its approximation. The key here is that Bayesian and frequentist procedures provide different sorts of guarantees. Frequentists optimize for \theta, the possible set of things that could describe all of the data X, while Bayesians will assume a single describing function \theta (this might come from an "expert") and simply optimize the expectation conditioned on the data. Neither is "wrong" but in the case of the bootstrap, the result is calibrated in a way that Bayesian inference simply never will be (if it were, it would be frequentist). EDIT: As per your second question, actually I think it's not more interesting. A classifier is a type of estimator, so all of the general frequentist guarantees actually still apply to decisionmaking.
- equark 14y agoFrequentist statistics is about determining the repeated sampling properties of a procedure/statistic/estimator. It's about evaluation not estimation. "Optimizing \theta" or whatever you're envisioning is just one possible procedure you might be interested in evaluating. You can use the repeated sampling properties of your procedure to do frequentist inference or to evaluate other properties like unbiasedness, consistency, risk, etc. Typically the goal is to find procedures that have "good" frequentist (repeated sampling) properties. Most Bayesian-inclined statisticians would tend to argue that many frequentist properties are not important to applied data analysis or optimal decision making.
- taliesinb 14y agoA possible example given in the lecture is figuring out if two sets of numeric data were sampled i.i.d from the same distribution (being hand-wavy about what this precisely means). I don't see how you can sensibly approach this kind of problem in a Bayesian way. From a frequentist perspective you're sort of spoiled for choice about how to approach this problem.
- mjw 14y agoThe useless answer is that they both do different things, so it depends what of those things you want :) One aspect of frequentist techniques that perhaps others haven't emphasised so much, is that they tend to give guarantees about expected behaviour which hold uniformly over all possible values of the unknown parameters. Whereas the Bayesian approach, the guarantees you obtain will only hold in an 'averaged-out' sense over the prior distribution you specify. If you're a bit paranoid and you want a probabilistic bound on what might happen in the worst case, you might sometimes find the former a little more comforting than the latter. In particular if you don't have much data, the influence of the choice of prior will be bigger and so the distinction will matter more. Hope that helps, and that any stats PhDs will correct me if I've over-simplified things here.
- wamatt 14y agoThis is an attempt to give a pragmatic overview, without the math. Bayesian methods rely on the concept of "priors". A prior is a known probability about a fact, or event, or whatever it is you are modeling. Priors "seed" the network. Bayesian networks generally need fewer samples of data to make predictions, and as a result the downside is the sensitivity of the data (aka priors) increases. Whereas frequentist approaches rely on much larger data sets that can handle noise, more effectively. Think about GMail's spam filter for example (A Bayesian approach), if you train HAM as SPAM, that is going to have devastating effect on the efficacy of your filtering. Thus practically speaking, if you have tons of data and looking for a signal, use frequentist. If you're building, for example, an expert diagnostic engine (think Sherlock Holmes), requiring few pieces of information to make a prediction, consider using Bayesian. I'm over-simplifying of course, but that seems to be the gist of it.
- othermaciej 14y agoI don't think that's right. If you have enough data points, your prior gradually gets more relevant. And Bayesian statistics has the concept of an "ignorance prior", which mathematically represents the position where all possibilities are equally likely. Bayesian statistics also offers the possibility of adding more data after running your experiment, and computing a new answer in a consistent way. Whereas doing this with frequentist statistics completely is completely invalid.
- wamatt 14y ago"Ignorance priors" are not the common use case for Bayesian in general, at least in my experience. Furthermore, who said adding more data is not allowed? Of course it is, and I agree fully, but do struggle to see how that contradicts what was said. Think about a Bayesian spam filter, you keep training it with more "priors", over time and it gets better.