4 ms·
Don't think that's quite what the author was getting at. A "frequentist" who makes a decision based on a p-value or confidence interval, and a "Bayesian" who ma
by keithwinstein 12y ago
Don't think that's quite what the author was getting at. A "frequentist" who makes a decision based on a p-value or confidence interval, and a "Bayesian" who makes a decision based on a posterior or predictive probability distribution, can both be viewed as making a decision according to a procedure that minimizes some statistic of a loss function.
Both of these broad families of methods are diverse, but to generalize broadly, decisions informed by "frequentist" tools are often concerned with the worst-case value of the loss function within some universe, and decisions based on "Bayesian" methods can be viewed as caring about the expected value of the loss function given some prior.
Both of those families (and many others) are "the ground truth," in that they both make statements that can be proved as theorems of mathematics. A "frequentist" confidence interval will always achieve its guaranteed coverage no matter what, if the likelihood function is true. A "Bayesian" credibility interval will include the true value of the parameters at exactly the advertised rate, when averaged over each possible value of the parameters weighted according to the prior, and assuming the likelihood function is true.
The author says, and I agree, that the thing worth arguing over is the "norm we use to choose our optimal procedure" and the nature of the loss function. Whether you care about controlling the worst case or the expected value or something else depends on what matters. (The author points out that in an adversarial situation where your strategy is known to your opponent, minimizing the worst case may be advisable...)
Some days, we care about the worst-case performance of QuickSort, and some days we care about its average-case performance (given an assumption that all input orderings are uniformly likely). It's ok to care about different things depending on the application; we don't have to split into warring tribes over it.
More here: http://blog.keithw.org/2013/02/q-what-is-difference-between-bayesian.html http://blog.keithw.org/2013/02/q-what-is-difference-between-...
http://jsteinhardt.wordpress.com/2014/02/10/a-fervent-defense-of-frequentist-statistics/ http://jsteinhardt.wordpress.com/2014/02/10/a-fervent-defens...
http://cs.stanford.edu/~jsteinhardt/stats-essay.pdf http://cs.stanford.edu/~jsteinhardt/stats-essay.pdf
- Matumio 12y agoDo Bayesian methods imply a choice of loss function? I think it's a framework for observational statistics, not for policy making or decision theory. You can use it as a tool during policy or decision making, where you should indeed have a debate about picking a loss function that corresponds to your goals.
- mjw 12y agoIf you're reporting point or interval estimates (rather than the entire posterior) then you are implicitly or explicitly optimising some kind of loss function. Also, worth a reminder that continuous estimation problems can be viewed as decision problems too, it's not just about discrete decision-making. I don't think it's great to take the view that, because I'm not making a decision based on this estimate myself, I don't need to worry about which loss function I'm implicitly optimising for when choosing an estimator. Someone else may need to.
- jules 12y agoBayesian statistics gives you a posterior distribution. What you do with that distribution is up to you. If you want to find the decision that minimizes the maximum loss instead of the expected loss, that fits perfectly well within the Bayesian framework. The posterior gives you all the information that you need to make a decision, whatever your loss function. Frequentist statistics on the other hand gives you no such thing. You can't make decisions based on frequentist statistics. Frequentist statistics is all about reasoning about things that did not happen. Things that did not happen are irrelevant for making decisions in a situation where that thing did happen. So while frequentist statistics is mathematically correct, it's also strictly speaking useless in practice. It's only useful insofar as it gives us heuristics for decision making when the Bayesian approach is impractical. Here's an example. Suppose there are two types of berries: edible and poisonous. We devise a statistical procedure where you measure some properties of the berry (lets say we look at the color), and the procedure should help you decide whether to eat that berry. Now in frequentist statistics, you'll get a procedure that gives you the correct answer with probability at least p regardless of what your measurement was. Suppose that p=90% and we observe that the color is blue, and the procedure says: this berry is edible. Can we now eat the berry? No! This says absolutely nothing about the probability of the blue berry being poisonous or not. For example suppose 95% of berries are edible and red, and 5% of the berries are poisonous and blue. Then a valid procedure would be one that says that all berries are edible. It's correct >90% of the time, so yay! But if the berry we are holding in our hands is blue, it's incorrect 100% of the time. The fact that the procedure would have given us the right answer if the berry we found was red is irrelevant for making the decision of whether to eat the berry in the situation where the berry we found was blue. Things that did not happen are irrelevant for making decisions in a situation where that thing did happen! tl;dr: frequentist vs bayesian is not about worst case vs average case. It's about P(measurement | true value) vs P(true value | measurement). The former is irrelevant for making decisions, the latter is exactly what you want.
- deleted 12y ago[deleted]
- mjw 12y agoThis isn't just about a difference in the choice of loss function to optimise. It's a difference in what sort of guarantees you seek about that loss function. Bayesian analysis seeks an estimator which minimises posterior expected loss, conditioning on the data and with the expectation taken over the parameters under a particular prior. A frequentist analysis might seek an estimator for which uniform bounds on the worst-case expected loss are available, which hold in expectation over the data, given any value of the parameters. Both approaches fit into a decision theoretic framework and there are good reasons why you might care about frequentist properties when making decisions. I agree that this isn't only about average case vs worst case -- as you point out it's also about whether you take expectations over data given params or over the params given data, and that's important too. But I think the average case vs worst case aspect of this is an important part of what this is all about and gets to the heart of what the trade-offs are when choosing between these methods. I disagree that the sampling distribution is "irrelevant for making decisions", that's quite an extreme view which I don't think many applied Bayesian statisticians would take. Frequentist properties are something people often validly care about when deciding on a statistical procedure to use in an experimental design context, i.e. before collecting the data -- and especially if you're choosing an estimator which you intend to use many times for many experiments, even if they're not all exact replicates of each other.
- ellyagg 12y agoWhat is not what the author was getting at? It's clear from your comments here--continually scare quoting "frequentist" and "Bayesian"--and your answer on Quora, that you're contemptuous of this debate, like the author. OP is saying that there's more to frequentist vs Bayesian than simply picking the right tool for the job, which you are both suggesting. It won't be settled in this comments section.
- bjterry 12y agoEverything that you've written is technically correct. But it seems like in the large majority of situations where people are using/publishing statistics, the expected value is what we care about but we are using frequentist methods. So if we are going to correct the "norm we use to choose our optimal procedure," it should imply a many-fold increase in uses of Bayesian methods. A huge amount of published papers that include any statistics are reporting means and p-values in very simple statistical analyses (essentially all papers published in the social sciences, health sciences, etc.). In many of these situations researchers could provide subjective priors based on the state of the research and knowledge beforehand and it would be philosophically valid and beneficial. Because of this mismatch, people who support Bayesian methods perhaps overstate their case. Maybe this is for the same reason that partisan political actors argue fervently for a single side of an issue, ignoring nuance. Nuance historically has not led to social change, and fixing statistics in science requires changing people's minds and "raising awareness."
- keithwinstein 12y agoI hear you, and my understanding is that this largely depends on the discipline. In fields like medicine and psychology, the flourishing of classical methods in the 30s and 40s and 50s did lead to a sort of dogmatism about p-values and a disdain for the old "inverse probability" (what we now call Bayesian methods) of the 18th and 19th centuries. These fields seem to be still recovering from this. My impression is that that's where you often find self-styled Bayesian rebels with the faith of the converted. In areas like radar or image processing or communications or information theory or ad placement or AI in general, I think the field has long had a more nuanced understanding of the underlying decision-theoretic concerns. When you have to win World War II and there is a cost to falsely identifying a Nazi aircraft as Allied, and a cost to falsely identifying an Allied aircraft as Nazi, you develop a notion of an ROC curve pretty quickly. Ditto when you want to talk to a Voyager probe and you're not sure if it just sent a 0 or a 1. In my view, the Bayesian vs. frequentist "debate" has little to say to these fields. Although for a contrary view, see what Jaynes writes about Shannon in "Probability Theory: The Logic of Science" in the last chapter ("Introduction to communication theory").