7 ms·
Holes in Bayesian Statistics
- remarkEon 7y agoUpvoted this mostly because I love seeing these debates here on HN about statistical methods. I don't do this kind of work in my day job, unfortunately, but studied it in college and always look back to a certain "fork in the road" moment that, had I adjusted my priors differently (heh), would've very likely led me to an academic life instead of in business.
- WhompingWindows 7y agoSeems like a weak post, very short on evidence and reasoning for such a massive topic.
- mbil 7y agoThe post is just the abstract for a paper, which is linked in the post. Here's the paper: http://www.stat.columbia.edu/~gelman/research/unpublished/bayes_holes_2.pdf http://www.stat.columbia.edu/~gelman/research/unpublished/ba...
- currymj 7y agothe first two words are a hyperlink to a much longer paper.
- tgflynn 7y agoThat definitely could have been made clearer. I looked all over the page for a link to an actual paper and missed it. I never would have guessed it was the link on what appeared to me to be an author's name.
- throwawayjava 7y agoThe linked paper is a very nice overview. Of course there problems are known and there are people trying to fix all of the issues (mostly in the relative obscurity of non-overhyped corners of academia), but the concise example-guided description of these problems is great. Somehow I think the most fundamentally damning critique, and causality shares this problem, is also the most vague. That applied scientists/experimentalists look at the "automation" that these approaches are supposed to enable and say "that's either doing the trivial part of the job or giving you BS answers".
- unishark 7y agoI've always felt bayesian statistics got more attention from researchers than was warranted (including today despite deterministic methods taking over the world) because it has a nice principled "theory of everything" starting point. But then of course you have to approximate the heck out of it to be able to solve it. Often far more than with other methods.
- steerablesafe 7y agoYou have to approximate the heck out of the Schrödinger equation as well, otherwise we would be stuck describing the Hydrogen and maybe the Helium atoms and nothing more.
- unishark 7y agoYes, however as I said in the subsequent sentence, the approximation is "often far more than other methods". For the most obvious example, a point estimate like MAP doesn't need to compute the denominator in Bayes law. That's two (generally easier) terms to approximate rather than three. Those using Bayesian methods point out the value in providing a full distribution, but the necessary additional approximations to get it means the location of its maximum can now actually be less accurate than a simple MAP estimate. But what always bugged me is multivariate problems where the Bayesian paper presumes everything is independent and Gaussian. Great after getting all psyched by that intro talk about the value of getting a distribution, we get the simplest imaginable one, just a mean and variance for each variable.
- mbil 7y ago...the inferential procedure of Bayesian statistics is to assume a prior distribution and a probability model for data and then use probability theory to determine the posterior. But if these steps, or something approximating them, are necessary, if you can’t just look at your data and come up with a subjective posterior distribution, then how is it reasonable to suppose that you could able to come up with an unassailable subjective distribution before seeing the data? Is the point to end up with an "unassailable subjective distribution"? I believe the power of Bayesian thinking is that you can take a subjective prior, which is necessarily assailable in its subjectivity, and then combine it with data. The result is something that is better than either the subjective prior alone or the likelihood estimator gleaned from data alone.
- skat20phys 7y agoI haven't read the linked paper yet, but the blog post points a little to what's on my mind. To borrow a bit from a different paper and anonymous reviewer on one of my papers, inference serves different aims or philosophies. Sometimes it serves more of an estimation function, to increase information about some quantity, or to improve the estimate of that quantity. But sometimes it serves an evaluative, competitive function, in the Popperian sense of affording risky tests of one or more theories or models. In this latter Popperian aim of inference, priors are to be minimized, which is in many ways the opposite scenario that is assumed with standard subjective Bayesian methods. And even with the former "estimation" inferential aims, there may be situations where you truly have no information or don't feel comfortable assuming it. What's nice about Bayesian statistics is it still provides a framework for this scenario, in the form of reference priors, in that if nothing else your design and model supporting the parameter(s) to be inferred about implies some kind of assumptions about what you're making inferences about. That in turn can be transformed into a "least informative" prior. So it allows an objective Bayesian framework. However, in that framework, in many cases you're still often left with uniform priors, which then reduce to frequentist statistics. And in a broader sense, frequentist methods are even further removed from making assumptions in that they completely eliminate the prior from inferential consideration. There's a tension then, in that in small samples your priors will bias your estimates. If you use least informative priors, you're often doing something akin to frequentist methods anyway. And in large samples the likelihood dominates the posterior so it matters still less. From a certain perspective, ultimately with Bayesian methods you're making a bet that your priors are accurate enough that the increased bias in estimates will be small enough to be offset by decreased variance due to use of a prior. It's a gamble though, the risks of which will probably vary depending on the costs and benefits of different types of error. It's nice to see a paper trying to be honest about the problems with Bayesian inference, as it's a bit overhyped at the moment imho.
- magoghm 7y agoFrom the the paper linked in the post: "A probability model is a tool for learning, not a suicide pact."
- astrophysician 7y agoEvery time I see an article like this, I eagerly read it hoping to find a cogent and coherent criticism of Bayesian stats, and it always ends up being a straw man or a very fair critique of somebody not correctly applying or interpreting an application of Bayes’ theorem. Bayesian stats is nothing more than a rigorous way to transform beliefs + data into a posterior. Yes, flat priors are not always the best choice. That’s not a criticism of Bayesian stats. It’s a statement about how actually formulating a prior is often times the hardest part of a problem. Is Bayesian stats useful for describing or understanding QM? Idk, again, not really it’s job... Use Bayesian stats, not with an air of suspicion, but a respect for the fact that it will give you the results implied by your data and prior, under your model assumptions. Nothing more, nothing less. By the way, what is the alternative if you find yourself in a situation where your result depends strongly on your prior and you aren’t really sure how to choose your prior? Wave your hands and find an ad hoc frequentist approach? How about just admitting to yourself that your data isn’t enough to make up for the fact that you can’t really quantify your true prior belief? If you read this and disagree, I sincerely implore you to comment — I just don’t understand the “debate” aspect of Bayesian stats. Some people misunderstand what Bayesian stats is sometimes (very understandable) but I have yet to see a legitimate philosophical or mathematical critique of the Bayesian approach that really made any sense to me. I would like to know if I am wrong though...
- jedbrown 7y ago> it always ends up being a straw man or a very fair critique of somebody not correctly applying or interpreting an application of Bayes’ theorem. I can't tell from your comment if you're aware that Andrew is the first author of the leading textbook on Bayesian methods. http://www.stat.columbia.edu/~gelman/book/ http://www.stat.columbia.edu/~gelman/book/
- astrophysician 7y agoThat’s awesome, and maybe I misunderstand his argument (very possible, in fact my main reason for posting this comment). Am I not understanding the arguments being made here? I have not been following these developments regarding using Bayesian stats in QM, it’s just that if Bayesian stats doesn’t help us understand QM it doesn’t mean there’s a “hole” because Bayesian stats makes no claim whatsoever that it has any applicability to QM. Am I just not understanding?
- blt 7y agoI'm turned off by the invocation of quantum mechanics here. Most applications of Bayesian statistics have nothing to do with QM, so "it's not compatible with QM" doesn't seem like a strong argument. QM is the number-one favorite topic of crank scientists and pseudoscience. When using it in the discussion of a seemingly-unrelated topic, authors should take extra care to motivate why QM is relevant. (Of course probability theory and QM are not unrelated topics, but probability exists independently of QM.)
- pc2g4d 7y agoI think they're assessing how universally applicable Bayesian reasoning is, and think they're seeing a limit in the quantum realm.
- syntonym2 7y agoWhile reading the blog post/abstract I had the same thought, but the full paper makes it clear that the author intends something different: > The second challenge that the uncertainty principle poses for Bayesian statistics is that [...] we routinely treat the act of measurement as a direct application of conditional probability. Furthermore it states that this problem might also arise for other applications of Bayesian statistics: > If classical probability theory needs to be generalizedto apply to quantum mechanics, then it makes us wonder if it should be generalized for applicationsin political science, economics, psychometrics, astronomy, and so forth. It’s not clear if there are any practical uses to this idea in statistics, outside of quantum physics. For example, would it make sense to use “two-slit-type” models in psychometrics, to capture the idea that asking one question affects the response to others?
- mikekchar 7y agoOne of the comments at the site finally cracked a little bit of light for me on Bell's theorem. To quote from the comment: "What his theorem says is that the world can not simultaneously be local and have hidden variables. His own position was that it seemed exceedingly likely that the world was non-local and that there were hidden variables (such as where is the photon at any given time)." I don't know why, but I kept imagining non-local scenarios and thinking, "This seems like it should be OK, so I don't understand what's going on". Having it spelled out is tremendously helpful. I still don't understand what's going on, but at least I don't feel completely crazy ;-)