7 ms·
It’s time to adopt modern Bayesian data analysis as standard procedure
- hessenwolf 16y ago1. Your list of references proves that Bayesian statisticians have been writing papers across a variety of disciplines. 2. Bayesian methods are not readily computed with today's hardware and software, and my desk is a counter-example. 2.1 Last time I fitted a Bayesian model, I had 600 processors with infinibandy things joining them up and left it for a week. 2.2 None of the software I use does Bayesian by default. 3. In practice, a lot of statistics in industry can be done by barcharts. I know that that is hard to hear. It is a big leap from there to Bayesian; bigger by far than from there to chi-sq tests. Data are so sparse, expert judgement so rich, and time so short... 4. Priors introduce subjectivity - no doubt about it. However, so do utility functions, pretty much cancelling it in my opinion. It is inappropriate to use p-values as a decision making framework for various reasons, but a lot of scientific papers are about recording experimental observations, not decision-making. Policy-decisions should use utility functions and priors; but I am happy with my science papers frequentist.
- wisty 16y ago5. Bayesian analysis is too easy. You don't have to transform the data in any convoluted way, you just describe your model and crank the handle. Publishers will no longer be able to sort the sheep from the goats.
- hessenwolf 16y agoHa ha. Yes, it is a bit like that. If somebody could just replace that pesky mcmc convergence crap with something that worked, or give me a quantum computer, then I would convert my models Bayesian overnight.
- lazyjeff 16y agoI agree with the general theme of the OP. Bayesian methods are technically better than frequentist methods. However, like Betamax vs. VHS, there is more than just the technical correctness. Your point in 2.2 is the big one -- if there was a simple way to switch to Bayesian methods in existing statistics software like SPSS, that would be quite revolutionary. Right now, null hypothesis testing is too easy to do and widely accepted, even though the results may be completely wrong.
- nazgulnarsil 16y agopriors acknowledge subjectivity.
- hessenwolf 16y agoAlso, this is true, but I think it doesn't disagree with my point. I just dread the length of my Own Risk and Solvency Assessment (ORSA) after Solvency II (the regulations for insurance companies in Europe) takes hold next year, if I have to explain the origins of my priors. I actually think we could do some fuck-awesome work in evaluating our risk capital requirements using priors on all of of our inputs, and the computational requirements would not be 'that' frightening, given the valuations are all Monte Carlo anyway. It's not happening, yet, though.
- Sniffnoy 16y agoI think about this point someone should link to [share likelihood ratios, not posterior beliefs](http://www.overcomingbias.com/2009/02/share-likelihood-ratios-not-posterior-beliefs.html http://www.overcomingbias.com/2009/02/share-likelihood-ratio...)... the summary is that you can do your Bayesian analysis without specifying the prior and just report the resulting likelihood ratios, telling everyone, "Here are the likelihood ratios, update your beliefs appropriately." Though that may lack some practicality.
- noelwelsh 16y agoCounterpoint: Last time I fitted a Bayesian model, I had 1 processor with some interpreted code and left it for ten minutes. YMMV.
- bluekeybox 16y agoAbsolutely. Not all applications of Bayesian analysis are computationally-intensive. In some cases (an example: finding of single-nucleotide polymorphisms in next-gen sequencing data), Bayesian analysis comes down to multiplying prior probability of a SNP (for humans, 0.001 per genome position) by a few other numbers from the data itself to obtain posterior probability, which can be done in a linear time in a few minutes on tens of gigabytes of NGS data. And the best part is, no Bonferroni adjustment bullshit!
- hessenwolf 16y agoHey, don't forget the Benjamini-Hochberg replacement for Bonferroni. The situation is not 'that' bad.
- tel 16y agoA point by point rebuttal. 1. The op knows this but is implicitly believing his audience is different from the cited fields often enough to make point (1) relevant. Moreover, it's perfectly correct to say that many journal reviewers are not interested in Bayesian methods. 2.1. This is highly anecdotal and not at all a strong point. I'm sure Efron has knocked out a 600 core cluster doing frequentist bootstrapping (1). I'm also certain that many, many Bayesian methods run near instantly on modern hardware. I'll concede that there are fewer closed form results, though. 2.2 This is very true. Entrepreneurs? 3. "Data so sparse, expert judgement so rich" is exactly where Bayesian analysis is most pertinent. Use a prior to clarify and quantify your expert opinion and then demonstrate that indeed your few observations are worthy to change someone's opinion. 4. Choice of frequentist testing regime introduces subjectivity, too! Moreover, since these have been heralded as "objective" for so long it's pretty difficult to get people to recognize as much. Oftentimes, a frequentist method will be equivalent to a Bayesian method under a maximally uninformed prior. This is still a subjective assumption (though there are benefits of such a prior). --- Frequentists test are oftentimes very necessary. They have already been highly optimized in many cases and thus are available on low resource computing platforms. They are definitely an important engineering solution! That said, Bayesian methods do a far better job being clear in their assumptions and simple in their logic. There is certainly room for better software (free or otherwise) to replace BUGS/JAGS/whatever for the largest use cases of statistics in many fields. Also, another point you make about Bayesian methods making life difficult during certification and publication is exactly right, and probably the largest (unspoken) reason why they're not going to be used in core scientific fields for a long while. But both of those reasons are distinctly practical and unscientific. Bayesian methods do a better job using your data. They do this by allowing expert knowledge to enter into statistics in a sensible fashion. Finally, they introduce an easily understood interpretation on the answers to your statistical questions. You might not personally want to use them today for practical reasons, but the author of this article is very much in the right to try to encourage more scientists in more fields to take a look. (1) Sorry, I'm actually not at all sure if this is the case. Bootstrapping is still more computationally efficient than MCMC, I think. I just used the example because I think it's ridiculous to make either point.
- hessenwolf 16y ago
- yannickmahe 16y agoCould somebody explain what is Bayesian analysis and how it works or point to an article explaining it ? The wikipedia article regarding Bayesian analysis is cryptic to me.
- jules 16y agoHere's a wikipedia page with an example http://en.wikipedia.org/wiki/Checking_whether_a_coin_is_fair http://en.wikipedia.org/wiki/Checking_whether_a_coin_is_fair The problem is as follows. I give you a sequence of coin flips, for example TTTHHTTT, and your task is to determine whether this coin is fair. Obviously you can't answer this question with certainty. The frequentist approach seems rather ridiculous to me. They pretend to be able to answer this question without knowing anything about coins. Instead of making the assumptions up front, they hide the assumptions in the method of determining whether the coin is fair. For example one would think that to decide whether a coin I give you is fair, you'd need some idea about what kind of coins I will give you. If we were talking about reality, you would be able to say "no it's not fair" without even looking at the data sequence, because in reality there are no fair coins. If we are in a different universe where fair coins do exist, you'd need some idea how many are fair. So obviously whether a coin is fair depends on which universe you live in, but the frequentist method is not parameterized by the universe. The assumptions are implicitly made inside the method. The bayesian approach asks you first to state your prior ideas about coins. Then you give the bayesian your data, and he will compute for you the probability that this coin is fair. This cleanly separates the correct mathematical derivation from the subjective assumptions, instead of hiding these assumptions inside the method. Moreover, the bayesian approach is automatic, in principle. Once you make your assumptions clear, the rest is just mechanical derivation. The frequentist approach requires divine inspiration, and is very easy to get wrong. For example you're not allowed to look at the data before formulating your hypothesis. Of course nobody does this right in practice. I've often seen people cherry pick a frequentist statistical test that proves their hypothesis. Given any data and any hypothesis, I can devise a correct frequentist statistical test that proves the hypothesis with arbitrarily large confidence.
- hessenwolf 16y ago
- araneae 16y agoWhile I like Bayesian statistics it is NOT a substitute for maximum likelihood in many situations. Additionally, some of his criticism of NHST is unfair because he criticizes weaknesses Bayesian also has. In the part where he gives the example of the pollster and how uncertain the p value would be because you have to incorporate sample design- well, you have to do that in Bayesian stats too. Statistics is a big field. Obviously people should expand their tool set, and maybe Bayesian is underused, but that doesn't mean Bayesian stats are right for every experiment and experimental design.
- jules 16y agoMaximum likelihood estimation falls out of Bayesian statistics with the right utility function. In practice though, that's often not the utility function you want, hence general Bayesian statistics.
- stochastician 16y agoCan you give your favorite example where ML methods lack a (useful) corresponding Bayesian substitute/replacement?
- nycticorax 16y agoI more-or-less agree with the OP: I'm a neuroscientist, and it would be nice if I could use Bayesian analyses in my papers without it being a point of contention with reviewers. But it seems like one of the big advances in frequentist statistics in the last fifty years is the introduction of nonparametric methods, which don't require you to make strong assumptions about the distribution of your data. My understanding is that the field of Bayesian nonparametric inference is still in its infancy. Also, this paper seems relevant: http://stat.stanford.edu/~ckirby/brad/papers/2005NEWModernScience.pdf http://stat.stanford.edu/~ckirby/brad/papers/2005NEWModernSc... (Bradley Efron is in the running for Greatest Living Statistician.)