9 ms·
Does someone have an ELI5 for frequentist vs. bayesian?
by zump 10y ago
Does someone have an ELI5 for frequentist vs. bayesian?
- raverbashing 10y agoI'll try They both yield the same results when your knowledge of the given probabilities is exact But Frequentists will look at a "top down" view, Bayesians will look "bottom-up", more importantly, as based on Bayes's theorem, they will look at "if a then b" kind of probabilities. This sounds like a good explanation (the 1st answer) https://stats.stackexchange.com/questions/22/bayesian-and-frequentist-reasoning-in-plain-english https://stats.stackexchange.com/questions/22/bayesian-and-fr... also obligatory XKCD: https://xkcd.com/1132/ https://xkcd.com/1132/ (Though frequentists are not so naive usually)
- Retric 10y agoThat XKCD is a joke and not how frequentists view statistics. The confusion comes from the way you read papers. A frequentist looks at a paper as evidence but not truth. Bayesian on the other hand gets to the same place with slightly different math.
- marcosdumay 10y agoThe joke is about priors selection. I'd say it's spot on, since, as you say, the only reason people are not fooled like that is because they evaluate the frequentist results as evidence in an ad-hock Bayesian model. The truth is that nobody thinks exclusively on frequentist or Bayesian terms. But that's not comics-grade material, and mixing them would hide instead of surface their differences.
- Retric 10y agoThe problem is it's a single sample. If the output was Yes, Yes, Yes, Yes, Yes, No, Yes, then you can do frequentist statistics, but when the output is just 'Yes' then the sample size is one.
- raverbashing 10y agoHaving only one sample is not a problem if you're Bayesian.
- Retric 10y agoWhich is it's own problem. A Bayesian is often happy to look at any data that agrees with their own interpretation which is why it's not useful for papers. The idea that A cause cancer is ridiculous. Collect data, well the A group has 10x as much cancer, but that's ridiculous so I conclude there is no relationship between A and cancer.
- red75prime 10y agoCollect data, well the A group has 10x as much cancer, now it is about 10 times less ridiculous. This should be about right.
- yummyfajitas 10y agoThere is no principled reason you can't compute a p-value from a sample size of 1.
- deleted 10y ago[deleted]
- Retric 10y agoWhat's the standard deviation of a sample size of one? Now, the standard counter example is a composite statistic. Like roll 100 6 sided dice get 600 and assume they are not fair. But, importantly there is a standard deviation assumed in the experiment and there was more than one dice roll. However, if you combine that with something else then your sample size drops back to one.
- yummyfajitas 10y agoThe p-value is defined as p=P(evidence seen | null hypothesis). The standard deviation is only relevant if it is required to compute that number. You can run NHST on distributions without a p-value, e.g. a Cauchy distribution. You might need a standard deviation if you want to do some naive Z-test based on the CLT approximation (since the normal distribution requires a standard deviation), but that's not what XKCD was describing. XKCD was describing an exact test using the true distribution.
- MarkMc 10y agoThis is a very interesting comment for someone like me who has little knowledge of statistics. The stack exchange answer shows how stupid Bayesians can be, and the XKCD shows how stupid Frequentists can be. Yet I find this particular criticism of Bayesians not fully convincing. The Bayesian approach is to take the existing knowledge (where the phone was often list in the past) with new knowledge (where the beeping is coming from) to come up with probabilities (where to look for the new phone). This seems to me to be the correct approach in general, it's just that in the case of the phone the new knowledge almost entirely outweighs the existing knowledge.
- eximius 10y agoMy favorite explanation is that frequentist methods answer the question "If I assume a model, A, what is the likelihood this data came from it?" while bayesian methods answer "Given this data, what is the model?" Frequentist methods rarely directly answer the question we actually have. But they're generally far easier to compute. Bayesian methods are often much more intuitive but are far more complicated and less performant.
- gbrown 10y agoMeh, most Bayesian techniques still assume a model. It's more like: Assuming a model M characterized by parameters T and giving rise to data Y, what is P(T|Y,M) To be sure, you can compare the probability of models as well, and there are Bayesian semiparametric techniques, but models are still really important.
- StClaire 10y agoThey start to diverge on the issue of what a probability actually is: frequentists see them as long run averages, and Bayesians see them as degrees of beliefs. If you have a coin that comes up heads 60% of the time, a frequentist looks at that as, "as the number of times I flip my coin goes towards infinity the proportion of heads I get goes to 60%." A Bayesian thinks, "absent other evidence on the how the coin gets flipped, I'm about 60% sure the outcome of the flip will be heads." This lets Bayesians talk about the probability of single events, like basketball games, where frequentists can't. Bayesians also see conditional probabilities everywhere. A conditional probability says, "well, if I know something about the situation, I should include it in my beliefs." Circumstances matter. It doesn't make a ton of sense to talk about the chances I get hit by a car. They very wildly depending on whether I'm standing on the highway or eating in my kitchen. Another basketball example. The chances that Spurs win changes dramatically if they play the Warriors or if they play the Kings. I might say "they have a 90% chance of winning given they play the Kings, but a 40% chance of winning given they play the Warriors. I also need a likelihood function. What are the odds I saw my data given my hypothesis is true. If I got hit by a car, what are the chances I was standing in the street? Given the Spurs won, what are the chances they played the Warriors? We use something called Bayes Rule which allows us to pile on more and more information on something we call a "prior belief," what we thought about our hypothesis before we saw our data. As we pile on data, we expect to change our beliefs. We become more sure of what we thought, maybe we become less sure, maybe we can totally change our minds. I want to use Bayes' example since I think it's so good. Imagine you came out of Platos cave and saw the sun rise. You'd think, that's weird, I bet that doesn't happen again. The sun goes down, and you spend some time in the dark. The next morning the sun comes up again. Now you're less sure that sun rises are fluke events. Plus you found some people who aren't freaking out about the whole "big ball of fire in the sky" thing. Maybe now you don't expect the sun not to rise tomorrow. Maybe it will, maybe it won't. As you see more and more sun rises, eventually you get to the point where you are extremely confident that the sun rises every morning. You saw more sunrises and updated your beliefs. We need one last piece of information: a prior. That's that initial belief you're cave-escaping-self had that sun rises are weird and you probably won't see another one. We can estimate them through population data--percentage of games the Spurs won against the Warriors--or we could just make them up. This is just our belief about the truth of our hypothesis; the chance I get hit by a car regardless of where I am. We take all this put it into Bayes rule, a blender that gives us the probability that our hypothesis is true given we saw our data. We can use this as a new prior too. One last example. I'm 70% sure the earth is round. I see a picture of the horizon taken from a hot air balloon and I think there's a 90% chance that I would see that if the earth were truly round. Without going into the calculation, I'm now give or take 80% sure the earth is round. I saw data to support my hypothesis and my belief got stronger. Why do Bayesian analysis? Because someone once published a study that says that frogs can sense earthquakes some time before they occur. That may be true, but I'm skeptical. My skeptical prior would only get moved slightly to become less skeptical, but it would still need more information, a replication of the study by other people, to actually convince me.
- tnone 10y agoNobody seems to be capable of explaining this properly. It's like monad tutorials, they explain what happens while mistakenly thinking they are telling you why it happens. I keep trying to fit this idea into my head and I can't because the information is not given. - Where did this difference come from? When did it develop? - What are the basic premises that a Bayesian believes that a frequentist doesn't, and vice versa? Reason it all the way through front and back. - What does the B/F's model look like? What are the pieces they use, how are they arranged, what are the dependencies, how does causality flow? - Why are the choices made by one invalid for the other's model? Where do they agree deliberately despite this? - What are the consequences in the real world? Give me a real example on why this difference matters? "Real" meaning I don't care about dice, I care about engineering and science. Instead you get some bullshit about fitting a distribution you don't understand to a model you can't see, while relying on understanding the nuances between words like probability and likelihood which is what you are trying to learn in the first place. Plus I swear the numbers agree in 99% of the "examples" given, with some handwaving "but it's different" to excuse it. Fucking explanations, how do they work? Not in academia.
- willis77 10y agoI agree, and suspect many people (outside of the statisticians who have had the time and space to digest the philosophical underpinnings) who claim to love Bayesian methods do so because they have been told it's the right thing to love. There are a lot of hand-wavy explanations out there that tell you what each side believes, but I have yet to see something that truly ELI5s it.
- gbrown 10y agoWell, maybe I'm an exception since I am a statistician, but Bayesian techniques allow me to do things which are simply impossible with frequentist tools. To be fair, I use both approaches just about every day.
- mattkrause 10y ago"I use both approaches just about every day" is the only sane answer here. Different tools for different jobs. What would think if you met carpenters who described themselves as "Hammerists" or "Sawsallitarians"?
- stdbrouw 10y agoAn example from Leonard Mlodinow: imagine a friend doesn't pick up the phone when you call. Now, you might think that maybe they're upset at you, because if they are indeed upset, the the probability that they wouldn't pick up the phone is very high. On the other hand, there might be many other reasons your friend doesn't pick up the phone – battery's dead, they went on an impromptu holiday, they're having a bad day, they didn't hear it ring. Frequentist statistics deals in the first kind of probabilities (the probability of seeing what you saw given a particular hypothesis) whereas Bayesian statistics is the other way around (the probability of a particular hypothesis given what you saw).
- xapata 10y agoExcept the Frequentist would test the null hypothesis: what's the probability the friend doesn't answer if the friend is not upset?
- stdbrouw 10y agoGiven that the alternative hypothesis is usually defined as the complement of the null hypothesis (HA = 1 - H0) this doesn't make much of a difference, though.
- xapata 10y agoFrequentist: H1: answer ~ upset + error H0: answer ~ error I find many people have trouble expressing a reasonable hypothesis / null-hypothesis pair. In fact, I'd bet that a good chunk of folks would try to make "upset" be the dependent variable in the phone call scenario.
- tel 10y agoStatistics is about finding a model of the world that we can trust. A model in this circumstance must be one that makes predictions about the world and therefore we trust it when we expect that its predictions will be largely correct—or at least more correct than any other model we have. Frequentists and Bayesians disagree on the processes for building and evaluating models. Their techniques are often complementary and are, in current and historical practice, almost always used together by professionals. I would call them two sides of the same coin, although some take philosophical perspectives which are more dogmatic. The theoretical foundation which divides them (confusingly called Bayes' Law) states that there is a relationship between "the probability of seeing some event when a model is true" and "the probability of a model being true given some event happened". In short, Frequentists tend to build their processes off of the first notion and Bayesians off the second. In practice, Frequentist methods build their models via whatever tools they like. These often include basic optimization tools for picking the best set of "parameters" of a model. They then evaluate the performance of these models by asking "how unlikely was reality given this model was true?" and rejecting models which fail to predict what happened. Bayesian methods are more fixed but also more dramatic in their ways of constructing models. They tend to create vast models with many moving parts using what's known as a "generative story". This is acceptable since they use the data they observe to compute "probability of truth" for all possible permutations of their model. This is considered a final result since someone might want to ask "how much more likely is model A to be true than model B?" but Bayesians will also at this point use optimization techniques to find "the most probable model". In many cases these two approaches arrive at the same places. In times that they do not they provide interesting questions about what we really mean when we say that we "trust a model" and this leads to endless discussion. It's also often the case that "avowed Frequentists" have historically used Bayesian methods to discover their basic models and then evaluated those models in a Frequentist fashion for publication (Fisher was known to do this). This arose because at a certain point in statistical history Bayesian methods were not socially acceptable. Finally, it's a pretty good idea for Bayesians to evaluate their models in Frequentist forms in order to have more ways to discuss how their models perform. Probably the last and most practical difference between the two is that Frequentists methods are often built taking into account their time and space complexity. Frequentists are more likely to evaluate the performance of various extremely simple estimation techniques and pick the best. The Bayesian process nearly always results in an extraordinarily difficult to evaluate integration problem that requires modern computers to get results out of. That said, Frequentists often arrive at their best results via "strokes of genius" while Bayesians can usually chug through any modeling problem and arrive with a decent (again, computational only) model.
- prandalr 10y agoFrequentists are Bayesians who use a uniform prior.
- apathy 10y agoAnd refuse to serially update it ;-) Which is why no working statistician is really 100% Bayesian (intractable in many cases) or 100% frequentist (obviously wrong in some cases). We all use Bayes' Rule (don't need to be "a Bayesian" to do that) and we all are forced to do Newman-Pearson-style power calculations now and then (holds nose). Even the latter have their uses, in preventing the worst of the worst abuses of frequentist techniques (it's not frequentism that is inherently bad per se, it's the profound abuses that turn it into a magnet for bad science; sometimes a Bayesian formulation of a problem is simply intractable).
- nazka 10y agoFrequentists like to give you a perfect answer the first time by computing everything at once. Seeing the world like pure math. For example giving an unbiased dice of 6 faces, the likelihood to have a 6 the first time is = 1/6. Bayesians like to walk in the park, and see step by step how things go. The more they walk, the more accurate the result will be. For example at step 5 with 1st roll: 1, 2sd: 6, 3rd: 3, 4th: 1, 5th: 2, you will have (1 + 0 + 0 + 1 + 0) / 5 = 0.4. At step 6 with a roll 5 you will have (1 + 0 + 0 + 1 + 0 + 0) / 6 = 0.333. The answer being closer and closer to the true answer after each roll, each step. Ultimately with enough rolls, bayesians will start to give you an answer close to the frequentists' one.
- kruschke 10y agoYes! Here: http://doingbayesiandataanalysis.blogspot.com/2017/02/the-bayesian-new-statistics-finally.html http://doingbayesiandataanalysis.blogspot.com/2017/02/the-ba...