4 ms·
I would care to interject. First of all, you are right on several points. * Most of mathematics is the same in both schools of though, and the interpretations
by ivan_k 11y ago
I would care to interject.
First of all, you are right on several points.
* Most of mathematics is the same in both schools of though, and the interpretations is not different.
* Some basic ideas (i.e. the nature of probability) are quite different, and this is where most of the argument (Frequentist vs. Bayesian) comes from.
However, this second point has a major impact on calculations. So I disagree here:
* The notion of a _prior_ (probability of the hypothesis, P(H)) is essentially nonsensical in the Frequentist view. Any frequentist would just call it 'bias'. However, for a Bayesian, this is the degree of belief that you put in you system before you do any measurements. Practically, it is either non-informative (you don't give more belief to any hypothesis a-priori), or comes from earlier data. The prior gives you a natural way to incorporate multiple experiments. I think that is a large difference in calculation (or at least the structure of calculation).
More importantly, Bayesian statistics inspires (and is enabled by) Markov Chain Monte Carlo inference. It is the main mathematical machinery used for today's Bayesian data analysis, and is impossible in the frequentist framework. This approach allows you to scale to very complicated (i.e. feature-rich, multiparameter) data, explore very complex (i.e. non-convex, hard to optimize) probability surfaces. All of this stuff is very hard (if at all possible) in the frequentist framework.
So there are difference. But people don't really argue about them. There is no grand flame war, or anything of that sort.
Oh, and contrary to what the article suggests, no statistician likes p-values.
- TuringTest 11y ago> The notion of a _prior_ (probability of the hypothesis, P(H)) is essentially nonsensical in the Frequentist view. Any frequentist would just call it 'bias'. However, for a Bayesian, this is the degree of belief that you put in you system before you do any measurements. Thanks, I was having problems with that point. Stating that "a priori, all hypothesis are equally likely" looks like a too strong assumption to make from complete lack of information. If you interpret it instead as "lacking information, I don't have a reason to prefer any hypothesis over the others" it seems more reasonable. However, that doesn't solve my qualms with the Bayesian approach as explained in this article. I understand the justification of Bayes Theorem from a frequentist approach, as starting with all the possible outcomes, and filtering that initial probability through the lens of available information; i.e. removing facts that we know can no longer be true, and counting those who can. In such context, the theorem seems intuitively true. However, if the a priori probability is interpreted as a lack of knowledge, the form of the theorem looks much more arbitrary. Why would that particular computation be the best way to increase our confidence, if the starting point is arbitrary and the shape of the formula is not related to the number of cases that can be true or false in the current state of the world? I understand that Bayesian analysis counts with well-developed and practical tools. But what I get from this article is that their particular form seems to come from tradition rather than any intrinsic property of that model - if you reject frequentism, any counting model might a priori work as well as the Bayesian one. Edit: Apparently Wikipedia agrees with me in this point.[1] There are other rational models for updating your probabilistic belief, and Bayesian is used primarily for being computationally convenient, rather than theoretically incontestable. Or am I reading too much into it? I'm certainly not expert in probability. [1] https://en.wikipedia.org/wiki/Bayesian_inference#Alternatives_to_Bayesian_updating https://en.wikipedia.org/wiki/Bayesian_inference#Alternative...
- ivan_k 11y agoI am not sure what you mean by "the form of the theorem looks much more arbitrary". The derivation of Bayes law comes from the axioms of conditional probability. Given two events A, B; we have: P(A^B) = P(A|B) * P(B) Probability of A and B = Probability of A given B happened times probability of B Symmetrically, we can say: P(A^B) = P(B|A) * P(A) Now we have: P(A|B) * P(B) = P(B|A) * P(A) Rearranging, we get: P(A|B) = P(B|A) * P(A) / P(B) So I do not see this as being particularly arbitrary. While other rational models are possible, I find this one rather practical and satisfying. [Edit]: markup
- TuringTest 11y agoWhat I mean is that those axioms of conditional probability seem intuitively true because of their frequentist interpretation, i.e. counting the possible cases that satisfy each probability. If you devoid them from the combinatorics that justify their meaning, there's no special reason to accept these particular axioms nor the law derived from them.