10 ms·
The hardest part of bayes is 1) understanding what 'likelihood' actually is or represents 2) understanding what a 'partition function' actually is or represen
by platz 4y ago
The hardest part of bayes is
1) understanding what 'likelihood' actually is or represents
2) understanding what a 'partition function' actually is or represents
I always forget what these mean since I don't use bayes in anger
Also the above reveal there is more elaborate structure in bayes than just the idea of 'updating your prior with new information', which is simply a platitude among STEM folk
- DiogenesKynikos 4y ago> understanding what 'likelihood' actually is or represents The probability of observing the data you observed, assuming that the model parameters take on a certain value. > understanding what a 'partition function' actually is or represents The probability of observing the data you observed, this time averaging over all possible parameter values (weighted by the priors).
- psi75 4y ago> The probability of observing the data you observed Yes, in discrete cases. In continuous cases, you have to work with a probability density. I think this is one of the hurdles people encounter when they're first exposed to Bayesian stats. The probabilities, technically speaking, are zero. The important insight in Bayesian work is that it's often not the probabilities themselves that matter but the ratios thereof, since from those alone you can compute posteriors.
- DiogenesKynikos 4y agoIndeed, but it's a bit tedious to always say, "probabilities (in the discrete case) or probability densities (in the continuous case)." In general, you have sums in the discrete case and integrals in the continuous case, but most formulas are otherwise the same.
- psi75 4y agoThat's quite true. Also, one could argue that continuous probabilities in practice are discrete probabilities due to finite resolution--we just don't care to specify what the resolution is.
- analog31 4y agoI once had a TA job for an undergrad stats course. This was the "non calculus" course for the psych majors. I had also taken the "math" version of the same course, where we spent two semesters and proved everything. I honestly never came up with a satisfactory layman's explanation why continuous distributions are necessary, or what "continuous" is. I knew that we used calculus to derive the formulas that they were faced with memorizing, but that would have been irrelevant to them. The best explanation I can think of today is: Use the one that makes the math easier or more readable.
- DiogenesKynikos 4y agoMany measurements are continuous. What is the age of a rock? That's not a discrete quantity: it could be anything.
- analog31 4y agoPlus or minus what? I know the importance of continuous sets in the study of statistics as a branch of math. But I don't know of any measurements that can't be represented by integer multiples of a unit for all practical purposes. And the students in the non-calculus stats course can't grasp what continuity is anyway.
- DiogenesKynikos 4y ago> Plus or minus what? That's what the probability density specifies. > But I don't know of any measurements that can't be represented by integer multiples of a unit for all practical purposes. You can always discretize any real number, but why would you? I don't see how integers are easier to deal with than real numbers. Calculus can be viewed as the limit in which you discretize numbers infinitely finely. Once you know how things work in that limit, it's generally easier to use calculus than to work with discretized quantities. One example: summations are often more difficult than integrals, and one way of approximating sums is to turn them into integrals. From my perspective, calculus is a basic part of mathematics that everyone should be expected to learn in school. In the US, calculus is often viewed as some sort of intimidating subject that only extremely clever people can grasp, and then only in late high-school or in university, but in East Asia, it's taught to children as a matter of course.
- melling 4y agoI thought likelihood doesn’t necessarily sum to one so it’s not a probability.
- MontyCarloHall 4y agoIt will always sum (or integrate) to one with respect to the data. For example, given likelihood p(x1, x2, ..., x_N|params), summing (or integrating) over all possible values of x_1 ... x_N will indeed yield 1.
- MontyCarloHall 4y agoLikelihood: given a probabilistic model and its parameters, what is the probability of observing some data under that model? For example, given a coin with some probability f of getting heads and thus 1-f of getting tails (model parameter), a set of coin flips (data), and the assumption that each coin flip is independent of all others (model), the probability of seeing an arbitrary sequence with H total heads and T total tails after H+T flips is p(H,T|f) = f^H(1-f)^T [0] Note that this likelihood is normalized (i.e. sums up to 1) with respect to all possible sequences of flips of length H+T, e.g. for length 2, there are 4 possible sequences of coin flips: p(hh|f) + p(ht|f) + p(th|f) + p(tt|f) = f^2 + 2f(1-f) + (1-f)^2 = 1 (Also note that the likelihood is only a probability for discrete data; it’s a probability density for continuous data, since the underlying terms would no longer be probabilities but rather densities. In that case, replace the previous sum with an integral ranging over the entire domain of your continuous data.) Partition function (usually called a "marginal likelihood"): what if we instead normalize the likelihood with respect to the model parameters? Then it would no longer be a probability distribution with respect to the data, but rather a probability distribution with respect to the parameters. The marginal likelihood is just this normalizing constant. In the previous example, p(f|H,T) = p(H,T|f)/<constant that would normalize p(H,T|f) to integrate to 1 over all possible values of f> You can optionally weight this likelihood by a prior p(f), in which case the numerator would be p(H,T|f)p(f), with the partition function updated accordingly. >Also the above reveal there is more elaborate structure in bayes than just the idea of 'updating your prior with new information', which is simply a platitude among STEM folk I could not agree more. The core philosophical tenet of Bayesian inference is that this re-normalization of the likelihood returns something probabilistically meaningful. The debate over the validity of priors pales in comparison to the debate over whether the likelihood ought to be treated as proportional to a probability distribution over the model parameters. [0] Note that this is distinct from the probability of seeing any sequence with H heads and T tails. That would require normalizing with respect to the number of total sequences with H heads and T tails (H+T choose H), yielding the binomial distribution.
- oxff 4y agoThe 'hardest' part is realizing that basically everyone reasons this way unless taught otherwise. The books do a lot of hard work to cover this up.
- jltsiren 4y agoThe hard part is that the intuitive understanding of probability you learned as a kid is almost certainly wrong. You need to unlearn it and replace it with proper axiomatic treatment of probability based on measure theory. Once you have done that and built a new intuition, Bayesian thinking should feel pretty natural. Once you accept that probability is just the "size" of a set relative to the size of a superset and it's up to you to attach meaning to the sets, much of the confusion goes away.
- nerdponx 4y agoI think you can get most of the way there with straightforward geometric intuition, treating the sample space as a rectangle of area 1 and proceeding from there into joint and conditional probability.
- lisper 4y agoNo, the hardest part is realizing that there are all kinds of tacit assumptions that we bring to bear merely by formulating a problem for Bayesian analysis. Take the example used in chapter 2 of the book. It assumes that news articles can be classified as "real" or "fake", and there is no middle ground. It assumes that the initial prior produced by experts is reliable. And most of all it assumes that certain features, like exclamation points in the title, are causally related to realness or fakeness. To illustrate this last point, consider analyzing the titles for a different features, the presence of the letter "z" rather than the presence of exclamation points. If it turned out that fake news in the corpus used to generate the priors just turned out to have more zees in their titles, would you then be justified in concluding that a news article about Zanzibar was more likely to be fake because it contained two zees? That example might seem contrived, but if you use a Bayesian spam filter it is actually plausible that the presence of the letter "v" is dignostic of spam because of the prevalence of spam concerning viagra. But again, there is a causal model of this: viagra is a product that is often the subject of spam marketing, and the word "viagra" happens to have a v in it, which is a priori an uncommon letter in English. But all this can fall apart depending on the circumstances. If you one day joined an email discussion of Stradavarius violins, the v signal could suddenly fail. An even more dramatic example: suppose you are an academic who starts to do research on spam filters and you have a collaborator who starts to send you examples of hard-to-filter spam. Now you have a very strong signal, but no straightforward textual analysis will allow you to extract it. The hard part of Bayesian analysis is deciding what features to even look at. All the rest is borderline trivial by comparison.
- bumby 4y ago>the hardest part is realizing that there are all kinds of tacit assumptions that we bring to bear merely by formulating a problem for Bayesian analysis. I think the common argument is this is a strength of Bayesian analysis. Namely, that your priors make you explicitly state your assumptions. All models integrate assumptions, but not all of them make you explicitly quantify them like Bayesian analysis does.
- nerdponx 4y ago
- deleted 4y ago[deleted]
- analog31 4y ago>>> ... the idea of 'updating your prior with new information' ... ... dates back to antiquity. I'm not a statistician, but my impression is that Bayesian methods are a way of formalizing that approach. On the other hand, use of the term "prior" is somewhat misleading because the formulas do not specify a time sequence for acquiring information. I tend to have a lesser view of sprinkling Bayesian terminology into blogs about social issues. That's what I call Bayes Theater.