4 ms·
Another thing to note is that you're multiplying probabilities together. Since each probability is between 0 and 1, youre always shrinking the likelihood with e
by ronald_raygun 3y ago
Another thing to note is that you're multiplying probabilities together. Since each probability is between 0 and 1, youre always shrinking the likelihood with each new data point. When you're doing this kind of analysis, the question you're asking is "given a model with these parameters, what's the probability I get exactly this sample?" Which, when you phrase it that way, it becomes more apparent why the likelihood is so small.
- cobbal 3y agoMultiplying them together certainly magnifies the effect, but it would magnify it the other way if the likelihoods were larger than one. (Easy to get, just tweak the variance of the normal distributions to be smaller). Likelihoods are more like infinitesimal fractions of a probability, that need to be integrated over some set of events to get back a probability. In the case of the joint distribution of 50 Gaussian, you can think of the likelihood having "units" of epsilon^50.
- j7ake 3y agoWait how do you get likelihoods greater than one? Definitely won’t work for likelihoods like Poisson or other count based models.
- blackbear_ 3y agoFor discrete distributions you indeed cannot, but for continuous distributions all you need is sufficiently small variance. Try for example a Gaussian with variance 1e-12
- KeplerBoy 3y agoThe value of a continuous probability density distribution at a specific point is pretty meaningless though; You have to talk about the integral between two values and that won't go above one.
- blackbear_ 3y agoIn context of maximum likelihood the value of the density at the maximizer is actually quite useful, for example for model comparison.