4 ms·
Statistician here. There's a deep idea called the likelihood principle https://en.wikipedia.org/wiki/Likelihood_principle https://en.wikipedia.org/wiki/Likeli
by waldrews 3y ago
Statistician here. There's a deep idea called the likelihood principle https://en.wikipedia.org/wiki/Likelihood_principle https://en.wikipedia.org/wiki/Likelihood_principle that says all the information we can get from the data about model parameters is contained in the likelihood function.
We're talking about the whole likelihood surface here, not just the single point that's the maximum likelihood estimator. The MLE is a method for choosing a valid point estimator from the likelihood function; it has some good properties, like being consistent (if you have enough data it converges to the truth) and asymptotically efficient (converges smallest possible variance) so long as some criteria are met.
But the MLE is not the only choice; for any given model, other procedures can be admissible estimators https://en.wikipedia.org/wiki/Admissible_decision_rule https://en.wikipedia.org/wiki/Admissible_decision_rule - it's just they also have to be procedures based on the likelihood function. In other words, your procedure doesn't have to be "take the likelihood function and find its maximum" but it has to be "take the likelihood function and... do something sensible with it."
So the MLE is popular in the frequentist world where you have to make the decision rules using the likelihood directly; in the Bayesian world, you take the likelihood and combine it with a prior, to make an actual probability distribution. Then you get things like like MAP (mode of the posterior) or the Bayes estimate (expectation of the posterior) - alternatives to MLE that still use the likelihood surface.
Of course this all works only if the underlying probabilistic model is literally true. So, in the machine learning world where the models are judged on being useful on usefulness and not expected to reflect mathematical reality, you're allowed to do things inconsistent with likelihood principle, like regularization tricks. In some physics situations (astronomical imaging comes to mind) where the probability model really is governed by the rules of nature, sticking to likelihood principle actually matters.
As to the question of being small, well, the likelihood is the probability (density) of the exact data you observe given parameters. Let's say you know the true parameter (the mean and standard deviation) and you observe a thousand draws from a normal distribution. Of course the probability of observing the very same pattern of a thousand values again is overwhelmingly unlikely. But if the mean was way different, that pattern would be proportionally even more unlikely.
We should only care about relative probabilities. What's the probability that the universe evolved in exactly such a way that your cat will have exactly this fur pattern? Astronomically small. What's the probability that the universe evolved in such a way, and some of that fur ends up on your furniture? Another unimaginably small number. But what's the probability that, in a universe where you and your cat exist as you are, his fur will get everywhere? That's pretty much a certainty.