3 ms·
Generally, when we construct models we do so by defining what probability they give to the data. That's a function that takes in your data set and returns some
by tel 2y ago
Generally, when we construct models we do so by defining what probability they give to the data. That's a function that takes in your data set and returns some number, the higher the better.
Technically, these functions need to satisfy a bunch of properties, but those properties matter mostly for people doing the business of building and comparing models. If you just have a model someone already made for you, then "the higher the better" is good enough.
It's also the case that these models have "parameters". As a simple example, the model of a coin flip takes in "heads" or "tails" and returns a number. The higher that number, the more probable it claims that outcome to be. When we construct that model, we also choose the "fairness" parameter, usually setting it so that both heads and tails are equally likely.
So really, it's a function both of the data and of its parameters.
Now, "maximum likelihood estimation" (MLE) is just the method where you fix the data inputs to the model to whatever your training data is and then find the parameter inputs that maximize its output. This kind of inverts the normal mechanism where you pick the parameters and then see how probable the data was.
Presumptively, whatever parameterization of your model makes the data the most likely is the parameterization that best represents your data. That doesn't have to be true, and often is only approximately true, but that presumption is exactly what makes MLE popular.
Finally, it's worth describing the origin of the name. When we look at our model after fixing the data inputs and consider it a function of its parameters instead we call that function a "likelihood". This is just another name for "probability" except it's used to emphasize that likelihoods don't meet all the technical properties I skipped up above. So "maximum likelihood estimation" is just the process of estimating the parameters of your model by maximizing the likelihood.