3 ms·
Nice write-up! Minor nitpick: ML/MAP estimators don't _require_ observations to be independent. At least, in my field we're looking at a single observation of a
by stilley2 8y ago
Nice write-up! Minor nitpick: ML/MAP estimators don't _require_ observations to be independent. At least, in my field we're looking at a single observation of a multivariate distribution, and we don't need to assume the elements are independent (ie, we permit a non-diagonal covariance matrix). My intuition says this is equivalent to assuming multiple correlated scaler observations, but I'd have to sit down with some paper. Also, you use "trough" where I think you mean "through."
- nerdponx 8y agoDependence between elements of the same observation is irrelevant. The point is that different observations must be independent and identically distributed for the standard formulation of the likelihood to be valid. Typically we write the likelihood function as L = Π P(y | θ) If you didn't have identically-distributed observations, the functional form of P would be different for each observation. And if you didn't have independent observations, then you're basically screwed in the general case. That expression for L is basically the definition of probabilistic independence: a finite set of random variables is mutually independent if and only if their joint probability function is equal to the product of the individual variables' probability functions. If you have dependence between observations, you lose the ability to write L in that nice form. This is a non-negotiable consequence of basic probability theory. The only way to do MAP estimation without iid observations is to know the joint distribution of your entire dataset, and be able to maximize that distribution with respect to θ given an arbitrary data set. This is possible but it's not quite the same thing as dumping your data into a GLM.
- conjectures 8y agoThe post this is a reply to was correct, and this is not. E.g. a simple counter example is finding the autocorrelation parameter in an AR(1) model for an economic time series. Under your suggested definition of MLE this can't be done, which is simply not the case. In fact, not approaching the more general case is liable to confuse learners as they may think that independence assumption is somehow baked into MAP/MLE, which it is not.
- nerdponx 8y agoI never suggested a definition of MLE. You need independence to use the "L = Π P(y | θ)" formulation, full stop.
- stilley2 8y agoTrue. But that form is a convenience, not a requirement.
- conjectures 8y agoYes, you do need independence to assume the likelihood factorises. You do not need independence to find a MLE/MAP.
- stilley2 8y agoBut that form is not required. A quick counter example. I'm trying to estimate a value from N measurements. The measurements experience Gaussian noise with some general covariance matrix K (i.e., they are not independent) Therefore, y is a sample from N([1, 1, ..., 1]^T u, K). The MLE is then ([1, 1, ..., 1]K^{-1}[1, 1, ..., 1]^T)^{-1} [1, 1, ..., 1] K^{-1} y. Or in words, multiply y by the inverse covariance matrix, sum the result, and divide by the sum of all the elements in the inverse covariance matrix. As a sanity check, when the measurements _are independent, this reduces to a weighted average, where the observations are weighted by their inverse variances.
- stilley2 8y agoI suppose my example meets the case of knowing the joint distribution.
- deleted 8y ago[deleted]