4 ms·
This is a terrific explanation for the uninitiated - covers pretty much all the reasons I would think of. The author mostly describes nice mathematical propert
by TTPrograms 11y ago
This is a terrific explanation for the uninitiated - covers pretty much all the reasons I would think of.
The author mostly describes nice mathematical properties, so I'd recommend considering absolute deviation in practical analysis in many cases as well. Then your error is dominated less by few outliers in your model. Specifically, least-squares is best when you expect your error to be gaussian distributed, which isn't always the case.
See for example:
http://matlabdatamining.blogspot.com/2007/10/l-1-linear-regression.html http://matlabdatamining.blogspot.com/2007/10/l-1-linear-regr...
You might say the example is contrived, but if the outliers were gone both estimates would give nearly the same result. The main time to prefer L_2 would be when your error is gaussian and your variance is huge, but that's going to be an uphill battle regardless.
- benkuhn 11y agoThanks for the kind words! I actually agree that L1 error is often more practical, and didn't mean the post to come off as prescriptive. (In fact, part of my motivation for writing it was that L1 error seemed to have much better practical properties, and I was curious why people used L2 despite this!)
- repsilat 11y agoIt's a good post. If you're interested, Nassim Taleb (shudders) wrote a counterpoint for Edge.org that seems to have disappeared. It exists on the Internet Archive, though: https://web.archive.org/web/20140606195420/http://edge.org/response-detail/25401 https://web.archive.org/web/20140606195420/http://edge.org/r... Well, that's not quite accurate -- Taleb doesn't argue that we should do away with squared error, he says we should do away with the standard deviation in favour of expected deviation from the mean. (Hrm, I'm going back and forth on whether there's a substantive difference between the concepts of variance/standard deviation as measures of a probability distribution and squared/RMS error measures... Obviously they coincide when the "prediction" is the mean of the data, but I don't think that's terribly convincing. I think it'd be perfectly reasonable to use squared error in your model fit and parametrise your Gaussian distributions by their average absolute deviation from the mean.) Discussed on HN https://news.ycombinator.com/item?id=7064435 https://news.ycombinator.com/item?id=7064435
- stdbrouw 11y agoWhen talking about deviation, you're describing the spread of a distribution. It's perfectly fine to describe your data using mean (or median) average deviation, but then fit a model on top of that minimizes the sum of squared errors. Taleb's point is that squared deviations make for a lousy descriptive statistic as it leads you to underestimate the average distance of observations from the mean.