3 ms·
One interesting property of least squares regression is that the predictions are the conditional expectation (mean) of the target variable given the right-hand-
by c7b 1y ago
One interesting property of least squares regression is that the predictions are the conditional expectation (mean) of the target variable given the right-hand-side variables. So in the OP example, we're predicting the average price of houses of a given size.
The notion of predicting the mean can be extended to other properties of the conditional distribution of the target variable, such as the median or other quantiles [0]. This comes with interesting implications, such as the well-known properties of the median being more robust to outliers than the mean. In fact, the absolute loss function mentioned in the article can be shown to give a conditional median prediction (using the mid-point in case of non-uniqueness). So in the OP example, if the data set is known to contain outliers like properties that have extremely high or low value due to idiosyncratic reasons (e.g. former celebrity homes or contaminated land) then the absolute loss could be a wiser choice than least squares (of course, there are other ways to deal with this as well).
Worth mentioning here I think because the OP seems to be holding a particular grudge against the absolute loss function. It's not perfect, but it has its virtues and some advantages over least squares. It's a trade-off, like so many things.
[0] https://en.wikipedia.org/wiki/Quantile_regression https://en.wikipedia.org/wiki/Quantile_regression
- deleted 1y ago[deleted]
- easygenes 1y agoYeah. Squared error is optimal when the noise is Gaussian because it estimates the conditional mean; absolute error is optimal under Laplace noise because it estimates the conditional median. If your housing data have a few eight-figure outliers, the heavy tails break the Gaussian assumption, so a full quantile regression for, say, the 90th percentile—will predict prices more robustly than plain least squares.
- c7b 1y agoTrue. But it's worth mentioning that normality is only required for asymptotic inference. A lot of things that make least squares stand out, like being a conditional mean forecast, or that it's the best linear unbiased estimator, hold true regardless of the error distribution. My impression is that many tend to overestimate the importance of normality. In practice, I'd worry more about other things. The example in the OP, eg, if it were an actual analysis, would raise concerns about omitted variables. Clearly, house prices depend on more factors than size, eg location. Non-normality here could be just an artifact of an underspecified model.
- lupire 1y agoHow does an upcoming college student, or worse an already graduate, learn statistics like this, with depth of understanding of the meaning of the math, vs just plug an chugging cookbook formulas and "proving" theorems mechanically without the deep semantics?
- ayhanfuat 1y agoStatistical Rethinking is quite good in explaining this stuff. https://xcelab.net/rm/ https://xcelab.net/rm/
- disgruntledphd2 1y agoBasically all of the Andrew Gelman books are also good. Data Analysis... https://sites.stat.columbia.edu/gelman/arm/ https://sites.stat.columbia.edu/gelman/arm/ Regression and Other Stories: https://avehtari.github.io/ROS-Examples/ https://avehtari.github.io/ROS-Examples/ Wasserman's All of Statistics is a really good introduction to mathematical statistics (the Gelman stuff above are more practically and analytically focused). But yeah, it would probably be easier to find a good statistics course at a local university and try to audit it or do it at night.
- monkeyelite 1y agoDont take the “for engineers” version. > and "proving" theorems mechanically I think you’ve have a bad experience because writing a proof is explaining deep understanding.
- JadeNB 1y ago> I think you’ve have a bad experience because writing a proof is explaining deep understanding. I think your wording is the key—coming up with a proof is creating deep understanding, but writing a proof very much need not be explaining or creating deep understanding. Writing a proof can be done mechanically, by both instructor and student, and, if done so, neither demonstrates nor creates understanding. (Also, in statistics more than in almost any other mathematically based subject, while the rigorous mathematical foundations are important, a complete theoretical understanding of those foundations need not shed any light on the actual practice of statistics.)
- levocardia 1y agoQuantile regression is great, especially when you need more than just the average. A quantile model for, say, the 10th and 90th percentiles of something are really useful for decision-making. There is a great R package called qgam that lets you fit very powerful nonlinear quantile models -- one of R's "killer apps" that keeps me from using Python full-time.