7 ms·
Why does a least squares fit appear to have a bias when applied to simple data?
- deleted 9mo ago[deleted]
- evanb 9mo agoor 3) there's a misunderstanding about ordinary least-squares.
- bschmidt25002 9mo ago[dead]
- charlieyu1 9mo agoIf you plot the regression line of y against x, and also x against y, you would get two different lines. I found it in the middle of teaching a stats class, and feel embarrassed. I guess normalising is one way to remove the bias.
- lambdaone 9mo agoYou are absolutely correct that the difference between y against x and x against y fitting perfectly demonstrates why the bias exists, but the correct way to remove the bias is not normalization, but to use a coordinate-independent regression technique. See the other comments by many other commenters for details.
- dllu 9mo agoYou can think of it as: linear regression models only noise in y and not x, whereas ellipse/eigenvector of the PCA models noise in both x and y.
- analog31 9mo agoThat brings up an interesting issue, which is that many systems do have more noise in y than in x. For instance, time series data from an analog-to-digital converter, where time is based on a crystal oscillator.
- GardenLetter27 9mo agoThis fact underlies a lot of causal inference.
- randrus 9mo agoI’m not an SME here and would love to hear more about this.
- jjk166 9mo agoWell yeah, x is specifically the thing you control, y is the thing you don't. For all but the most trivial systems, y will be influenced by something besides x which will be a source of noise no matter how accurately you measure. Noise in x is purely due to setup error. If your x noise was greater than your y noise, you generally wouldn't bother taking the measurement in the first place.
- bravura 9mo ago“ If your x noise was greater than your y noise, you generally wouldn't bother taking the measurement in the first place.” Why not? You could still do inference in this case.
- jjk166 9mo agoYou could, and maybe sometimes you would, but generally you won't. If at all possible, it makes a lot more sense to improve your setup to reduce the x noise, either with a better setup or changing your x to be something you can better control.
- 9mo ago
- bschmidt25013 9mo ago[dead]
- sega_sai 9mo agoThe least squares and pca minimize different loss functions. One is sum of squares of vertical(y) distances, another is is sum of closest distances to the line. That introduces the differences.
- ryang2718 9mo agoI find it helpful to view least as fitting the noise to a Gaussian distribution.
- LudwigNagasena 9mo agoOLS estimator is the minimum-variance linear unbiased estimator even without the assumption of Gaussian distribution.
- rjdj377dhabsn 9mo agoYes, and if I remember correctly, you get the Gaussian because it's the minimum entropy (least additional assumptions about the shape) continuous distribution given a certain variance.
- porridgeraisin 9mo agoAnd given a mean.
- contravariant 9mo agoBoth of these do, in a way. They just differ in which gaussian distribution they're fitting to. And how I suppose. PCA is effectively moment matching, least squares is max likelihood. These correspond to the two ways of minimizing the Kullback Leibler divergence to or from a gaussian distribution.
- MontyCarloHall 9mo agoThey both fit Gaussians, just different ones! OLS fits a 1D Gaussian to the set of errors in the y coordinates only, whereas TLS (PCA) fits a 2D Gaussian to the set of all (x,y) pairs.
- gpcz 9mo agoYou would probably get what you want with a Deming regression.
- tomp 9mo agoLinear Regression a.k.a. Ordinary Least Squares assumes only Y has noise, and X is correct. Your "visual inspection" assumes both X and Y have noise. That's called Total Least Squares.
- emmelaich 9mo agoYep, to demonstrate, tilt it (swap x and y) and do it again. Maybe this is what TLS does?
- srean 9mo ago>(swap x and y) and do it again. This is a great diagnostic check for symmetry. > Maybe this is what TLS does? No, swapping just exchanges the relation. What one needs to do is to put the errors in X and errors in Y in equal footing. That's exactly what TLS does. Another way to think about it is that the error of a point from the line is not measured as a vertical drop parallel to Y axis but in a direction orthogonal to the line (so that the error breaks up in X and Y directions). From this orthogonality you can see that TLS is PCA (principal component analysis) in disguise.
- esafak 9mo agoThere is an illustration in https://en.wikipedia.org/wiki/Total_least_squares https://en.wikipedia.org/wiki/Total_least_squares
- a3w 9mo agoMy head canon: If the true value is medium high, any random measurements that lie even further above are easily explained, as that is a low ratio of divergence. If the true value is medium high, any random measurements that lie below by a lot are harder to explain, since their (relative, i.e.) ratio of divergence is high. Therefore, the further you go right in the graph, the more a slightly lower guess is a good fit, even if many values then lie above it.
- theophrastus 9mo agoHad a QuantSci Prof who was fond of asking "Who can name a data collection scenario where the x data has no error?" and then taught Deming regression as a generally preferred analysis [1] [1] https://en.wikipedia.org/wiki/Deming_regression https://en.wikipedia.org/wiki/Deming_regression
- jmpeax 9mo agoFrom that wikipedia article, delta is the ratio of y variance to x variance. If x variance is tiny compared to y variance (often the case in practice) then will we not get an ill-conditioned model due to the large delta?
- kevmo314 9mo agoIf you take the limit of delta -> infinity then you will get beta_1 = s_xy / s_xx which is the OLS estimator. In the wiki page, factor out delta^2 from the sqrt and take delta to infinity and you will get a finite value. Apologies for not detailing the proof here, it's not so easy to type math...
- moregrist 9mo agoMost of the time, if you have a sensor that you sample at, say 1 KHz and you’re using a reliable MCU and clock, the noise terms in the sensor will vastly dominate the jitter of sampling. So for a lot of sensor data, the error in the Y coordinate is orders of magnitude higher than the error in the X coordinate and you can essentially neglect X errors.
- sigmoid10 9mo agoThat is actually the case in most fields outside of maybe clinical chemistry and such, where Deming became famous for explaining it (despite not even inventing the method). Ordinary least squares originated in astronomy, where people tried to predict movement of celestial objects. Timing a planet's position was never an issue (in fact time is defined by celestian position), but getting the actual position of a planet was. Total least squares regression also is highly non-trivial because you usually don't measure the same dimension on both axes. So you can't just add up errors, because the fit will be dependent on the scale you chose. Deming skirts around this problem by using the ratio of variances of errors (division also works for different units), but that is rarely known well. Deming works best when the measurement method for both dependent and independent variable is the same (for example when you regress serum levels against one another), meaning the ratio is simply one. Which of course implies that they have the same unit. So you don't run into the scale-invariance issues, which you would in most natural science fields.
- em500 9mo agoSorry for my negativity / meta comment on this thread. From what I can tell the stackexchange discussion in the submission already to provides all the relevant points to be discussed about this. While the asymmetry of least squares will probably be a bit of a novelty/surprise to some, pretty much anything posted here is more or less a copy of one of the comments on stackexchange. [Challenge: provide a genuinely novel on-topic take on the subject.]
- oh_my_goodness 9mo agoThis answer is too grown-up for the forum.
- deleted 9mo ago[deleted]
- bee_rider 9mo agoThe stackexchange discussion already provides a good answer. I think there is not much to be said. It is not a puzzle for us to solve, just a neat little mathematical observation.
- SubiculumCode 9mo agoBut bringing it up as a topic, aside from being informative, allows for more varied conversation that is allowed on stack exchange, like exploring alternative modeling approaches. It may not have happened, but the possibility can only present itself given the opportunity
- efavdb 9mo agoMany times I've looked at the output of a regression model, seen this effect, and then thought my model must be very bad. But then remember the points made elsewhere in thread. One way to visually check that the fit line has the right slope is to (1) pick some x value, and then (2) ensure that the noise on top of the fit is roughly balanced on either side. I.e., that the result does look like y = prediction(x) + epsilon, with epsilon some symmetric noise. One other point is that if you try to simulate some data as, say y = 1.5 * x + random noise then do a least squares fit, you will recover the 1.5 slope, and still it may look visually off to you.
- fluidcruft 9mo agoMaybe comparing plots of residuals makes it clearest.
- paulfharrison 9mo agoA note mostly about terminology: The least squares model will produce unbiassed predictions of y given x, i.e. predictions for which the average error is zero. This is the usual technical definition of unbiassed in statistics, but may not correspond to common usage. Whether x is a noisy measurement or not is sort of irrelevant to this -- you make the prediction with the information you have.
- deleted 9mo ago[deleted]
- noncovalence 9mo agoThis problem is usually known as regression dilution, discussed here: https://en.wikipedia.org/wiki/Regression_dilution https://en.wikipedia.org/wiki/Regression_dilution
- thaumasiotes 9mo agoIs it? The wikipedia article says that regression dilution occurs when errors in the x data bias the computed regression line. But the stackexchange question is asking why an unbiased regression line doesn't lie on the major axis of the 3σ confidence ellipse. This lack of coincidence doesn't require any errors in the x data. https://stats.stackexchange.com/a/674135 https://stats.stackexchange.com/a/674135 gives a constructed example where errors in the x data are zero by definition. Unless I'm misunderstanding something?
- bgbntty2 9mo agoI haven't dealt with statistics for a while, but what I don't get is why squares specifically? Why not power of 1, or 3, or 4, or anything else? I've seen squares come up a lot in statistics. One explanation that I didn't really like is that it's easier to work with because you don't have to use abs() since everything is positive. OK, but why not another even power like 4? Different powers should give you different results. Which seems like a big deal because statistics is used to explain important things and to guide our life wrt those important things. What makes squares the best? I can't recall other times I've seen squares used, as my memories of my statistics training is quite blurry now, but they seem to pop up here and there in statistics relatively often, it seems.
- djaouen 9mo agoSquares are preferred because that is the same as minimizing Euclidean Distance, which is defined as sqrt((x2-x1)^2).
- lucketone 9mo agosqrt((x2-x1)^2) == x2-x1 I think you meant sqrt(x^2+y^2)
- deleted 9mo ago[deleted]
- stirfish 9mo agoI haven't done it in a while, but you can do cubes (and more) too. Cubes would be the L3 norm, something about the distance between circles (spheres?) in 3d space? I need to read about norms again to tell you why or when to choose that, but I know the Googlable term is "vector norms" I remember one is Manhattan distance, next is as-the-crow-flies straight line distance, next is if you were a crow on the earth that can also swim in a straight line underwater, and so on
- qnleigh 9mo agoLeast squares is guaranteed to be convex [0]. At least for linear fit functions there is only one minimum and gradient descent is guaranteed to take you there (and you can solve it with a simple matrix inversion, which doesn't even require iteration). Intuitively this is because a multidimensional parabola looks like a bowl, so it's easy to find the bottom. For higher powers the shape can be more complicated and have multiple minima. But I guess these arguments are more about making the problem easy to solve. There could be applications where higher powers are worth the extra difficulty. You have to think about what you're trying to optimize. [0] https://math.stackexchange.com/questions/483339/proof-of-convexity-of-linear-least-squares https://math.stackexchange.com/questions/483339/proof-of-con...
- djaouen 9mo agoThis is probably obvious, but there is another form of regression that uses Mean Absolute Error rather than Squared Error as this approach is less prone to outliers. The Math isn’t as elegant, tho.
- Ericson2314 9mo agoYes people want to mentally rotate, but that's not correct. This is not a "geometric" coordinate system independent operation. IMO this is a basic risk to graphs. It is great to use imagery to engage the spatial reasoning parts of our brain. But sometimes, it is deceiving — like this case —because we impute geometric structure which isn't true about the mathematical construct being visualized.
- ModernMech 9mo agoThis is why my favorite best fit algorithm is RANSAC.
- taylorius 9mo agoI think the linear least squares is like a shear, whereas the eigenvector is a rotation.
- deleted 9mo ago[deleted]
- anArbitraryOne 9mo agoShear with one component always in the y direction
- kleiba 9mo ago> So, instead, I then diagonalized the covariance matrix to obtain the eigenvector that gives the direction of maximum variance. ...as one does...
- abanana 9mo agoWithout knowing the meaning of that level of mathematical jargon, it feels like a "reticulating splines" sort of line. Makes me want to copy it and use it somewhere.
- anthk 9mo agoFrom T3X, intro to Statistics: https://t3x.org/klong-stat/toc.html https://t3x.org/klong-stat/toc.html Klong language to play with: https://t3x.org/klong/ https://t3x.org/klong/
- PythonPeak 9mo ago[dead]