5 ms·
> that's really the most obvious cost function To me Deming regression is the most obvious. It actually took me a long time to realize that x~y and y~x are in
by pps43 6y ago
> that's really the most obvious cost function
To me Deming regression is the most obvious. It actually took me a long time to realize that x~y and y~x are in most cases different lines.
- quietbritishjim 6y agoOops, you're absolutely right. I remember back at secondary school a teacher going through exactly the process I described, with him asking us what the best choice of line would be. The first choice someone (not me sadly) suggested was indeed this one. For others that, like me, don't already know it (at least by name): the Deming cost function is the sum of perpendicular distances to the points from the line, as opposed to measuring only the vertical component. Edit: Actually it looks like Deming regression is a bit more subtle and statistical than just sum of perpendicular distances. A treatment of linear regression from a statistical perspective of trying untangle some random noise is very worthy, but I'd save it for a second lesson, following a first lessons that's just an unmotivated "let's choose the line that's somehow closest to all the points".
- em500 6y agoDeming regressions have some serious drawbacks compared to OLS, most importantly that it's not scale/unit-invariant: if you rescale your (e.g. use house area in square metres instead of square feet to predict house prices) you get different results. This may be less of an issue for the machine leaner who's only interested in point predictions, but it's a serious concern for us old fashioned statisticians who are more interested in inference.
- autokad 6y agoi wouldn't dare to say its a 'serious drawback'. as soon as you add ridge to regression (something very common for statisticians to do), it's also no longer scale/unit-invariant. Edit: this is also not just true for ridge, but lasso as well
- em500 6y agoPersonally I consider the lack of scale-invariance one of the main drawbacks of most common regularizers too. Again, not a big deal if all you're after are y-hat, a bit more concerning if you're interested in beta-hat.
- pps43 6y agoI would not call it a drawback. In regular regression you make an assumption that x is known exactly, and all errors are in y. In Deming regression you make another assumption, about the ratio of their variances (δ). If you change units, you need to change δ as well, and then there's no difference.
- tomrod 6y agoThis is also called Total Least Squares -- IIRC it was developed for use when you have measurement error in X (least squares assumes no measurement error in X, all error is in Y).