3 ms·
Is there any fundamental reason to measure discrepancy by abs(s - x_i)^2 rather than say by abs(s - x_i)^1.5? Is something special about 2 in this context, or i
by RoboTeddy 9y ago
Is there any fundamental reason to measure discrepancy by abs(s - x_i)^2 rather than say by abs(s - x_i)^1.5? Is something special about 2 in this context, or is it just a social convention that seems to work pretty well?
- moultano 9y agoYes! Lots of them actually. 1. The gaussian distribution is important (all sums converge to it) and the squared difference from the mean characterizes it. 2. Euclidean space is important, and squared errors stay the same if you rotate everything. (Other errors don't). 3. Linear regression with squared error has a closed form solution. Other types of models converge very fast when using it because the further you are from optimal, the bigger your gradient is.
- deleted 9y ago[deleted]
- nalourie 9y agoTo expand on the other answers, if you want a notion of angle between things then you need the distance to be defined by an inner product (dot product). This leads to the square. So, in some sense squared error generalized our geometric intuitions the best.
- ssalazar 9y agoAnother thought- x^2 is differentiable for all x, so its easier to reason about and work with e.g. gradient descent than abs(x)^p for arbitrary p.
- FabHK 9y agoIt's also pretty and easy to minimise, as the derivative of the square is just linear.
- jules 9y agoThe ^2 gives the normal distribution, which is special for many reasons. With ^1 you get the Laplace distribution. This is widely used when the data is sparse (has many elements exactly 0). This is less sensitive to ourliers and in image reconstruction it gives sharper images than ^2.