3 ms·
What prevents you from using a different distance than Euclidean distance?
by CardenB 5y ago
What prevents you from using a different distance than Euclidean distance?
- danuker 5y agoSuppose you use prices and square meters, to cluster apartments. Suppose you also use a simple metric, like: distance <- sqrt(d_price*2 + d_area*2) Without normalizing the units somehow, so that they have a meaningfully similar numeric expression, you are essentially only clustering by price, which is in the tens- or hundreds-of-thousands. You use very little information about the area, because that is in the hundreds, at most thousands.
- jacquesm 5y agoThat is not a proper distance function.
- danuker 5y agoWhat do you mean? Is it not the Euclidean distance? https://en.wikipedia.org/wiki/Euclidean_distance https://en.wikipedia.org/wiki/Euclidean_distance
- nerdponx 5y agoThis has nothing to do with using non-Euclidean distance, though. This is a separate issue related to having features on different scales.
- danuker 5y agoIt was an example metric with coefficients of 1 for each feature. The point I was making was that the different scale of the features influences the result by a lot.
- deleted 5y ago[deleted]
- amilios 5y agoThe way k-means is constructed is not based on distances. K-means minimizes within-cluster variance. If you look at the definition of variance, it is identical to the sum of squared Euclidean distances from the centroid
- nerdponx 5y agoDo you lose convergence guarantees if you minimize "sum of squared distances from the centroid" using your other-than-Euclidean distance metric?