8 ms·
The examples are all in 2D because it allows someone to visualise what's going on. In higher dimensions things get messier and you have to rely on cluster quali
by lmcinnes 10y ago
The examples are all in 2D because it allows someone to visualise what's going on. In higher dimensions things get messier and you have to rely on cluster quality measures ... which are often bad (or, more often, defined to be the objective function that a particular clustering algorithm optimizes, and hence give a false sense of how "well" the clustering has done). I've worked with HDBSCAN quite successfully on mid-range dimensionality data (50 to 100 dimensions). I agree that it doesn't work as well on truly high dimensional data, but then little can once the curse of dimensionality truly kicks in. Your best bet, at that point, is the assume that your data actually lies on some lower dimensional manifold (i.e. the intrinsic dimensionality is much lower) and apply suitable dimension reduction techniques (t-SNE, Robust PCA, etc.) and then cluster.
- visarga 10y agoOr to apply neural nets to compute some kind of vector representation that contains semantic information.