4 ms·
Very interesting read. I interpret this in the way that clustering (eg HDBSCAN) on UMAP-projected data makes some sense at least (contrary to tSNE), are there a
by jmrko 7y ago
Very interesting read. I interpret this in the way that clustering (eg HDBSCAN) on UMAP-projected data makes some sense at least (contrary to tSNE), are there any differing opinions on this? Interesting related discussions: https://stats.stackexchange.com/questions/263539/clustering-on-the-output-of-t-sne https://stats.stackexchange.com/questions/263539/clustering-...
- jointpdf 7y agoHere’s a pretty comprehensive answer on the topic from the original UMAP author: https://github.com/lmcinnes/umap/issues/25 https://github.com/lmcinnes/umap/issues/25 Clustering the output of UMAP is also given a nice tutorial in the docs: https://umap-learn.readthedocs.io/en/latest/clustering.html https://umap-learn.readthedocs.io/en/latest/clustering.html Basically, the answer is yes you can do this, but verify and analyze the output to ensure it makes sense (e.g. coloring points by known features/labels). For example, if you have a small number of points in the dataset (<1000), UMAP tends to display a dense cluster that is quite separated from the remaining data. However, this apparent cluster is spurious and contains noisy data points that UMAP couldn’t “figure out what to do with” (they are similar in their dissimilarity to the other data).