3 ms·
https://en.wikipedia.org/wiki/T-distributed_stochastic_neighbor_embedding https://en.wikipedia.org/wiki/T-distributed_stochastic_neigh... In case someone wonde
by throwaway_2047 6y ago
https://en.wikipedia.org/wiki/T-distributed_stochastic_neighbor_embedding https://en.wikipedia.org/wiki/T-distributed_stochastic_neigh...
In case someone wondering, like me, what TSNE is.
Which I still don't understand after reading
- polm23 6y agoIf you have a machine learning model and you want to see what things it thinks are similar, you can use TSNE to visualize that by rendering similar points close together in two or three dimensions. UMAP is another method used for similar purposes.
- ArnoVW 6y agoit's an algorithm for projecting data to lower dimension. I.e. you have an Excel sheet with 20.000 lines (representing customers for ex) and 200 columns (representing blood pressure, height, weight, etc). What you want to do is "visualise" those 20.000 points in 2D or 3D so you can get an idea of how the data is distributed. So you use t-SNE to "compress" those 200 columns to 2 or 3, and you display that. Traditionally you would use Primary Component Analysis, but that only uses linear projection, and will not be able to project data that has non-linear relationships in the distributions. Another algorithm, sometimes more powerfull and scalable is LargeViz.
- lmeyerov 6y agoUMAP has largely replaced t-SNE in our toolkit as one of our top go-to viz pipelines. Unlike most examples out there, we post-process with k-nn to expose the graph of correlations over arbitrary data sets -- bank accounts fraud scores, cancer protein mutations, twitter bots, malware files, etc. -- and then investigate. Algorithms like UMAP figure out this connectivity anyways (see also: TDA), and useful for guiding subsequent explorations. If you're doing an interactive analysis, like looking at data in a Jupyter notebook, super powerful to expose that inferred connectivity and make it interactive (on-the-fly filtering, clustering, etc.) on it. Tool-wise, we do it in a few lines over tables with many rows/columns via end-to-end GPU acceleration using https://www.RAPIDS.ai https://www.RAPIDS.ai (GPU dataframes + UMAP) + Graphistry (GPU viz, which we make).
- dumb1224 6y agoDo you mean principal component analysis? My naive understanding after reading the original paper is that the algorithm is training a transformation to project the high dimensional data into low dimensions by best preserving both the global and local proximity. So that the samples similar at high dim space should also be close in the low dim one. It has some assumption of the distribution of the data in low dim so it won't be a random guess. It's using the t-distribution at low dim hence the name t-SNE. Correct me if any mistakes.
- jlg23 6y agoThis google tech talk is a pretty gentle introduction: https://www.youtube.com/watch?v=RJVL80Gg3lA https://www.youtube.com/watch?v=RJVL80Gg3lA