4 ms·
I implemented the new t-SNE in sklearn, so I've got some experience in reading these diagrams. Unfortunately, as wonderful as the algorithm is, it's extremely h
by juxtaposicion 11y ago
I implemented the new t-SNE in sklearn, so I've got some experience in reading these diagrams. Unfortunately, as wonderful as the algorithm is, it's extremely hard to interpret what it means rigorously. I've seen many diagrams that look like this one -- and they were generated from actual noise. So take the plots with a big grain of salt :)
I'd be interested in seeing more direct evidence, like SVD factorizing the PMI matrix (which is what similar to what word2vec is doing) and seeing how much of the variance is explained by the first components. If you want to do this, check out: https://minhlab.wordpress.com/2015/06/08/a-new-proof-for-the-equivalence-of-word2vec-skip-gram-and-shifted-ppmi/ https://minhlab.wordpress.com/2015/06/08/a-new-proof-for-the...
- bearzoo 11y agoi once read that high perplexity can generate embeddings that are very tightly bound on a unit circle...not sure if this is what is going on
- Houshalter 11y agoThe diagram shown is only a visualization. The actual word vectors have many dimensions. To reduce them to 2 dimensions, they use a method which tries to keep vectors that are similar as close to each other as possible, but also unsimilar words apart. This creates the shape seen on the scatter plot. Just looking at the scatter plot by itself doesn't tell you anything about the underlying data.
- andreasvc 11y ago> Just looking at the scatter plot by itself doesn't tell you anything about the underlying data. Well if that were the case it would be perfectly pointless to make such a visualization... The goal of dimensionality reduction is to provide a useful summarization of the data; it is a valid question to ask to what degree it is successful at that.
- bearzoo 11y agoI did not claim anything about the underlying data.. The 2 dimensional embeddings were forced into the unit circle because the 'perplexity' hyper parameter for the t-sne was set too high. From the guy who helped make t-sne: When I run t-SNE, I get a strange ‘ball’ with uniformly distributed points? This usually indicates you set your perplexity way too high. All points now want to be equidistant. The result you got is the closest you can get to equidistant points as is possible in two dimensions. If lowering the perplexity doesn’t help, you might have run into the problem described in the next question. Similar effects may also occur when you use highly non-metric similarities as input.
- perone 11y agoHi, thanks for the feedback. There are three main points that makes me believe that the clusters aren't artificial: the first one is that I've made the clusters with DBSCAN on the original data (100-d word vectors) and not after the t-SNE embedding. The second point is that I manually inspected some clusters and they make sense when compared to the similarity queries on the word2vec model (take a look on the star names for instance). The third point: I took the folios from the two main clusters (red/blue) (https://www.reddit.com/r/MachineLearning/comments/419e5a/voynich_manuscript_word_vectors_and_tsne/cz0nwlw https://www.reddit.com/r/MachineLearning/comments/419e5a/voy...) and they seem to match with the folios from the two languages hypothesis.
- haddr 11y agoI had the same experience. Whenever I use t-SNE with noisy data or data grouped densely around some point, they tend to have a very similar visual structure as the example given in this experiment: a circular shape with some points making small "clusters" but with rather no significant meaning.