3 ms·
I'm not making an arbitrary point. The end of my comment mentions word embeddings, which are learned from analyzing language and its usage in an unsupervised ma
by simulate-me 4y ago
I'm not making an arbitrary point. The end of my comment mentions word embeddings, which are learned from analyzing language and its usage in an unsupervised manner. You can use a word embedding to measure similarity between words. For instance, I just downloaded a pre-trained model of Stanford's GloVe embedding that was trained on 840 billion (yes, billion with a b) tokens. The cosine similarity can be used to measure how close two words are in the embedding. The cosine similarity of million and billion is 0.89. The similarity of duck and fuck is 0.27. The data indicates million and billion are much more similar than duck and fuck.
This result is intuitively obvious to me, which this article illustrates. Even if the resulting sentence is not factually true, a true-sounding sentence can be constructed by taking a sentence with the word million and replacing it with the word billion (in a majority of cases). This isn't true with duck and fuck, and duck is a noun and a verb used in generally different contexts than the word fuck.