4 ms·
It's not only about the letters. "Million" and "billion" are both numeric concepts. Numeric concepts are only a small subset of things expressible by language.
by simulate-me 4y ago
It's not only about the letters. "Million" and "billion" are both numeric concepts. Numeric concepts are only a small subset of things expressible by language. So million and billion are close because they are spelled similarly and they deal with the same conceptual "thing." Duck and Fuck don't occupy the same conceptual category, so the distance between the two words is large despite being spelled similarly.
One interesting application of neural networks is the creation word embeddings. ML models are trained to place words in a vector space, which is useful for measuring distance between words, performing arithmetic on words, or finding the closet word. Using an embedding allows your to formalize the "distance" between words, and perform fun tricks like King + Woman = Queen.
- moate 4y ago>> So million and billion are close because they are spelled similarly and they deal with the same conceptual "thing. Citation needed. Is this based on any sort of formal linguistic/anthropological reasoning or just like, your opinion? Also, the article's whole point is "people are bad at numbers" and your argument that "linguistically, all numbers close" isn't really true. Why don't rhyming words in English feel "close"? Also, "duck" and "fuck" are both verbs involving a thing you would do with your body, so why isn't that closeness? If you're going to point out that language feels arbitrary by making arbitrary points, you're going to be "right" but you're not actually saying much. Language can be both specific and arbitrary (it's a means of expressing both objective and subjective concepts) so an argument in favor of doing your best when seeking to be objective seems pretty reasonable.
- simulate-me 4y agoI'm not making an arbitrary point. The end of my comment mentions word embeddings, which are learned from analyzing language and its usage in an unsupervised manner. You can use a word embedding to measure similarity between words. For instance, I just downloaded a pre-trained model of Stanford's GloVe embedding that was trained on 840 billion (yes, billion with a b) tokens. The cosine similarity can be used to measure how close two words are in the embedding. The cosine similarity of million and billion is 0.89. The similarity of duck and fuck is 0.27. The data indicates million and billion are much more similar than duck and fuck. This result is intuitively obvious to me, which this article illustrates. Even if the resulting sentence is not factually true, a true-sounding sentence can be constructed by taking a sentence with the word million and replacing it with the word billion (in a majority of cases). This isn't true with duck and fuck, and duck is a noun and a verb used in generally different contexts than the word fuck.