4 ms·
I respect your skepticism :) It's easy to imagine exceptions to the idea of a simple numerical word-scoring algorithm... Of course, a word like "bad" might be
by benjismith 6y ago
I respect your skepticism :)
It's easy to imagine exceptions to the idea of a simple numerical word-scoring algorithm...
Of course, a word like "bad" might be used ironically, or in some other slang-sense, with a different literal meaning on the page...
But that's totally fine. In principle, the word2vec algorithm is designed to cope with ambiguities like that.
When you analyze billions of words of prose, you can build a model of word-associativity that captures the superposition of all those different word-senses, and the contexts where they tend to appear on the page.
After a big crazy machine-learning process, each word is modeled as a vector in 300-dimensional space, with a vast network of associations and relationships between the other words in the vector-space, based on the way those words are used together in typical English grammar.
When we score the emotional valence of a particular word, we use a "word-vector" technique where those ambiguities are basically already priced into the scoring calculation. Words with a "less ambiguous" sentiment score (joy, paradise, ..., agony, depression) have their lack-of-ambiguity baked into the formula already.
Extreme scores are reserved for words with unambiguous intensity.
But the important thing is: we're not really as concerned about the numerical scores of individual words as we are with the shifting balance of those sentiment scores over the course of a long document.
It's not a perfect way of scoring sentiment of individual words, but it's REALLY reliable for estimating the basic structure of a narrative.