4 ms·
One sentiment analysis, there seems to be a lot of false negatives - at least when parsing Techcrunch. For example, all these were tagged as 'negative': - Tim
by photorized 12y ago
One sentiment analysis, there seems to be a lot of false negatives - at least when parsing Techcrunch. For example, all these were tagged as 'negative':
- Timely Turns Your Calendar Into A Time Tracker
- Audi Tests Self-Driving Cars On Florida’s Roads
- Twitter Acquires Password Security Startup Mitro, Open Sources Its Product
- doppenhe 12y agoWe used the Stanford NLP library with their training data set which is considered to be one of the better ones. I did notice false negatives as well but it can definitely be trained to be more accurate.
- photorized 12y agoI decided there's no such thing as a good NLP library. In my other apps, I usually use several libraries at once, then make them vote. :) Works out better than Stanford NLP. I do need to add some intelligence to skim.io though.
- doppenhe 12y agoYes 100% agreed. One of our goals with Algorithmia is to have all these libraries already there , preloaded, available and standardized around a similar api signature. This way you can do exactly what you described or quick A/B testing on them.
- hnriot 12y agoThis is called ensemble classification where you feed the outputs of multiple classifiers as features into an ensemble classifier that produces the final result. How are you using the Stanford NLP? That's all GPL? There are alternatives you could look at for sentiment analysis but short "documents" like those referenced will always produce poor results because there's just not enough signal to work with. The training models need to have vocabulary overlap with the documents (at least for word features); try TextBlob which uses a lexicon approach rather than a classifier, or try rolling your own with an off-the-shelf SVM and pull labeled training data from one of the many sources (or generate your own using Crowdflower.) Small documents (tweets/titles etc) pose unique challenges, especially when there's irony or sarcasm involved or implicit sentiment through pragmatic knowledge. For example knowing Sarah Palin and how she's regarded automatically gives a person a head start in determining the sentiment of a short document with her name. This kind of pragmatic knowledge is hard for classifiers to learn.
- walterbell 12y agoIn the example above, could social network analysis (e.g. https://en.wikipedia.org/wiki/NodeXL https://en.wikipedia.org/wiki/NodeXL) be used to profile Sarah Palin, then combined with text classification?
- hnriot 12y agoIt's an option, not sure about NodeXL, wikipedia is already available in structured form in Freebase and DBPedia, but understanding something so complex as a reputation is beyond current machine learning. Bringing the pragmatic background knowledge to sentiment analysis is going to be one of the differentiations. We got 65% easily with a lexicon, we got 75% with SVMs and gobs of training data, we've gone past that with hierarchical aspect models and technologies like word vectors, but the problem of making improvements gets exponentially harder. Before social networks can play a role in sentiment analysis likely we'll see breakthroughs in coreference and similar problems that will help eek out more signal from training data. We can certainly use the "hive mind" to assist in this problem, even something as simple as collaborative filtering can help.
- JSno 12y agowhat 'negative' are you guys talking about? I just noticed your system gave several keywords for one article. What's about positive/negative? thanks
- photorized 12y agoI was referring to the smileys displayed next to each article, which were probably meant to indicate sentiment.
- yaeger 12y agoI just checked the arstechnica one and all items there were also tagged as negative.