3 ms·
I decided there's no such thing as a good NLP library. In my other apps, I usually use several libraries at once, then make them vote. :) Works out better tha
by photorized 12y ago
I decided there's no such thing as a good NLP library. In my other apps, I usually use several libraries at once, then make them vote. :) Works out better than Stanford NLP.
I do need to add some intelligence to skim.io though.
- doppenhe 12y agoYes 100% agreed. One of our goals with Algorithmia is to have all these libraries already there , preloaded, available and standardized around a similar api signature. This way you can do exactly what you described or quick A/B testing on them.
- hnriot 12y agoThis is called ensemble classification where you feed the outputs of multiple classifiers as features into an ensemble classifier that produces the final result. How are you using the Stanford NLP? That's all GPL? There are alternatives you could look at for sentiment analysis but short "documents" like those referenced will always produce poor results because there's just not enough signal to work with. The training models need to have vocabulary overlap with the documents (at least for word features); try TextBlob which uses a lexicon approach rather than a classifier, or try rolling your own with an off-the-shelf SVM and pull labeled training data from one of the many sources (or generate your own using Crowdflower.) Small documents (tweets/titles etc) pose unique challenges, especially when there's irony or sarcasm involved or implicit sentiment through pragmatic knowledge. For example knowing Sarah Palin and how she's regarded automatically gives a person a head start in determining the sentiment of a short document with her name. This kind of pragmatic knowledge is hard for classifiers to learn.
- walterbell 12y agoIn the example above, could social network analysis (e.g. https://en.wikipedia.org/wiki/NodeXL https://en.wikipedia.org/wiki/NodeXL) be used to profile Sarah Palin, then combined with text classification?
- hnriot 12y agoIt's an option, not sure about NodeXL, wikipedia is already available in structured form in Freebase and DBPedia, but understanding something so complex as a reputation is beyond current machine learning. Bringing the pragmatic background knowledge to sentiment analysis is going to be one of the differentiations. We got 65% easily with a lexicon, we got 75% with SVMs and gobs of training data, we've gone past that with hierarchical aspect models and technologies like word vectors, but the problem of making improvements gets exponentially harder. Before social networks can play a role in sentiment analysis likely we'll see breakthroughs in coreference and similar problems that will help eek out more signal from training data. We can certainly use the "hive mind" to assist in this problem, even something as simple as collaborative filtering can help.