4 ms·
For those interested in machine learning, the "self improvement" technique that the author talks about falls under semi-supervised learning, specifically self-t
by randomwalker 16y ago
For those interested in machine learning, the "self improvement" technique that the author talks about falls under semi-supervised learning, specifically self-training, which is apparently "still extensively used in the natural language processing community." http://en.wikipedia.org/wiki/Semi-supervised_learning http://en.wikipedia.org/wiki/Semi-supervised_learning http://pages.cs.wisc.edu/~jerryzhu/icml07tutorial.html http://pages.cs.wisc.edu/~jerryzhu/icml07tutorial.html
Semi-supervised learning is a good idea in this type of situation, given that unlabeled samples are far more abundant than labeled samples, but there are gotchas to watch out for. In general, SSL helps when your model of the data is correct and hurts when it is not.
Here's an example of what can go wrong in this particular application: let's say the word 'better' is mildly positive, but when it appears in high-confidence samples, it's usually because it appears together with the words 'business' and 'bureau', as in "I just reported Company X to the Better Business Bureau", i.e., strongly negative. This means that the new self-training samples containing the word better will all be negative, which will bias the corpus until eventually 'better' is treated as a strongly negative feature.
Occasional random human spot-checks of the high-confidence classifications would be useful :-) Also, self-training gives diminishing returns in accuracy, whereas the possibility for craziness remains, so turning it off after a while might be best.
A survey of semi-supervised learning:
http://www.cs.wisc.edu/~jerryzhu/pub/ssl_survey.pdf http://www.cs.wisc.edu/~jerryzhu/pub/ssl_survey.pdf
- spxdcz 16y agoThanks very much for the links (I wrote the article). I'm new to this area of analysis, so knowing the names for things really helps! I'll check out those links and delve further down the rabbit hole - thanks again.
- akshayubhat 16y agogiven that unlabeled samples are far more abundant than labeled samples It is also related to Transfer learning. This paper provides a good background http://www.stanford.edu/~hllee/icml07-selftaughtlearning.pdf http://www.stanford.edu/~hllee/icml07-selftaughtlearning.pdf It contains a nice graphic describing difference between Supervised, Semi-Supervised, Transfer and Self-Taught learning. Since twitter data consists of short texts. Unlabeled data from sources other than twitter could also be used, e.g. Google Buzz.
- jacquesm 16y agoSingle words may not be the best way to attack this problem though, multi-word expressions would do a better job. And then you could label the occurrence of 'reported*better business bureau' as a strong negative. Context is everything in natural language processing, and by dropping all context the problem becomes harder to solve.
- _delirium 16y agoLearning becomes problematic if you get too complex in your features though, especially if you go up to learning things like regexes. Simple ramping up in complexity, from e.g. word counts to word-pair counts (or other low-n n-grams) does give you gains sometimes, but there've been a number of cases where increasing the space like that gave surprisingly little/no gains, which is one reason the simple models keep being used (besides simplicity and speed).
- jacquesm 16y agoThat's true. I built a 'chatbot' long ago that used simple regexes (words+wildcards) to match incoming patterns. It worked well because at the higher levels of the conversation you'd use single words to guide to a portion of the conversation tree, and lower down you could make decisions on very specific differences in the input. For a classifier that's a less useful approach, but I think single words is too narrow. 3-grams is probably the sweet spot for something like this.
- StavrosK 16y agoThe problem with semi-supervised learning (or rather the way it's used here) is that, if you don't label the low-confidence samples yourself, but just leave it alone instead, it might diverge and start producing worse and worse results as it thinks it knows it guessed correctly, but in fact didn't. Basically, the problem is that you can't make a closed system learn from itself, without any outside feedback. The information has to come from somewhere. It's a bit like someone giving you two Chinese phrases and their translation (without you knowing any Chinese beforehand), and then leaving you to translate a whole book. You will start guessing, based on what you already know, and by the end you'll have arrived to a (totally incorrect) interpretation of what you think each ideogram means.