3 ms·
I'm surprised that no one mentioned this paper that first evaluated this approach to using Wikipedia data: http://www.aaai.org/Papers/IJCAI/2007/IJCAI07-259.pdf
by cybernytrix 16y ago
I'm surprised that no one mentioned this paper that first evaluated this approach to using Wikipedia data: http://www.aaai.org/Papers/IJCAI/2007/IJCAI07-259.pdf http://www.aaai.org/Papers/IJCAI/2007/IJCAI07-259.pdf
That said, the major drawback of using Wikipedia is the size. If this approach is to be used for all words (not just Apple) then the total training corpus will be several GBs. Definitely not practical...
- nl 16y agoWhat's not practical about it? GBs of data are pretty easy to handle these days
- cybernytrix 16y agoGBs and TBs of data is common, not for this task. All you are doing is Word Sense Disambiguation and there are algorithms to do WSD that work with much much smaller training sets. Just don't think that the exponential increase in training data is justified...