3 ms·
Weird because POS tagging takes linear time in the number of words.
by botexpert 10y ago
Weird because POS tagging takes linear time in the number of words.
- gok 10y agoThat's assuming a lot... there are many, many ways to do part of speech tagging. You could imagine a slow implementation that exhaustively considered every possible POS tagging for an input sequence (would be O(k^N) where k is # of parts of speech)
- botexpert 10y agoThere's not really that much assumption. Given the fact that good and simple linear time algorithms exist for decades it's very unlikely that they use some ridiculous permutation enumeration algorithm. Even a Microsoft Research project has a fairly complex reductions implementation of a POS tagger that is blazingly fast and production ready (learning to search interface). [1] [1]: https://github.com/JohnLangford/vowpal_wabbit/wiki https://github.com/JohnLangford/vowpal_wabbit/wiki
- nl 10y agoYou could imagine a slow implementation that exhaustively considered every possible POS tagging for an input sequence (would be O(k^N) where k is # of parts of speech) You could.. if you were crazy. No sane implementation would do that, and POS tagging isn't some brand new thing where people make that kind of mistake.
- kylebgorman 10y agoAnything that uses the standard decoding strategy (Viterbi) is quadratic in the number of "states" (i.e., tags), though, and that easily dominates. (Nowadays, many people just use greedy---linear in the number of states---decoding because it rarely is much worse.)