4 ms·
I ran the systems over the CoNLL 2000 PoS data (see training time), consisting of roughly 10k sentences, split into 8k training and 2k testing sentences (by the
by fnl 12y ago
I ran the systems over the CoNLL 2000 PoS data (see training time), consisting of roughly 10k sentences, split into 8k training and 2k testing sentences (by the organizers of the ST).
However, this review is not about "who has the best accuracy", as I believe that is mostly feature-dependent, and you can find entire tomes in the literature and Web discussing this point back and forth (to little objective avail, unless part of a shared task with unseen data, IMO).
My interest was in training times, ease of defining features, model flexibility, and token throughput. And on these points, the systems I looked into have very large and important differences.