4 ms·
Using accuracy to measure PoS taggers makes results look good, but is ensnaring due to their huge bias: Tagging every word by the majority tag found during trai
by fnl 11y ago
Using accuracy to measure PoS taggers makes results look good, but is ensnaring due to their huge bias: Tagging every word by the majority tag found during training and everything else as either NNP or NNPS (with suffix -s) means the statistical baseline is already well beyond 90% accuracy. However, my point was that from the results shown its not clear to me if the gains in attachment score you saw when using Gold Standard PoS tags would be lost in a "real-world" usage when you have to rely on the tagger's own PoS tags. In such a case, it could be that your embeddings contribute much less "new knowledge" than what you see in your results, using independent (Gold) PoS tags. This might be mitigated by using two independently trained and set up PoS taggers, however. But this finally gets us back to my initial concern: How much performance gain really is in there from all this added complexity and is that "worth the effort"?
- williamtrask 11y agoGenerally, the industry benchmarks dependency parsing using gold standard POS tags. However, your point is well taken. Personally, I have little doubt that it would still yield the same level of improvement, but fortunately a bit of experimentation can settle it for sure :) Perhaps also relevant to this conversation, the disambiguation for pre-training did in fact use "real-world" tags (not gold standard). Thus, sense2vec as an algorithm was able to sort through the noise generated by mistakes in the part-of-speech tagger to still generate meaningful embeddings.