4 ms·
So your results demonstrate that directed dependency labeling works better with vectors learned from PoS-tagged words than with PoS-tagged vectors (learned from
by fnl 11y ago
So your results demonstrate that directed dependency labeling works better with vectors learned from PoS-tagged words than with PoS-tagged vectors (learned from untagged words)? And if so, why are you sure you are not overfitting on the corpus or that the "unseen" (in your case: label +) word (pairs) issues will in the end do more harm than what you gain when using this approach on truly independent data/text?
EDIT: Sorry, this question above is probably too convoluted to understand. As I understand, the evaluation of the UAS in the paper was made by letting the parser use the gold PoS labels from the UD treebank (plus either word embeddings). But what would happen if the PoS labels for evaluating the dependency parser came from a PoS tagger, as would be the case when working on unseen data? I might imagine that "plain" embeddings could maybe produce a better UAS in that case, because they are not as "overfitted" as the "enriched" embeddings (as those are derived from the PoS-tagger labeled words in the first place).
- williamtrask 11y agoPerhaps, but seeing as POS taggers are ~97% accurate (at least in English), I'd expect this to be minimal. Furthermore, the baseline neural network also has access to the gold standard POS tags, so the comparison of adding POS disambiguated embeddings is pretty clean. It's the difference between "words + pos tags" as features and "pos-disambiguated word + pos tags".
- fnl 11y agoUsing accuracy to measure PoS taggers makes results look good, but is ensnaring due to their huge bias: Tagging every word by the majority tag found during training and everything else as either NNP or NNPS (with suffix -s) means the statistical baseline is already well beyond 90% accuracy. However, my point was that from the results shown its not clear to me if the gains in attachment score you saw when using Gold Standard PoS tags would be lost in a "real-world" usage when you have to rely on the tagger's own PoS tags. In such a case, it could be that your embeddings contribute much less "new knowledge" than what you see in your results, using independent (Gold) PoS tags. This might be mitigated by using two independently trained and set up PoS taggers, however. But this finally gets us back to my initial concern: How much performance gain really is in there from all this added complexity and is that "worth the effort"?
- williamtrask 11y agoGenerally, the industry benchmarks dependency parsing using gold standard POS tags. However, your point is well taken. Personally, I have little doubt that it would still yield the same level of improvement, but fortunately a bit of experimentation can settle it for sure :) Perhaps also relevant to this conversation, the disambiguation for pre-training did in fact use "real-world" tags (not gold standard). Thus, sense2vec as an algorithm was able to sort through the noise generated by mistakes in the part-of-speech tagger to still generate meaningful embeddings.
- williamtrask 11y agoRe-read this and perhaps identified a mis-understanding. Above you mentioned that "directed dependency labeling works better with vectors learned from PoS-tagged words than with PoS-tagged vectors (learned from untagged words)". The method does not pos-label vectors or words in the model. Instead, it pos-labels the text that sense2vec is trained on. In this way, you get multiple vectors for words with multiple-POS usages (as if they were different words altogether). The change to the syntactic parser was just based on using POS to select which word embedding (of the several available for each word) to use as input. Sense2vec was trained on predicted POS labels. The parser used gold standard POS tags in both normal word2vec and sense2vec for even comparison.