3 ms·
Next, try the taggers on a more realistic setting than the standard corpuses -- e.g. a product review that compares several products, and you'll instantly see h
by nxb 11y ago
Next, try the taggers on a more realistic setting than the standard corpuses -- e.g. a product review that compares several products, and you'll instantly see how incredibly poor the current state of the art NER is.
Technology is really going to advance once we have anything that comes close to human level on NER and relation extraction. Kind of like self driving cars, the basic ideas have been around for decades, but performance in realistic adverse conditions remains awful for almost everywhere that it could theoretically be used.
- boomzilla 11y agoThat is because the taggers are not trained on the same data. You can't expect taggers trained on wikipedia data to do well in anything but other wikipedia articles. On the other hand, if one has access to Amazon review data, (with links to the product catalog), I am pretty sure a tagger that does well on Amazon data can be trained.
- zeerakw 11y agoWell that depends, if you somehow manage to link well across different domains it can be done. Take a look at the Lowlands project from Copenhagen University (http://lowlands.ku.dk http://lowlands.ku.dk), which deals specifically with cross domain adaptation. You are right that reasonable domains are required though.