4 ms·
It wouldn't, because final scores are only evaluated against the "validation" set. As for turning current contests into an exercise in overfitting the "test" s
by solve 11y ago
It wouldn't, because final scores are only evaluated against the "validation" set.
As for turning current contests into an exercise in overfitting the "test" set - we already reached that point long ago. Test vs validation scores often diverge wildly in these contests.
Edit - Replying to arnsholt:
Completely true. The huge problem I see, is that all the classic NLP tagging corpuses are created from the very narrow domain of news articles, and a few good corpuses now appearing for biology texts, and that's about it. Want to do, e.g. NER for product reviews or chat logs? - Incredibly bad results. There's a huge corpus problem in NLP today.
- arnsholt 11y agoNot to mention out-of-domain performance. I'm not familiar with computer vision, but in NLP taggers are hovering around human-level performance, and parsers are quickly approaching that level. But if you take a state-of-the-art system and test it on a slightly different corpus (even something as simple as text from the same newspaper, but a year later!) performance drops by a lot.