4 ms·
Here's the thing: the authors, and most people in the NLP community (as opposed to people who can use off-the-shelf tools and nothing more) know how to get bett
by sqrt17 8y ago
Here's the thing: the authors, and most people in the NLP community (as opposed to people who can use off-the-shelf tools and nothing more) know how to get better performance on other domains. It's just that "it requires some adjustments and manual labour" doesn't add anything to the "deep learning solves the problem" narrative that is predominant in research articles and most blog posts on the topic.
On the other hand, you can bet that actual practice in Zalando (the authors are all from Zalando's reasearch lab) involves more regexes and retraining models on proprietary datasets and less using off-the-shelf models and hope they stick.
No one claims either that you can solve every Vision problem with a model trained on ImageNet - you'd do transfer learning, or for non-understanding problems (estimating colors and contrast or anything else that's unrelated to objects in the image) you'd use something else that doesn't involve deep learning models at all.
- sweezyjeezy 8y agoRight, but my point is - once you start needing to add ad-hoc retraining, or regex hacks, it's not clear to me that shaving a point off baseline f1 scores is really all that relevant anymore.
- yorwba 8y agoYou'd have to do the same modifications to the baseline models to adapt them to a different domain. If they managed to shave off percentage points on a large number of benchmarks, then it's likely that using their models will also help you with the task you care about.
- sweezyjeezy 8y agoNot convinced. Pretty much all baseline NER datasets are on news corpuses, which are written in well formatted prose, tend not to have spelling mistakes, abbreviations, bad punctuation, etc. etc. Why do you think that a better performance on these kinds of datasets will translate to better performance in other domains? I wouldn't even be surprised if it's the opposite, maybe it relies more heavily on these assumptions. The truth is there is no way to know a priori - you need a different kind of benchmark to test this.
- yorwba 8y agoWNUT-17 [1] is not a news corpus. It has a lot of badly-formatted prose, spelling mistakes, abbreviations, bad punctuation etc. etc. Accordingly, it's the dataset where they get the worst F1 of 50.20, but still better than the previous best of 45.55. In general, I'd be surprised if they hard-coded reliance on the specific regularities of news texts into the model assumption, so if a model is able to exploit those regularities better, training it on a corpus with different regularities should also enable it to perform well on that corpus. [1] https://noisy-text.github.io/2017/emerging-rare-entities.html https://noisy-text.github.io/2017/emerging-rare-entities.htm...
- pdyck 8y agoI had the same experience when trying to do NER on customer support requests. My model performed great for research datasets but it was mediocre at best for my own dataset. Do you have any suggestions on how to achieve better results in domains where mistakes, bad punctuation, etc are common?
- edraferi 8y agoLabel more training data. Do more clustering. Label more training data. Strip out more garbage. Label more training data. PS you can get an idea of how much value additional training data will give you by training models on various subsets of your dataset (e.g. 10%, 20%...), evaluating them against the same test dataset, and plotting the results.
- rpedela 8y agoI have found pre-trained models for language detection and Wikipedia word embeddings useful. Everything else I have to train from scratch. I have successfully done NER on search keywords where there were many examples of misspelled words or weird punctuation. I used Spacy but I had to train the model from scratch.
- srean 8y agoI dont think there is anything contentious here. If the data that a model is trained on is nothing like what the model will be applied on the results will suck. Well, duh! isn't that obvious ? If one is trained on screwing light bulbs that training would not be very helpful in composing music. If there is some common structure between the train and the test scenarios there would be some point in learning that from the default train set. Then you use your domain specific training set to unlearn the things that do not apply and learn the other things that do. As long as there is something worth learning from the default train set it will be of some use.