3 ms·
I'm not sure what to make of this post. There is always a degree of uncertainty with the experimental design and it's not surprising that there are a couple of
by muds 3y ago
I'm not sure what to make of this post. There is always a degree of uncertainty with the experimental design and it's not surprising that there are a couple of buggy questions. Imagenet (one of the most famous CV datases) at this point is known to have many such buggy answers. What is surprising is the hearsay that plays out on social media that blows the proportion of the results out of the water and leads to opinion pieces like these targeting the authors instead.
Most of the damning claims in the conclusion section (Obligatory: I haven't read the paper entirely, just skimmed it.) usually get ironed out in the final deadline run by the advisors anyway. I'm assuming this is a draft paper for the EMNLP deadline this coming Friday published on arxiv. So this paper hasn't even gone through the peer review process yet.
- iudqnolq 3y agoImageNet has five orders of magnitude more answers, which I would assume makes QA a completely different category of problem. The authors could probably have carefully review all ~300 of their questions. If they couldn't they could have just reduced their sample size to say 50.
- muds 3y agoI admit that Imagenet isn't the best analogy here. But I'm pretty confident that this data cleaning issue would be caught in peer review. The biggest issue which I still don't understand was the removal of the test set. That was bad practice on the authors' part.
- iudqnolq 3y agoIn general with evaluations LLVMs I keep seeing issues that would be caught by a human carefully looking at the results. To a nonexpert who occasionally peeks at things it seems like having a small dataset that's just slightly too big for you to manually review is a bad practice. It also seems like 100% accuracy should have raised red flags, especially if you know your dataset isn't perfectly cleaned.