4 ms·
This is a bit like saying "F1 doesn't produce useful cars", isn't it? The point of a competition is to meet specific parameters as well as possible and push th
by Sileni 7y ago
This is a bit like saying "F1 doesn't produce useful cars", isn't it?
The point of a competition is to meet specific parameters as well as possible and push the boundaries of what can be done. It's not meant to create a "daily driver".
I realize he argued against this with the coin flip test, but that is why you'd ideally want to have many of these competitions over time. If you start to see the same names popping up at the top regularly, you know there's some sort of significance to them. Teams would ultimately want to trend towards whatever wins competitions most consistently, so they'd want to rely on models they think are the most likely to perform in a real world test. They wouldn't want to simply rely on a coin flip.
And in large competitions, you have the chance of batching the top performers together and seeing what is common between them. Presumably there's a reason these models pan out over the rest in aggregate; are they worth pursuing a bit?
I think we agree at the end though, even within my analogy. A huge value of F1 racing is the publicity the teams give their sponsors. They might learn some information that can be pushed down to their consumer vehicles, but it's marginal compared to team winnings and the value of saying "See? Our engineers are the best".
- xenocyon 7y agoThe article isn't actually that trivial; it makes a very provocative and bold statement towards the end, essentially declaring that all stated progress in image classification in the last 5 years is questionable. I don't know if I necessarily agree with that, but there is definitely a danger in evaluating models purely based on their predictive power when many models are being evaluated on a common dataset - which is exactly the mainstream practice in the world of deep learning research - and it is therefore wise to be wary, not just as individuals but as a community.
- lukeor 7y agoHi, author here. I didn't actually mean to suggest that the last 5 years of performance improvement could be spurious. That clearly isn't true. I use resnets/densenets etc in my day to day work! What the picture was trying to say is that, within a given year, the "winner" becomes less likely to be truly better than the second place team. Alexnet was clearly better than the alternative, even with Bonferroni adjusted significance thresholds. Less so by 2016/17. I'm writing a follow up on imagenet in particular to address some of the nuance. It is very clearly not a representative example of ML competitions, but the same effects still apply to some extent (imo).
- scoopertrooper 7y ago> Teams would ultimately want to trend towards whatever wins competitions most consistently Could this just surface the teams that just submitted the most models? Maybe some sort of wins per a submission score could help with this?
- nabla9 7y agoWe can't have meaningful discussions based on titles or misunderstanding of the context the article uses and overgeneralizing. My understanding is that the context is "usable models in clinical setting". Am I reading it wrong?