4 ms·
I appreciate the authors calling out #2. ImageNet and CIFAR are both "solved" benchmarks, to the extent that state of the art algorithms by now are likely overf
by dkislyuk 8y ago
I appreciate the authors calling out #2. ImageNet and CIFAR are both "solved" benchmarks, to the extent that state of the art algorithms by now are likely overfitting to the specific dataset details. In particular, objects in ImageNet are nearly always in canonical pose, with no occlusion, few confounding objects, illuminated, and in a semantically-obvious configuration. For some applications this is acceptable but as an industry benchmark ImageNet is not informative (user generated photo distributions are never this clean).
Even more damning is the recent BagNet paper (nice summary here: https://blog.evjang.com/2019/02/bagnet.html https://blog.evjang.com/2019/02/bagnet.html), which indicates that ImageNet can likely be solved with no global features (i.e. model doesn't have to learn anything truly abstract, just configurations of textures, shapes, colors). I thought the author of that blog post put it nicely:
"As someone who is deeply interested in AGI, I find ImageNet much less interesting now, precisely because it can be solved with models that have little global understanding of images."
- joshvm 8y agoThis also happened (and is happening?) in stereo imaging research. For years, Middlebury was what you tested on, and for years that's what got you published. Nowadays Middlebury is viewed as solved by the top algorithms. If you try those algorithms on your own data, good luck getting similar performance; at least I've not seen any kind of advantage in using anything other than SGM (outside of specific research contexts like my PhD). I'm more concerned that everyone is using KITTI as a (often the only) benchmark for deep-learning based stereo matchers, since those are all images of roads. At least with classical stereo you have some idea what* your cost function is. The other one people are increasingly using is Scene Flow, which is (entirely?) synthetic. Not a great situation. * KITTI is a widely used dataset of driving imagery