3 ms·
I wonder if the exact some models that failed this test would succeed if their training data included images of weird angles/unusual contexts? My guess is ever
by 2bitencryption 7y ago
I wonder if the exact some models that failed this test would succeed if their training data included images of weird angles/unusual contexts?
My guess is every "hammer" image in the training data set was "conventional" -- a convenient angle and orientation. If half the images of "hammers" were instead "unconventional", would the model adapt to realize "my existing model of a hammer is incomplete; there must be a way to consolidate these two different images"?
Or does this require an internal 3d modeling, and better inputs wouldn't help; instead the model itself would need to be more advanced?
- bonoboTP 7y agoThey show in section 4.3 that fine-tuning the last layer of ResNet-152 on half of ObjectNet (25k images) and testing on the other half increases the top-1 accuracy from 29% to 50%, while the corresponding accuracy on ImageNet is ~67%. Nevertheless, I agree with you. Given a huge dataset with millions of of unconventional images may be enough. Who knows. Things kind of go in cycles in machine learning (similar to other fields). There is nowadays growing dissatisfaction of having to use so much (labeled) data, and people want the models to be better "primed" to capture the variations and structures existing in the real world. Partially because labeling a lot of data is just very expensive, but partially it's also seen as inelegant and "black-boxy" or it's just not in their scientific taste. Other people argue that learning it all from data is fine and this kind of robustness shouldn't have to be baked in to models. Rather they should/could be learned from vast amounts of unlabeled data instead (Yann LeCun seems to be in this group.), with unsupervised/self-supervised methods.