3 ms·
> With a bit of experimentation on Imagenet-1k, we can reach 82.0% accuracy with a 176x176 training image size with no extra data, matching ConvNeXt-T (v1, with
by jerpint 3y ago
> With a bit of experimentation on Imagenet-1k, we can reach 82.0% accuracy with a 176x176 training image size with no extra data, matching ConvNeXt-T (v1, without pre-training a-la MAE) and surpassing ViT-S (specifically, the ViT flavor from DeiT-III).
This is of no surprise, and well known. ViTs excel when data is plentiful, but perform less well than convolutions in lower data regimes. ConvNets on the other hand (resnets for example) don’t do as well in high data regimes. Imagenet is considered “low data regime” compared to what ViTs get normally trained on, most SOTA ViTs on imagenet are typically pre-trained on MUCH larger datasets (e.g. JWT300)
- huevosabio 3y agoMy intuition for this is that CNNs impose a stronger prior on the function space. We rely on the previous experience of computer vision and even our own retinas to give it a useful prior. Transformers are more general, they impose a weaker prior and require more data. But by imposing a weaker prior they also are more flexible and adaptable, thus able to take advantage of more data.
- chrishare 3y agoAgree. However I feel that there may still be many priors that are useful in practice, like decomposing the vision task into a hierarchy of sub problems, but as a human - I could be biased.
- huevosabio 3y agoOh absolutely, I think having a prior makes these problems tractable. Otherwise we would just force feed data to fully connected networks!
- Legend2440 3y agoIt would be much better if we had a way to learn good priors -> extract them -> initialize new models with them. We shouldn’t be trying to guess or handcraft them.
- joefourier 3y agoNot all vision transformers have weak priors. Shifted-window transformers and neighborhood attention have priors well suited to images; the latter is showing extremely strong performance in image generation (such as the recent hourglass diffusion which allows pixel space diffusion to be trained orders of magnitudes faster than vanilla attention) and general image classification, and certainly does not need the same dataset size of classical ViTs.
- dkislyuk 3y agoYes, exactly. ViTs need O(100M)-O(1B) images to overcome the lack of spatial priors. In that regime and beyond, they begin to generalize better than ConvNets. Unfortunately, ImageNet is not a useful benchmark for a while now since pre-training is so important for production visual foundation models.
- fzliu 3y agoI recommend checking out this paper: https://arxiv.org/pdf/2310.16764.pdf https://arxiv.org/pdf/2310.16764.pdf. From the concluding section: "Although the success of ViTs in computer vision is extremely impressive, in our view there is no strong evidence to suggest that pre-trained ViTs outperform pre-trained ConvNets when evaluated fairly." Neural networks can be unpredictable, and there's evidence that questions how important transformers' lack of inductive bias (at scale) really is.