4 ms·
> ViT models are outperforming CNNs in terms of computational efficiency and accuracy, achieving highly competitive performance in tasks like image classificati
by mrintellectual 4y ago
> ViT models are outperforming CNNs in terms of computational efficiency and accuracy, achieving highly competitive performance in tasks like image classification, object detection, and semantic image segmentation.
Since then, this has been show to be untrue. Using more modern training techniques along with depthwise convolutions (https://arxiv.org/abs/2201.03545 https://arxiv.org/abs/2201.03545) results in equal if not better performance on vision tasks. Improved training methodologies have also been shown to boost the accuracy of ResNet50 - an 6-year-old pure convolutional architecture - on ImageNet-1k by over 5% (https://arxiv.org/abs/2110.00476 https://arxiv.org/abs/2110.00476).
Pure ViTs are also more difficult to train when compared with traditional convnets, although this has since then been somewhat remedied by Swin (https://arxiv.org/abs/2103.14030 https://arxiv.org/abs/2103.14030).