3 ms·
Masked Autoencoders Are Scalable Vision Learners – New SSL Algorithm
- conview 5y agoThe return of patch-based self-supervision! No ResNets but from one of the authors of the initial paper. Now with ViT, very simple self-supervised (SSL) pre-training shines again. It outperforms the contrastive learning counterparts and is simpler than BEiT, as its pixel-based (no tokenisation) needed. No need for special augmentation considerations like BYOL. i1k=87.8%