4 ms·
They often reference this paper as the motivation for that https://arxiv.org/pdf/2203.15556.pdf https://arxiv.org/pdf/2203.15556.pdf I.e. training with 10x data
by naillo 3y ago
They often reference this paper as the motivation for that https://arxiv.org/pdf/2203.15556.pdf https://arxiv.org/pdf/2203.15556.pdf I.e. training with 10x data and 10x longer can yield as good models as a gpt-3 model but with fewer weights (according to the paper) and the same principle applies in vision.