3 ms·
"As further proof, features from the model achieve state-of-the-art performance on a number of classification datasets and near state-of-the-art unsupervised ac
by jaredtn 6y ago
"As further proof, features from the model achieve state-of-the-art performance on a number of classification datasets and near state-of-the-art unsupervised accuracy on ImageNet."
Impressive stuff! This performs well even without domain-specific architecture choices.
- visarga 6y agoOne more step for ML. It used to be that we needed hand designed image features. Now we can learn even the image priors (spatial locality and translation invariance) from data. Transformers are basically learning relations between pairs of input tokens, moving the problem to a more abstract level than predicting directly on tokens. While CNNs excel at benefiting from those two forms of invariance, transformers have permutation invariance, they can predict on sets, graphs and non-euclidean spaces.
- gwern 6y ago> Now we can learn even the image priors (spatial locality and translation invariance) from data. Right. The attention layers even learn attention patterns which look like convolution layer kernels! But better, presumably: https://arxiv.org/abs/1911.03584 https://arxiv.org/abs/1911.03584