2 ms·
Not all vision transformers have weak priors. Shifted-window transformers and neighborhood attention have priors well suited to images; the latter is showing ex
by joefourier 3y ago
Not all vision transformers have weak priors. Shifted-window transformers and neighborhood attention have priors well suited to images; the latter is showing extremely strong performance in image generation (such as the recent hourglass diffusion which allows pixel space diffusion to be trained orders of magnitudes faster than vanilla attention) and general image classification, and certainly does not need the same dataset size of classical ViTs.