4 ms·
If I understand what you're asking, the Transformer isn't initially treating the image as a sequence of pixels like p1, p2, ..., pN. Instead, you can use a conv
by mliker 3y ago
If I understand what you're asking, the Transformer isn't initially treating the image as a sequence of pixels like p1, p2, ..., pN. Instead, you can use a convolutional neural network to respect the structure of the image to extract features. Then you use the attention mechanism to pay attention to parts of the image that aren't necessarily close together but that when viewed together, contribute to the classification of an object within the image.
- famouswaffles 3y agoVision Transformers don't use CNNs to extract anything first. (https://arxiv.org/abs/2010.11929 https://arxiv.org/abs/2010.11929). You could but it's not necessary and it doesn't add anything so it doesn't happen anymore. Vision transformers won't treat the image as a sequence of pixels but that's mostly because doing that gets very expensive very fast. The image is split into patches and the patches have positional embeddings.