3 ms·
This is nice because convolutional models seem better for some vision tasks like segmentation which are less obvious how to do with ViTs. Convolution seems like
by dontreact 3y ago
This is nice because convolutional models seem better for some vision tasks like segmentation which are less obvious how to do with ViTs. Convolution seems like something you fundamentally want to do in order to model translation invariance in vision.
- famouswaffles 3y agoSegment Anything is a transformer https://segment-anything.com/ https://segment-anything.com/
- lostmsu 3y agoBut what about rotational invariance? I feel like tilts are about as common as slides.