4 ms·
I don't get the point. A simple CNN with stride =1 should be able to solve it perfectly and generalize it to any size.
by Nopoint2 1y ago
I don't get the point. A simple CNN with stride =1 should be able to solve it perfectly and generalize it to any size.
- montebicyclelo 1y agoIt wasn't obvious that a transformer could do this, and learn to produce conv via attention
- artemisart 1y agoBut it is, as long as the positional embedding are sufficient, i.e. use relative positional embeddings here.