4 ms·
It wasn't obvious that a transformer could do this, and learn to produce conv via attention
by montebicyclelo 1y ago
It wasn't obvious that a transformer could do this, and learn to produce conv via attention
- artemisart 1y agoBut it is, as long as the positional embedding are sufficient, i.e. use relative positional embeddings here.