6 ms·
There’s some interesting work replacing scaled dot product attention and position embeddings with fixed format MLPs [0] - so I tend to lean towards thinking of
by rsfern 2y ago
There’s some interesting work replacing scaled dot product attention and position embeddings with fixed format MLPs [0] - so I tend to lean towards thinking of classic transformers as having a reasonable enough inductive bias and the scalability to actually realize the amount of compute that’s needed
0: https://arxiv.org/abs/2105.08050 https://arxiv.org/abs/2105.08050