3 ms·
I would argue that input scaling is not fundamental to Transformers. Recurrent neural network size is also independent of input sequence length. The successfu
by mrfox321 6y ago
I would argue that input scaling is not fundamental to Transformers.
Recurrent neural network size is also independent of input sequence length.
The successful removal of inductive bias is really what differentiates this from previous sequence-to-sequence neural networks.
- keithyjohnson 6y agoWhich inductive bias?
- nmfisher 6y agoPresumably that the output at step (n) is conditioned only the output of step (n-1).