4 ms·
This paper “were RNNs all we needed?” explores this hypothesis a bit, finding that some pre-transformer sequence models can match transformers when trained at a
by rsfern 1y ago
This paper “were RNNs all we needed?” explores this hypothesis a bit, finding that some pre-transformer sequence models can match transformers when trained at appropriate scale. Though they did have to make some modifications to unlock more parallelism
https://arxiv.org/abs/2410.01201 https://arxiv.org/abs/2410.01201