4 ms·
There was no hugely scalable approach before transformers, RNNs, the previous SOTA, were notoriously bad at scaling.
by bitL 3y ago
There was no hugely scalable approach before transformers, RNNs, the previous SOTA, were notoriously bad at scaling.
- tensor 3y agoYes, we needed clever ideas from scientists to make them scale. In fact, we still need clever ideas to make them scale because the current architectures still have all sorts of problems with length and efficiency.