3 ms·
That's something that Microsoft research wrote two decades ago. And those results were well known in the NLP community. Example: https://www.microsoft.com/en-u
by mrazomor 3y ago
That's something that Microsoft research wrote two decades ago. And those results were well known in the NLP community.
Example: https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/sigir2002-qarevised.pdf https://www.microsoft.com/en-us/research/wp-content/uploads/...
(Michele Banko published a few similar papers on that topic)
- bitL 3y agoThere was no hugely scalable approach before transformers, RNNs, the previous SOTA, were notoriously bad at scaling.
- tensor 3y agoYes, we needed clever ideas from scientists to make them scale. In fact, we still need clever ideas to make them scale because the current architectures still have all sorts of problems with length and efficiency.