6 ms·
Retentive Network: A Successor to Transformer for Large Language Models
- Anya200 3y ago[flagged]
- deleted 3y ago[deleted]
- lumost 3y agoThis a super interesting paper, but in my oppinion - they do not complete their core claim of offering a generalized mathematical structure for attention/recurrence. The specific structure they propose is very interesting and demonstrably efficient computationally - however they do not show that this approach produces similar accuracy as large LLMs. I’m anxiously awaiting the follow up where someone tries spending 1MM+ on demonstrating this approaches effectiveness in a large language model context.
- whimsicalism 3y ago> however they do not show that this approach produces similar accuracy as large LLMs. I think they have demonstrated their case pretty well, unless there is some serious degradation of the scaling - 7b is pretty big.
- turingfeel 3y agoInterestingly, I did see this tweet [0] mentioning a phase shift that occurs in transformers at exactly the scale RetNet stopped at. Probably simply coincidental but I was previously unaware of this phenomenon at such a scale in transformers. [0] https://twitter.com/gordic_aleksa/status/1682479676910870529 https://twitter.com/gordic_aleksa/status/1682479676910870529
- whimsicalism 3y agotim dettmers is such a resource, cheers for this
- wills_forward 3y agoMIT and Microsoft Researchers Introduce "RetNet" - An 8X Faster Transformer Alternative for AI
- mirekrusin 3y agoThe claim is parallelism for training which is not fixed speed up, different complexity for inference (constant time), and different complexity for large context inference (linear) - so nothing that can be summarised as 8x - or am I getting this summary wrong?
- whimsicalism 3y agothe words per second i believe from the first graph in the paper
- canjobear 3y agoWhat’s the MIT connection? The authors all seem to be affiliated with Tsinghua.
- whimsicalism 3y agoAmazing - 6.7 billion is significantly larger than I've seen any transformer alternative trained to so far (H3 only went up to 2.7b, e: oops - RWKV goes up to 14b)... cool to see that it appears to be scaling even better and the O(1) & O(N) scaling is great. Wish there was more consistency on source of training data. Training on just The Pile would enable more clean comparison with most promising transformer alternatives, like H3 and give a better sense of how robust the perplexity improvements cited are.
- bhy 3y agoRWKV has 14B version.
- whimsicalism 3y agogood point, for some reason i always leave out rwkv when thinking of the transformer models.. perhaps because it is more of a redux
- ttul 3y agoThe key differences between multi-head attention in Transformers and the proposed multi-scale retention (MSR) in RetNets are: - Retention replaces the softmax in attention with an exponential decay along the sequence dimension. This allows formulating retention in a recurrent form for efficient O(1) inference. - Retention heads use different decay rates (gamma values) for multi-scale modeling. Attention heads use the same softmax. - Retention outputs are normalized per-head with GroupNorm before concatenation. Attention uses LayerNorm on the concatenated output. - Retention can be computed in parallel, recurrent, or chunkwise recurrent modes. Attention is only parallel. - The recurrent form enables RetNets to summarize long previous context into a fixed-size state during inference. Attention recomputes on the full context each step. - So in summary, retention adapts attention to enable recurrent modeling and multi-scale decays. This provides efficiency benefits and competitive performance to Transformers.
- sp332 3y agoWith transformer models, it’s common to put instructions and system messages at the beginning of the input. But with this decay, the beginning of the input would always have the sparsest attention, right? Maybe the instructions should be moved to the end. But then again if it’s recurrent, you might want to prime it with a description of the task.
- ttul 3y agoHypothetically, fine tuned transformer networks would learn to attend more to the beginning and end of the sequence since that is where the instructions typically are. And lo and behold a recent paper demonstrated this to be true.
- sfriedr 3y agoI only spent a few minutes skimming thr paper, but: 1) there are a lot of papers claiming to be the successor to the Transformer, and not all of them are cited; e.g., the MetaFormer is missing https://arxiv.org/abs/2111.11418 https://arxiv.org/abs/2111.11418. Another candidate that wasn't compares against (or at least argued why it wouldn't make sense to compare against) are the Hopfield Networks https://arxiv.org/abs/2008.02217 https://arxiv.org/abs/2008.02217. So until a more solid Related Work section is written (their section is actually called "Relation to and Differences from Previous Methods") I reserve the right to be skeptical whether their model is the "best" successor to the Transformer. 2) they say in the abstract "We theoretically derive the connection between recurrence and attention" but I couldn't find a longer theorem-proof section. So either this is done only in a cursory manner, or the proof is very easy. Recurrence and attention have been around for a long time as concepts, so surely there are already proofs in similar contexts of this fact (I am not working in this particular area of Machine Learning, so I don't know the SOTA by heart, but I strongly suspect that these aspects have been discusses previously; thr Hopfield Network paper I linked to unearthes some theoretical facts about attention, for example). So -based on my very cursory reading- this paper seems like an interesting approach, but I do see some holes in thr execution. Time will tell whether Rentetive Network will become mainstream or not. Ok, this was my five minute review of the paper. Now I have to urgently return to completing my actual reviews for NeurIPS, haha.
- whimsicalism 3y ago> 1) there are a lot of papers claiming to be the successor to the Transformer, and not all of them are cited; e.g., the MetaFormer is missing https://arxiv.org/abs/2111.11418 https://arxiv.org/abs/2111.11418. Another candidate that wasn't compares against (or at least argued why it wouldn't make sense to compare against) are the Hopfield Networks https://arxiv.org/abs/2008.02217 https://arxiv.org/abs/2008.02217. Neither of those papers are NLP applicable? And I think it's perfectly fair to focus on the alternatives (ie. like H3 and RWKV) that have been able to scale up to LLM levels and perplexity, which neither of the alternatives you mention have. Should they just cite every 'is All You Need' paper?
- 3y ago