3 ms·
xLSTM code release by NX-AI
- pietz 2y agoCould someone provide a quick summary where they stand compared to transformer architectures? Do they have real world scale results that are competitive?
- barrell 2y ago- They outperform transformers at lower parameter counts. Time will tell if that hold up for more parameters - They scale linearly in complexity, which means with a longer context window they will be faster and cheaper than transformers - It's been mostly academic as far as I know, only just recently being published. I don't think there's been an opportunity to use them at 'real world scale' yet, although tbh I'm a little uncertain what you mean by it.
- oersted 2y agoIt does seem to match transformers, but I wouldn't say it meaningfully outperforms them in terms of quality vs parameters. Model: #Params (M), SlimPajama (15B) ppl ↓ - GPT-3: 356M, 14.26 - Llama: 407M, 14.25 - H3: 420M, 18.23 - Mamba: 423M, 13.70 - Hyena: 435M, 17.59 - RWKV-4: 430M, 15.62 - RWKV-5: 456M, 16.53 - RWKV-6: 442M, 17.40 - RetNet: 431M, 16.23 - HGRN: 411M, 21.83 - GLA: 412M, 19.56 - HGRN2: 411M, 16.77 - xLSTM[1:0]: 409M, 13.43 - xLSTM[7:1]: 408M, 13.48 There are more detailed perplexity and task benchmarks in the paper. Overall, all the architectures perform very similarly on every benchmark, sometimes xLSTM is slightly ahead but not always, and the difference is not really meaningful. This is great news though, it means we are not losing anything by switching to xLSTM and we get important advantages like the scalable context window. I'm quite excited about this because we can potentially have the LLM remember what you say and do few-shot persistent learning from user interaction (updating "itself", the state vector). It would be very interesting if LLMs were no longer static. Although I'm sure it will be a challenge to train the model to keep such learnings in its memory long-term. The paper: https://arxiv.org/abs/2405.04517 https://arxiv.org/abs/2405.04517
- 3abiton 2y agoLinear scaling for context is also a bit deal. Flash attention partially solved this for TF, but xLSTM seems promising!
- piecerough 2y ago> It would be very interesting if LLMs were no longer static. Little bit of a nightmare too. Instructions keep piling up for you that you no longer openly can access and remove
- wantsanagent 2y agoDeeper dive by Yannic if you want it: https://www.youtube.com/watch?v=0OaEv1a5jUM https://www.youtube.com/watch?v=0OaEv1a5jUM
- htrp 2y agoThis is exciting because it is an architecture that had so much promise, but we could never solve the gradient/parallelization problems better than transformers. This code will allow people yo experiment and see if it is a viable architecture at foundation/frontier model scale.
- ein0p 2y agoNote: GNU AGPLv3. Industry labs won’t touch this with a hundred foot pole. Given that they’re the only ones with access to serious resources, it could be a while before we see a large model of this architecture
- striking 2y agoReimplementation from paper is pretty common, though, no?
- ein0p 2y agoYes, that’s why it’ll take time. There’s so much stuff competing for researchers’ attention, and experimentation with this takes so much time and $$$, that if it wasn’t for Sepp Hochreiter on the list of authors this could get ignored entirely. IOW it’s not the seller’s market for novel architectures right now.
- htrp 2y agoYou can't outspend the industry labs given the compute inflation in transformer architectures (unless you are ridiculously well connected in the venture/sovereign funding communities). And realistically, do we need another GPT4 evaluation paper?
- ein0p 2y agoThat is by far not the only thing industry labs are working on currently. I work in one. My group might be unusual, but I can’t name a single currently active project here that is not a departure from Transformers one way or another. I expect a ton of such efficiency oriented work in the next 4-5 years. We can’t be burning money as inefficiently as we do it right now.
- nurple 2y agoHow does AGPLv3 impact a lab's ability to do research on an implementation?
- trextrex 2y agoI'm not clear on what advantage this architecture has over mamba/Griffin. They also have the linear scaling, better sequence parallelism and are competitive in performance with transformers.
- wave_1 2y agostate tracking...
- lalaland1125 2y agoThe whole field seems to be having issues with comparisons right now. We really don't even know how Mamba vs Griffin compare.
- ganzuul 2y agoAre there any studies on predicting neural architecture scaling? E.g. a small training dataset which indicates performance on a large training dataset?
- dang 2y agoRecent and related: xLSTM: Extended Long Short-Term Memory - https://news.ycombinator.com/item?id=40294650 https://news.ycombinator.com/item?id=40294650 - May 2024 (73 comments)
- brcmthrowaway 2y agoCongrats to the x.AI team!
- tripplyons 2y agoThis release was not from xAI, it was from NXAI.