9 ms·
xLSTM: Extended Long Short-Term Memory
- KhoomeiK 2y agoFor those who don't know, the senior author on this paper (Sepp Hochreiter) was the first author on the original paper with Schmidhuber introducing LSTMs in 1997.
- deleted 2y ago[deleted]
- ramraj07 2y agoAt least in biology, the first author of a paper is more often than not just a pair of gifted hands who did the experiments and plotted the graphs. Doesn’t always translate that they become good PIs later (though they get their chances from these papers).
- albertzeyer 2y agoIt seems Sepp Hochreiter has talked already about this model since Oct 2023: https://github.com/huggingface/transformers/issues/27011 https://github.com/huggingface/transformers/issues/27011 In the scaling law comparison, I wonder if it is reasonable to compare number of parameters between Llama, Mamba, RWKV, xLSTM? Isn't compute time more relevant? E.g. in the figure about scaling laws, replace num of params by compute time. Specifically, the sLSTM has still recurrence (memory mixing) in it, i.e. you cannot fully parallelize the computation. So scaling up Transformer could still look better when you look at compute time. It seems neither the code nor the model params are released. I wonder if that will follow.
- YetAnotherNick 2y agoRecurrence is less of issue with really large models training than it is with medium sized models. Medium sized transformer models are generally not trained with sequence parallelism, but sequence parallelism is getting more common with transformer training. And sequence parallelism is same for transformer or recurrent model. For really large models, it is in fact easier to achieve peak flops because computation required scales faster than memory bandwidth required(square vs cube).
- albertzeyer 2y agoWith sequence parallelism, you mean to increase the batch size, i.e. number of sequences in a batch? > Medium sized transformer models are generally not trained with sequence parallelism, but sequence parallelism is getting more common with transformer training Is there some word missing? You mean it's more common for large-sized Transformers? > computation required scales faster than memory bandwidth required (square vs cube) That is an interesting thought. I'm trying to understand what exactly you mean. You mean, computation time is in O(N^2) where N is the sequence length, while required memory bandwidth is in O(N^3)? Why is that?
- YetAnotherNick 2y agoNo, it means dividing the sequence into multiple chunks and processing them one by one, very similar to recurrence. See [1]. Sequence parallelism is needed when the sequence can't fit in a single GPU. Sequence parallelism is the hardest parallelism, but it is required for longer sequence. Many models just trains for smaller sequence length for majority of the training and switch to sequence parallelism for last few percentage of training. [1]: https://arxiv.org/pdf/2105.13120 https://arxiv.org/pdf/2105.13120
- logicchains 2y ago>Sequence parallelism is the hardest parallelism, but it is required for longer sequence In terms of difficulty of implementation it's arguably much easier than pipeline parallelism, which I'd argue is the hardest kind (at least to implement it efficiently without bubbles), and takes the most lines of code to implement (especially in Jax, where sequence parallelism is almost trivial).
- sigmoid10 2y agoAnother week, another paper that thinks they can revive recurrent networks. Although this time the father of LSTM is a co-author, so this paper should not come as a surprise. Sadly, the results seem to indicate that even by employing literally all tricks of the trade, their architecture can't beat the throughput of flash-attention (not by a long shot, but that is not surprising for recurrent designs) and, on top of that, it is even slower than Mamba, which offers similar accuracy at lower cost. So my money is on this being another DOA architecture, like all the others we've seen this year already.
- l33tman 2y agoTo put another perspective on this, lots of modern advancements in both ML/AI and especially computer graphics has come from ideas already from the 70-80s that were published, forgotten, and revived. Because underlying dependencies change, like the profile of the HW of the day. So just let the ideas flow, not every paper has to have an immediate payoff.
- KeplerBoy 2y agoTo be fair, Hochreiter seems pretty confident that this will be a success. He stated in interviews "Wir werden das blöde GPT einfach wegkicken" (roughly: We will simply kick silly GPT off the pitch) and he just founded a company to secure funding. Interesting times. Someone gathered most of the available information here: https://github.com/AI-Guru/xlstm-resources https://github.com/AI-Guru/xlstm-resources
- imjonse 2y agoWith all due respect for his academic accomplishments, confidence in this domain in the current climate is usually a signal towards potential investors; it can be backed by anything between solid work (as I hope this turns out to be) and a flashy slide deck combined with a questionable character.
- KeplerBoy 2y ago
- deleted 2y ago[deleted]
- WithinReason 2y agoI like the color coded equations, I wish they would become a thing. We have syntax highlighting for programming languages, it's time we have it for math too.
- imjonse 2y agoMath notation has different fonts, with similar goals as syntax highlighting. It also works well in black and white :)
- aeonik 2y agoObligatory link to BetterExplained color coded math equations. https://betterexplained.com/articles/colorized-math-equations/ https://betterexplained.com/articles/colorized-math-equation...
- elygre 2y agoI have no idea about what this is, so going off topic: The name XLSTM reminds me of the time in the late eighties when my university professor got accepted to hold a presentation on WOM: write-only memory.
- woadwarrior01 2y agoI think it's a fine name. The prefix ensures that people don't confuse it with vanilla LSTMs. Also, I'm fairly certain that they must've considered LSTM++ and LSTM-XL.
- pquki4 2y agoI mean, if you look it another way, XSLT is a real thing that gets used a lot, so I don't mind appending an M there.
- beAbU 2y agoI thought this was some extension or enhancement to XSLT.
- cylemons 2y agoSame
- GistNoesis 2y agoCan someone explain the economics behind this ? The claim is something than will replace the transformer, a technology powering a good chunk of AI companies. The paper's authors seems to be either from a public university, or Sepp Hochreiter's private company or labs nx-ai.com https://www.nx-ai.com/en/xlstm https://www.nx-ai.com/en/xlstm Where is the code ? What is the license ? How are they earning money ? Why publish their secret recipe ? Will they not be replicated ? How will the rewards be commensurate with the value their algorithm bring ? Who will get money from this new technology ?
- imjonse 2y agoShould all arxiv papers be backed by economic considerations or business plans?
- jampekka 2y agoOr any?
- AIsore 2y agoNope, they should not. It is academia after all. How would you even do that in, say, pure mathematics? Concretely, I would love to know what the business plan/economic consideration of Gower's 1998 proof of Szemeredi's theorem using higher order Fourier analysis would even look like.
- queuebert 2y agoComing soon to an HFT firm near you ...
- Der_Einzige 2y agoYes they should. Academia and peer review is so corrupt, gamified, and poor quality that I’d literally trust capitalist parasites more than the current regime of “publish or perish” and citation cartels. At least capitalists have something to fight over that’s worth fighting for (money). Academics will bitterly fight over the dumbest, least important shit. There’s a law about how the less something matters, the more political the fights over it will be.
- smusamashah 2y agoCan someone ELI5 this? Reading comments it sounds like it's going to replace transformers which LLMs are based on? Is it something exponentially better than current tech on scale?
- probably_wrong 2y agoLSTMs are a recurrent architecture for neural networks, meaning that your output depends both on your current input and your previous output. This is similar to how language works, as the next word in your sentence must fit both the idea you're trying to convey (your input) and the words you've said up until now (your previous output). LSTMs where very popular for a while (I think the first good version of Google Translate used them) but they had two critical downsides: their performance went down with longer outputs, and they where a bit annoying to parallelize because computing the output for the 10th word required first computing the output of the previous 9 words - no way to use 10 parallel computers. The first problem was solved with Attention, a scaffolding method that prevented degradation over longer sequences. Eventually someone realized that Attention was doing most of the heavy lifting, built an attention-only network that could be easily parallelized (the Transformer), and LSTMs lost the top place. Are xLSTMs better? On paper I'd say they could be - they seem to have a solid theory and good results. Will they dethrone Transformers? My guess is no, as it wouldn't be the first time that the "better" technology ends up losing against whatever is popular. Having said that, it is entirely possible that some inherently recurrent tasks like stock price prediction could get a boost from this technology and they may find their place.
- jasonjmcghee 2y agoThey reference "a GPT-3 model with 356M parameters" So GPT-3 Medium (from the GPT-3 paper) - feels pretty disingenuous to list that as no one is referencing that model when they say "GPT-3", but the 175B model. I wasn't aware that size of the model (356M) was released- what am I missing here? I also think it's relatively well understood that (with our current methods) transformers have a tipping point with parameter count, and I don't know of any models less than ~3B that are useful- arguably 7B. Compare these benchmarks to, say, the RWKV 5/6 paper https://arxiv.org/abs/2404.05892 https://arxiv.org/abs/2404.05892
- CuriouslyC 2y agophi3 mini is surprisingly capable given its size. You can teach small transformers to do stuff well, you just can't have good general purpose small models.
- jasonjmcghee 2y agoTotally. But they aren't fine tuning these afaict- but comparing general purpose capabilities.
- Der_Einzige 2y agoThe point still stands, Phi3 is an excellent model and shows that good models don’t need that many parameters You should see the work on ReFT coming from mannings group showing that you can instruction fine tune models by modifying like, 0.00001% of the parameters. By doing it this way, you significantly mitigate the risk of catastrophic forgetting.