4 ms·
Implementation of Google's Griffin Architecture – RNN LLM
- VHRanger 2y agoLike RWKV and Mamba, this is mixing some RNN properties to avoid the issues transformers have. However I'm curious about their scaling claims. They have a plot that shows how the model scales in training with the FLOPs you throw at it. But the issue we should rather be concerned with is the wall time of training for a set amount of hardware. Back in 2018, we could train medium sized RNNs, the issue was with wall time of training and training stability.
- whimsicalism 2y agotransformers were also just better at the LM task than 2018 RNNs for equal amount of flop training
- VHRanger 2y agoYeah, that's just the training stability part to my knowledge
- whimsicalism 2y agothey're also just less capable models. like just adding attention on top of an RNN made them a lot better
- SpaceManNabs 2y agoCalculating self-attention is still quadratic though. So you get the negatives of transformers there too.
- foota 2y agoDo you know the downside with RWKV? Based on how they present it, it seems like the best thing since sliced bread, but I would have assumed that it would have been widely adopted if that were the case.
- VHRanger 2y agoIt seems only OK as a model? Looking at the LLM chat leaderboard it's 71st and the 14B version is worse than a lot of 7B models: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboard https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar... Also, llama.cpp makes inference accessible for a lot of people, and it's not available for RWKV. Not to knock on the model, I'm sure it's good. I also like that it's a succesful example of citizen science. It's just not popular enough to have the inference infrastructure transformers have, not established enough to attract enough money to get 60B+ models trained, and so on.
- whimsicalism 2y agoi believe it is undertrained, at minimum
- WanderPanda 2y agoThis leaderboard is not the best for comparing model architectures, the dataset and finetuning have too much influence. I think perplexity on a particular dataset would be a better way to compare
- logicchains 2y ago>Also, llama.cpp makes inference accessible for a lot of people, and it's not available for RWKV. It absolutely is: https://github.com/RWKV/rwkv.cpp https://github.com/RWKV/rwkv.cpp .
- jimmyl02 2y agoFrom what I know about RWKV, it's mostly a one man effort and doesn't have the same data pipeline / resources as most major labs. It's a bit unfortunate but I'm curious about the performance given the same training corpus as OpenAI's GPTs. Maybe some labs have tried internally but haven't released results? On the other hand it makes sense to invest more money into transformer training runs as they have been proven to work. They really burst onto the scene and brought back RNNs in the world of transformers. The claim that RWKV isn't paralleizable during training also seems to be refuted in their readme. I'd guess it's generalizable performance as there is a difference between doing well on benchmarks and being usable. Personally I've tried running the weights a long time ago when it was first released and the results weren't usable but I'm sure there has been considerable progress since then.
- GaggiX 2y agoThe paper shows that the speed is comparable to transformer models, faster with smaller with "long" sequence length like 8k.
- riku_iki 2y agoI didn't get one detail: they selected 6B transformer as baseline and compared it to 7B Griffin Why wouldn't select equal size models?..
- szundi 2y agoThey probably had them for some reason and it was cheaper not to retrain one of them again
- riku_iki 2y agoIts just performance comparison is misleading then, they report marginal improvements which is expected just because of models size differences..
- GaggiX 2y agoIt also performs better on any other size.
- riku_iki 2y agoThey have baseline transformer of max size 6B in tables. Other models are trained on very different data and probably differently.
- GaggiX 2y agoAll the MQA transformers, Hawk and Griffin are trained on the same MassiveText dataset so no.
- riku_iki 2y agoYes, but MQA is limited to 6B size, while "other" larger non-RNN models in table(Llama-2) are not trained on the same dataset, and Hawk and Griffin are 7B. Sorry, I don't understand your point.
- spxneo 2y agoim not smart enough to know the significance of this...is Griffin like MAMBA?
- VHRanger 2y agoYes, like RWKV and Mamba this is a new generation of models that are more like big RNNs than pure transformers we have now
- boywitharupee 2y agoand is Griffin a state space model?
- imjonse 2y agoNo, it's a combination of RNN and Transformer.
- michwilinski 2y agoI mean, SSMs are in fact under the hood RNNs
- VHRanger 2y agoAt the end of the day, either you carry around a hidden state, or you have a fixed window for autoregression. You can call hidden states "RNN-like" and autoregressive windows "transformer-like", but apart from those two core paradigms I don't know of other ways to do sequence modelling. Mamba/RWKV/Griffin are somewhere between those two extremes.
- stri8ed 2y agoIsn't that how previous models were, before the attention is all you need paper?
- janwas 2y agoFor anyone interested in a C++ implementation, our github.com/google/gemma.cpp now supports this model.
- JyrkiAlakuijala 2y agoFun fact -- gemma.cpp uses highway, an amazing high performance computation library originally developed in the JPEG XL effort.