9 ms·
Eagle 7B: Soaring past Transformers
- viktor-ferenczi 3y agoAfter reading all these, I've landed at the following conclusion: Regardless of its architecture there is only a finite amount of information the language model can work with at any given time. It depends on the task at hand which way of "forgetting" causes the least problems. For coding and math a perfect context with a well defined maximum length of 16k..256k tokens paired with high quality ICL would work better than automated "random" forgetting. However, it requires a good strategy to present only the information relevant for the task to fit into the maximum context length. For free-form literature and other non-technical stuff automated forgetting is likely beneficial, because you don't need to come up with a strategy to choose what's important to keep in-context. What you get is automated gradual forgetting and "mixing up" past memories, just like in humans. Since I'm a software developer geek I strongly prefer the first one, but as you can see, it depends on the task at hand.
- htrp 3y ago> We are releasing RWKV-v5 Eagle 7B, licensed as Apache 2.0 license, under the Linux Foundation, and can be used personally or commercially without restrictions great on the team to actually set up the right incentives for testing and adoption.
- karterk 3y agoIt's interesting how all focus is now primarily on decoder-only next-token-prediction models. Encoders (BERT, encoder of T5) are still useful for generating embedding for tasks like retrieval or classification. While there is a lot of work on fine-tuning BERT and T5 for such tasks, it would be nice to see more research on better pre-training architectures for embedding use cases.
- jeremycochoy 3y agoI believe RWKV is actually an architecture that can be used for encoding: given a LSTM/GRU, you can simply take the last state as an encoding of your sequence. The same should be possible with RWKV, right?
- nickthesick 3y agoIs it possible to try this out on something like llama.cpp? Does the different architecture make a difference there?
- deleted 3y ago[deleted]
- kristjansson 3y agoThere's https://github.com/saharNooby/rwkv.cpp https://github.com/saharNooby/rwkv.cpp, which related-ish[0] to ggml/llama.cpp [0]: https://github.com/ggerganov/llama.cpp/issues/846 https://github.com/ggerganov/llama.cpp/issues/846
- coder543 3y agoIt’s cool that progress is being made on alternative LLM architectures, and I did upvote the link. However, I found this article somewhat frustrating. Showing the quality of the model is only half of the story, but the article suddenly ends there. If people are going to be motivated to adopt an entirely different architecture, then performance and context size deserve at least as much discussion. Given the linear nature being shown, it seems like the primary thing people are going to want to see is the next frontier of LLMs: context sizes of ~1M tokens. The word “context” does not even appear in this article, which is disappointing. If there were a discussion of context, it would be nice to see if it passes the passkey test. The article also appears to reuse a chart from RWKV-4 showing how awesome a linear function is compared to a quadratic one, but… cool story? It’s not even clear what this chart is truly showing. Is this chart only showing generated tokens, or is this including prompt tokens? As I have never used RWKV, I have no idea how the prompt processing speed compares to the token generation speed. Prompt processing speed has been a big problem for Mixtral, for example. As a reader, I want to see a couple of actual examples of X prompt tokens + Y generated tokens, and the tokens/s of X and Y for RKWV-5 and for Mistral on the same hardware. On the Mistral side, it is trivial to collect this information in llama.cpp, but I don’t know how the tooling is for RWKV.
- Harrisonv 3y agoFor linear transformers, the current metric is "perfect token recall", the ability for the model to recall a randomized sequence of data. You can find the limit of a particular model architecture by training a model of a particular size to echo randomized data, and I believe this was touched on in the zoo-ology paper. This doesnt prevent the model from retaining sequences or information beyond this metric, as information can easily be compressed in the state, but it anything within that window can be perfectly recalled by the model. Internal testing has placed the value for Eagle around the 2.5k ptr[perfect token recall] mark, while community fine tunes done on the partial checkpoints for long distance information gathering and memorization have been shown to easily dwarf that. prompt processing speed benefits from the same gemm optimizations as standard transformers, with the extra benefit of those gemm optimizations working for batch inference as well (no need for vllm as memory allocation is static per agent)
- 3y ago
- visarga 3y agoThis shows the model architecture, be it transformer, Mamba, SSM or RWKV - doesn't really matter when compared to the impact of the training set. We're spending too much time debating models when we should be talking about language data, a reservoir of human experience won at great sacrifice by humanity. And the same data when used to train humans creates modern capable people. Alone, without society and language, we would be mere shadows of ourselves. What does it say when AI acquires so many capabilities from language data? maybe intelligence was not centered in the brain. It's a social process.
- jack_pp 3y agoOf course it is in the brain, the brain created and evolved the language as a very powerful tool. If intelligence was in the language then other animals would be as intelligent as us
- klipt 3y agoScientific advancement requires both brains and knowledge transfer over generations. "If I have seen further, it is by standing on the shoulders of giants."
- deleted 3y ago[deleted]
- mediaman 3y agoKnowledge transfer over generations is a function of the brain. Other species have much more limited ability to transfer knowledge intergenerationally, and that is because the human brain's capability for symbolic language is much more advanced than other animals', who are not able to encode knowledge nearly as efficiently.
- klipt 3y agoThe point is it's a function of many connected brains, not just one brain.
- vessenes 3y agoMy experiments with RWKV-4 showed good inference speed but suuper slow tokenization speed; I’m not sure if this was RWKV specific or implementation specific: it was some time ago. Any guidance on this for rwkv-5?
- adt 3y agohttps://lifearchitect.ai/models-table/ https://lifearchitect.ai/models-table/
- black_puppydog 3y agoIIUC this model type makes the "ALScore" column completely pointless, because the quality of results isn't related to the number of weights in the same way as for regular transformers.
- ComplexSystems 3y agoCan someone knowledgeable about this stuff maybe explain what it means? What is the context and how does this compare to the usual transformer models? I don't get how to interpret some of these benchmarks. It looks like it's as good as Mistral 7B/mistral-tiny?
- ilaksh 3y agoForgive me for not searching around enough but so far I am not sure about: - how much RAM is needed - how many tokens per second with CPU only, like a typical VM/VPS for example
- column 3y agoit seems it does not support CPU only
- _hl_ 3y agoAttention is, after all, not what you need.
- jeremycochoy 3y agoWell, RWKV is using some linear (non-quadratic) form of attention so... strictly speaking... you still need a bit of attention :D
- 127361 3y agoThey've joined the Linux Foundation, does that mean the models are going to be eventually censored to satisfy the foundation's AI safety policies? That includes ensuring the models don't generate content that's non-inclusive or against diversity policies?
- lhl 3y agoIt is trivial to fine tune any model (whether a base model or an aligned model) to your preferred output preferences as long as you have access to the model weights.
- Al-Khwarizmi 3y agoNot trivial for the general public at all, and furthermore, you need much more memory for finetuning than for inference, often making it infeasible for many machine/model combinations.
- lhl 3y agoIf you are running a local LLM already (which no one in the "general public is") then the bar is really not that much higher for fine-tuning (either for an individual or community member to do). And you don't need any additional equipment at all. When I say trivial, I really do mean it - you can go to https://www.together.ai/pricing https://www.together.ai/pricing and see for yourself - a 10M token 3 epoch fine tune on a 7B model will cost you about $10-15 right now. Upload your dataset, download your fine tune weights (or serve via their infrastructure). This is only going to get easier (compare how difficult it was to inference local models last year to what you can do with plug and play solutions like Ollama, LM Studio, or Jan today). Note also that tuning is a one-time outlay, and merges are even less resource intensive/easier to do. To put things in perspective, tell me how much cost and effort it would be to tune a model where you don't have the weights at all in comparison.
- Al-Khwarizmi 3y agoRunning a local LLM - downloading LM studio, installing on Windows, using the search function to search for a popular LLM, click "download", click the button to load the model, chat. Fine-tuning - obtaining a dataset for your task (this in itself is not trivial), figuring out how the service you linked works (after figuring out that it exists at all), uploading the dataset, paying, downloading the weights - OK, now how do you load them into LM studio? It's all subjective, of course, but for me there's a considerable difficulty jump there.
- Harrisonv 3y agoTry rwkv-demo-api.recursal.ai if you want to try it and dont want to wait for gradio
- zurfer 3y agoThank you! I'm not experienced with 7B models. 3 things stand out to me: - it's absolutely not useable for the kind of use cases I solve with GPT-4 (code generation, information retrieval) - it could technically swallow a 50 page PDF, but it's not able to answer questions about it (inference speed was good, but content was garbage) - it is ok for chatting and translations (how is your day?)
- deleted 3y ago[deleted]
- lelag 3y agoI also tried the demo and I find it pretty much useless at most things even comparing it to a small 7b transformer model like mistral. From my albeit quick tests, what I found is that it knows clearly less things than mistral, it hallucinates much more, it does not follow instructions, has less reasoning capabilities and asking it to translate a Japanese text into English gave me a bad translated summary instead of the full translation. I don't see how this is soaring past transformers when clearly it's unable to do any of the useful tasks you can use a transformer model for today...
- samus 3y agoHas anybody else noticed the map indicating that there are barely any fluent English speakers outside of the "English-speaking" countries? Or is the threshold for fluency so high that no place even in Europe except the Ireland and the UK qualify at all?
- airspresso 3y agoThis caught my eye as well. I interpreted it as the map only shows countries that have English as primary language. But the article was not precise enough about this. I applaud them for focusing on multilingual performance though, as that is an important area of NLP which still has lots of room for improvement.
- samus 3y agoI highly welcome the effort as well*, but I don't see why they would have to mistake first-language ability for fluency to argue for that. The difference is vast and relevant: anyone with good English reading and writing skills can take advantage of a model and might prefer it over a worse model in their native language. *: Just sceptical whether there's enough content out there which isn't just (often badly or too straightforwardly) translated from English. Not an issue for the languages with let's say >10 Mio. speakers, but for everything smaller.
- pico_creator 3y agoYea, thats why I focused only on the top 25 languages, despite the model being trained for 100+ languages. Was not confident, on the languages beyond the 25th cut-off, until we build better datasets (which we are in works on with various regional groups!)
- vidarh 3y agoNorwegian has ca. 5m speakers, and ChatGPT does not just do fine with both the (mutually intelligible) Norwegian written languages, but also has no problem "translating" to/from several regional dialects when I've experimented. And that is, I presume - I could be wrong -, before anyone has tried to really mine the Norwegian national library, as even much of what is online is not easily accessible for crawling. I think there'll be plenty of content for even much smaller languages - especially anywhere with depositary laws -, but it's often going to require cumbersome collection efforts and negotiating access.
- ReptileMan 3y agoTheir map showing distribution of English speaking people is just terrible - I am fairly sure that there is at least one percent speaking English in India, Western Europe, Eastern Europe, Russia and China.
- samus 3y agoIt must set a terribly high threshold (like language certificate holders or graduates of English-speaking schools) or actually report the percentage of native speakers. But one only has to be fluent enough to write chat messages to use a text model!
- vidarh 3y agoYeah, that's weird. Depending on how you count least India, Pakistan, and Nigeria has more English speakers than the UK in absolute terms (Nigeria might end up either side depending on how strict you are), and Nigeria is on its path to overtake the UK as the country with the second largest number of native English speakers, as it's increasingly often the 2nd language of parents with different 1st languages. E.g. my ex's parents had Igbo and Yoruba as their 1st languages, but she and all her siblings has English as theirs.
- dizhn 3y agoFYI a lot of people are asking questions which one of the project members has been answering on Reddit the last few days. https://www.reddit.com/user/PicoCreator/ https://www.reddit.com/user/PicoCreator/
- jug 3y ago> [Mar 2024] An MoE model based on the v5 Eagle 2T model (note, approximate date) Hyped about this! This could strike a powerful balance between performance and reasonably retained low environmental/token cost impact. Would be cool with improved coverage of Scandinavian languages along with it, but I guess we'll see. And yeah, I think a true revolution will happen (or might already be) when we realize the value of training data and how to structure and balance its content in the most optimal way for training.
- senseiV 3y agoLooking into the nordic pile maybe? There are some datasets
- viraptor 3y ago> A common feedback we receive for the RWKV multi-lingual approach is: it hurts our English evaluation scores (...) Has anyone quantified that specifically? I'd love to read more details since I'd expect the concepts to start mapping between languages at some point. I.e. with enough language fluency I'd expect learning knowledge/reasoning in language to improve the result in another. But I can't find any paper talking about specifically about that.
- marmaduke 3y agoI think the group working on rwkv has yes, even if they don't show the details in this article. (I followed them on their discord channels for quite some time). The interesting take away for me was that training rwkv from zero to intelligible sentences for minority language was more faster (in units of tokens trained!) than other architectures, making it more accessible for cases where large corpus like the Pile don't or can't exist.
- sanxiyn 3y agohttps://arxiv.org/abs/2306.06688 https://arxiv.org/abs/2306.06688 studied this. (Check out figure 3.) In short, yes, multilingual transfer does happen. But importantly, not at 7B. Being English-centric is possibly okay for LLaMA 70B, but it really is not okay for LLaMA 7B.
- YeGoblynQueenne 3y ago>> An eagle, flying past a transformer-looking robot That hero image is a complete mess (e.g. look at the eagle's forward paw, the "transformer robot"'s right arm or the position of the eagle's left wing). Why is it that people put up such obviously messed-up images in their articles? Do they not see that level of detail, or do they just find it cool to have some "AI art" in their article, as a kind of an in-group code, like "we use AI"?
- pmx 3y agoI mean, They're an organisation creating AI models, releasing an AI model. Them using AI images isn't hugely shocking to anyone with the capacity to reason. The picture looks cool at first glance, and that is really all that matters. It's just a cool bit of eye-candy leading into the article.
- YeGoblynQueenne 3y ago>> Them using AI images isn't hugely shocking to anyone with the capacity to reason. Gee, thanks, that's so kind.
- RcouF1uZ4gsC 3y agoSounds like this hallucinates pretty easily From Reddit : https://www.reddit.com/r/LocalLLaMA/comments/1ad0j63/rwkv5_100_trained_released/ https://www.reddit.com/r/LocalLLaMA/comments/1ad0j63/rwkv5_1... User: Which is larger, a chicken egg or a cow egg? Assistant: To determine which is larger, a chicken egg or a cow egg, let's first look at their respective sizes and compare them. Chicken Egg: The average chicken egg size ranges from 2.5 to 3 inches (6 to 8 cm) in length and 1.5 to 2 inches (3.8 to 5 cm) in diameter.[1] Cow Egg: The average cow egg size is roughly 2 inches (5 cm) in length and 1 inch (2.5 cm) in diameter.[2] Based on these measurements, we can see that the chicken egg is slightly larger than the cow egg, with dimensions of around 3 inches by 1 inch versus 2 inches by 0.75 inches respectively (length x diameter). However, as chickens are much bigger than cows on average,[3] it would be more accurate to say that a chicken egg is roughly three times bigger than a cow egg based on their respective sizes within each species. So, the answer is: A chicken egg is three times larger than a cow egg in terms of size when taking into account their respective dimensions within each species' average range.
- yakireev 3y agoI absolutely should not anthropomorphise LLMs, but I can't get rid of the feelling that "it" was writing this answer with a mischievous smirk, and was having a lot of fun in the process. The future is weird.