3 ms·
Quantisation thankfully is applicable to RWKV as much as transformers. Most notably in our RWKV.cpp community project: https://github.com/saharNooby/rwkv.cpp ht
by pico_creator 3y ago
Quantisation thankfully is applicable to RWKV as much as transformers. Most notably in our RWKV.cpp community project: https://github.com/saharNooby/rwkv.cpp https://github.com/saharNooby/rwkv.cpp
Tooling/Ecosystem is something that I am actively working on as there is still a gap to transformers level of tooling. But i'm glad that there is a noticeable difference!
And yes! experiments are important, to ensure improvements in the architecture. Even if "Linear Transformers" replaces "Transformers". Alternatives should always be explored, to learn from such trade-offs to the benefit of the ecosystem
(This was lightly covered in the podcast, where I share IMO that we should have more research into text based diffusion networks)
- lhl 3y agoGlad to hear that you're focused on improving tooling, as you point out, as there is definitely a gap (that maybe I was too circumspect about it originally, but to be clear, difference wasn't in RKWV's favor). Some of it may just be the docs. I think one easy improvement is that when you search for RWKV, you end up on the Github page, but the README basically goes from a description of the model, to a short snippet of sample inference code, to a list of implementations, to a bunch of random research notes? It's a bit scattered and pretty hard to read, but should be easy to clean up. From a practical perspective for example, there's no list/description in the repo or the model cards on all the different versions at https://huggingface.co/BlinkDL https://huggingface.co/BlinkDL (world, raven, pile, novel, pileplus, etc) and while it says RWKV-4-World is the "best" model, it's only a 7B and Raven 14B is listed above and says it's fine-tuned (while World is not)? Which is actually better? Unsure, no benchmarks (even the spreadsheet screenshots eventually listed don't give any info, nor do the model cards) - even as someone returning to RWKV this was very confusing. Compare it to what the Llama community is doing these days with model cards (eg https://huggingface.co/Open-Orca/OpenOrcaxOpenChat-Preview2-13B https://huggingface.co/Open-Orca/OpenOrcaxOpenChat-Preview2-... or for quants https://huggingface.co/TheBloke/WizardCoder-Guanaco-15B-V1.0-GPTQ https://huggingface.co/TheBloke/WizardCoder-Guanaco-15B-V1.0...) Also, for those not familiar with the project at all, I think a slightly expanded/better organized "Usage" section I think would be helpful in terms of getting up and running (it doesn't help that multiple implementations each have their own quant formats). I feel like have a section that would get a user easily up and running (I think ExLlama and MLC both do a good job here) would be helpful. Again, some of this might just be a pass of cleaning up/better organizing the docs.
- pico_creator 3y agoThanks for clarifying, and pointing out where docs can be improved. For most parts I am trying to ensure everything do get properly reorganised under the official wiki page here. I have just added a guideline of sort to help better navigate which model should be used accordingly, based on your feedback. https://wiki.rwkv.com/#which-rwkv-models-should-i-be-using https://wiki.rwkv.com/#which-rwkv-models-should-i-be-using Eventually, from an SEO stand-point, the first landing for RWKV should eventually be optimised to be the wiki or rwkv.com page itself. And not blinkDL original trainer. This would be more aligned with the fact that the vast majority of search for the model would naturally be individuals who would want to try the model (and not train it / finetune from scratch).
- lhl 3y agoThe wiki looks like a much better organized starting point! I wonder if that could simply be linked after the intro in the BlinkDL repo until search rankings catch up: "For those looking to get started using or testing out RWKV models, visit: https://wiki.rwkv.com/ https://wiki.rwkv.com/" I do also think that having a recent benchmark table w/ the latest models and reference to a few similar sized transformer base models (but most importantly llama, llama2) would be super useful - maybe a table w/ memory size at various context lengths, or other ways to highlight RWKV's unique advantages. While local LLMs are still niche, I do think a lot of why RWKV gets overlooked is because if it has lower capabilities, doesn't inference faster, and is harder to set up, it's not exactly clear why to check it out (beyond as a curiosity).