4 ms·
I've counted three different Rust LLaMA implementations on r/rust subreddit this week: https://github.com/Noeda/rllama/ https://github.com/Noeda/rllama/ (pure
by adeon 4y ago
I've counted three different Rust LLaMA implementations on r/rust subreddit this week:
https://github.com/Noeda/rllama/ https://github.com/Noeda/rllama/ (pure Rust+OpenCL)
https://github.com/setzer22/llama-rs/ https://github.com/setzer22/llama-rs/ (ggml based)
https://github.com/philpax/ggllama https://github.com/philpax/ggllama (also ggml based)
There's also a discussion on GitHub issue on setzer's repo to collaborate a bit on these separate efforts: https://github.com/setzer22/llama-rs/issues/4 https://github.com/setzer22/llama-rs/issues/4
- hummus_bae 4y agocorresponding reddit threads: - https://www.reddit.com/r/rust/comments/6jm58w/rllama_a_rust_llama_implementation/ https://www.reddit.com/r/rust/comments/6jm58w/rllama_a_rust_... - https://www.reddit.com/r/rust/comments/6jm6gk/off_by_one_abstraction_a_Rust_LLama_instead_of/ https://www.reddit.com/r/rust/comments/6jm6gk/off_by_one_abs... - https://www.reddit.com/r/rust/comments/6jmpu3/llama_rs_simple_pure_rust_llama/ https://www.reddit.com/r/rust/comments/6jmpu3/llama_rs_simpl...
- comex 4y agoTwo of those links don’t work.
- comex 4y agoDo you know if any of them support GPTQ [1], either end-to-end or just by importing weights that were previously quantized with GPTQ? Apparently GPTQ provides a significant quality boost “for free”. I haven’t had time to look into this in detail, but apparently llama.cpp doesn’t support it yet [2] though it will soon. And the original implementation only works with CUDA. [1] https://github.com/qwopqwop200/GPTQ-for-LLaMa/ https://github.com/qwopqwop200/GPTQ-for-LLaMa/ [2] https://github.com/ggerganov/llama.cpp/issues/9 https://github.com/ggerganov/llama.cpp/issues/9
- adeon 4y agoAs far as I know, the two ggml ones are basically just llama.cpp ports that include the ggml source code so if the support is not in llama.cpp, I don't think it's in these implementations either. Although maybe that also means that they'll gain that ability as soon as llama.cpp does. I'm the author of the last one, rllama and it has no quantization whatsoever. I don't think any of these are improvements over llama.cpp for end-users at this time. Unless you really really really want your software to be in Rust in particular.