3 ms·
As far as I know, the two ggml ones are basically just llama.cpp ports that include the ggml source code so if the support is not in llama.cpp, I don't think i
by adeon 4y ago
As far as I know, the two ggml ones are basically just llama.cpp ports that include the ggml source code so if the support is not in llama.cpp, I don't think it's in these implementations either. Although maybe that also means that they'll gain that ability as soon as llama.cpp does.
I'm the author of the last one, rllama and it has no quantization whatsoever. I don't think any of these are improvements over llama.cpp for end-users at this time. Unless you really really really want your software to be in Rust in particular.