3 ms·
I have not, but I want to in near future because I'm really curious myself too. I've been following Rust community that now has llama.cpp port and also my OpenC
by adeon 4y ago
I have not, but I want to in near future because I'm really curious myself too. I've been following Rust community that now has llama.cpp port and also my OpenCL thing and one discussion item has been to run a verification and common benchmark for the implementations. https://github.com/setzer22/llama-rs/issues/4 https://github.com/setzer22/llama-rs/issues/4
I've mostly heard that, at least for the larger models, quantization has barely any noticeable effect. Would be nice to witness it myself.