4 ms·
Llama.cpp already uses an idea from it internally for the KV cache [0] So a quantized KV cache now must see less degradation [0] https://github.com/ggml-org/l
by kgeist 6mo ago
Llama.cpp already uses an idea from it internally for the KV cache [0]
So a quantized KV cache now must see less degradation
[0] https://github.com/ggml-org/llama.cpp/pull/21038 https://github.com/ggml-org/llama.cpp/pull/21038