3 ms·
Better keep the KV cache in full precision
by ggerganov 6mo ago
Better keep the KV cache in full precision
- freakynit 6mo agoWow.. the GOAT himself.. thank you sooo much for creating llama.cpp ... will re-deploy with full kv cache once requests stop coming.