3 ms·
LLama 3 405B had the most unoptimized kv cache usage by far. Deepseek v4 pro uses 2.4GB for the same context length[1]. [1]: https://vllm.ai/blog/2026-04-24-de
by YetAnotherNick 2mo ago
LLama 3 405B had the most unoptimized kv cache usage by far. Deepseek v4 pro uses 2.4GB for the same context length[1].
[1]: https://vllm.ai/blog/2026-04-24-deepseek-v4 https://vllm.ai/blog/2026-04-24-deepseek-v4
- philipportner 2mo agoGood point, thanks! I haven't been keeping up with most of the new model internals.