4 ms·
Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their u
by scrlk 2mo ago
Nice to see a provider being transparent about KV cache quantisation. I've been suspecting that some providers do this silently whilst heavily promoting their unquantised weights, even though KV quantisation can degrade quality more than weight quantisation.
However, I wish their testing were more detailed. Firstly, some model families are more sensitive to KV quantisation than others (only Kimi K2.6 was tested). Secondly, the evaluation suite they use to claim that FP8 KV quantisation is indistinguishable is noticeably lacking coding benchmarks; in long-running tasks, minor tool call errors compound over time.
- anonova 2mo agovLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCodeBench 6. https://vllm-project.github.io/2026/04/22/fp8-kvcache.html https://vllm-project.github.io/2026/04/22/fp8-kvcache.html
- deleted 2mo ago[deleted]
- Tepix 2mo agovLLM tested with Llama-3.3-70B-Instruct, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, and Qwen3.5-27B. I'm wondering if it needs to be tested with every other model or not.
- scrlk 2mo agoIMO, yes. For example, Qwen 3.x is insensitive to weight and KV cache quantisation, whereas Gemma 4 is more sensitive: https://localbench.substack.com/p/kv-cache-quantization-benchmark https://localbench.substack.com/p/kv-cache-quantization-benc...
- amluto 2mo agoThey made an extremely strong claim: > None of this would matter if it changed the model's answers If they want to assert that the answers don’t change, then perhaps they should calculate the statistical distance between the token probability outputs or something to that effect. I doubt the results would indicate that the answers don’t change by any reasonable interpretation. Maybe the results are still good enough.
- scrlk 2mo agoKL divergence is your friend when it comes to evaluating the effects of quantisation: https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_diver...
- stingraycharles 2mo agoIsn’t that already in detail by the research of these quantization techniques?