2 ms·
vLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCo
by anonova 2mo ago
vLLM's study also concluded that "FP8 can deliver meaningful latency and capacity gains with small or negligible accuracy loss". Their benchmarks include LiveCodeBench 6.
https://vllm-project.github.io/2026/04/22/fp8-kvcache.html https://vllm-project.github.io/2026/04/22/fp8-kvcache.html
- deleted 2mo ago[deleted]
- Tepix 2mo agovLLM tested with Llama-3.3-70B-Instruct, Qwen3-30B-A3B-Instruct-2507, Qwen3-30B-A3B-Thinking-2507, and Qwen3.5-27B. I'm wondering if it needs to be tested with every other model or not.
- scrlk 2mo agoIMO, yes. For example, Qwen 3.x is insensitive to weight and KV cache quantisation, whereas Gemma 4 is more sensitive: https://localbench.substack.com/p/kv-cache-quantization-benchmark https://localbench.substack.com/p/kv-cache-quantization-benc...