3 ms·
There was a study specifically related to Qwen3.8 27B that showed that kv cache quantization has almost no impact on this model all the way to q4: https://arxi
by skolos 26d ago
There was a study specifically related to Qwen3.8 27B that showed that kv cache quantization has almost no impact on this model all the way to q4:
https://arxiv.org/html/2609.04098 https://arxiv.org/html/2609.04098