3 ms·
This is my experience too. Qwen optimizes for a lot of scenarios which masks their weaker generalization compared to US frontier models. Never go below an fp1
by CMay 3mo ago
This is my experience too. Qwen optimizes for a lot of scenarios which masks their weaker generalization compared to US frontier models.
Never go below an fp16 kv cache unless you've already tested it in advance with your model on a verified task that you know it can successfully complete. People should also test the difference using the exact same seed value so they can see how the tokens diverge. If you have memory constraints, sometimes you can still use an fp16 kv cache and use storage for an agentic buffer to work your task with mixed abstractions rather than having everything in memory.
For 4-bit weight quants, Gemma 4 31B QAT is where people should be looking instead of Qwen 3.6.
- beacon294 3mo agoI find that fp8 cache can be pretty bad in vllm but works fine in llama.cpp. I don't know why, but I plan to review the implementations.
- CMay 3mo agoLlama.cpp implemented some rotation optimizations for quantized kv cache to improve the preservation of attention quality or similar, after everyone was talking about TurboQuant. It's not perfect and when you're talking about long form reasoning, little differences can make or break the results so it is situational.
- beacon294 3mo agoI'll read it. It could be the quants too. Some quants I try are inexplicably bad, some seem better than official (or unsloth) quants... even what should be run of the mill gguf quantization.