30 ms·
Looking at the huggingface page, the unsloth people haven't finished quantizing it yet, but I'm sure they're active on it right now. It'll be interesting to see
by walrus01 1mo ago
Looking at the huggingface page, the unsloth people haven't finished quantizing it yet, but I'm sure they're active on it right now. It'll be interesting to see how the capabilities and benchmark tests compare on system where it can fit in under 512GB of RAM with full context.
In terms of coding and command line capabilities I'm also very interested to see a head-to-head of it vs. qwen 3.8-flash-next Q8 which is something like 190GB of memory used when loaded into llama-server. It fits very well in all sorts of 256GB or under class machines.
- lowbloodsugar 29d ago3.8-flash-next fits on a single 6000 at Q4 if you offload the PLE. Crazy fast and still effective.
- agile-gift0262 29d agoSorry for the tangent, but how does Qwen3.8-flash-next compare to DeepSeek v4 Flash? I still haven't found the time to set it up, but I'm really happy with DeepSeek v4 Flash
- lowbloodsugar 29d agoIt was the best model given my constraints (RTX PRO 6000 96gb + 256GB DDR4), when run against rust programming benchmarks. For Qwen3.8-flash-next NVFP4 and the latest vllm container, the PLE is 100GB of main ram, and everything else runs on the GPU with room for a total of 560k tokens (two full 262k conversations). DeepSeek has to offload a ton to the CPU and it performed worse than Qwen in absolute terms and was a lot slower (not usable). If you have enough room to run DeepSeek v4 Flash comfortably then you can likely run the Q8 of the qwen model.