3 ms·
The benchmark only touches 8B-class models at 8-bit quantification. Would be interesting to see how it fares with models that use more of the card ram, and unde
by benob 2y ago
The benchmark only touches 8B-class models at 8-bit quantification. Would be interesting to see how it fares with models that use more of the card ram, and under varying quantization and context lengths.
- Maxious 2y agoThere's been some INT4/NVFP4 gains too https://hanlab.mit.edu/blog/svdquant-nvfp4 https://hanlab.mit.edu/blog/svdquant-nvfp4 https://blackforestlabs.ai/flux-nvidia-blackwell/ https://blackforestlabs.ai/flux-nvidia-blackwell/
- threeducks 2y agoI agree. This benchmark should have compared the largest ~4 bit quantized model that fits into VRAM, which would be somewhere around 32B for RTX 3090/4090/5090. For text generation, which is the most important metric, the tokens per second will scale almost linearly with memory bandwidth (936 GB/s, 1008 GB/s and 1792 GB/s respectively), but we might see more interesting results when comparing prompt processing, speculative decoding with various models, vLLM vs llama.cpp vs TGI, prompt length, context length, text type/programming language (actually makes a difference with speculative decoding), cache quantization and sampling methods. Results should also be checked for correctness (perplexity or some benchmark like HumanEval etc.) to make sure that results are not garbage. If anyone from Phoronix is reading this, this post might be a good point to get you started: https://old.reddit.com/r/LocalLLaMA/comments/1h5uq43/llamacpp_bug_fixed_speculative_decoding_is_30/m08poqr/ https://old.reddit.com/r/LocalLLaMA/comments/1h5uq43/llamacp... At time of writing, Qwen2.5-Coder-32B-Instruct-GGUF with one of the smaller variants for speculative decoding is probably the best local model for most programming tasks, but keep an eye out for any new models. They will probably show up in Bartowksi's "Recommended large models" list, which is also a good place to download quantized models: https://huggingface.co/bartowski https://huggingface.co/bartowski
- regularfry 2y agoUsing aider with local models is a very interesting stress case to add on top of this. Because the support for reasoning models is a bit rough, and they aren't always great at sticking to the edit format, what you end up doing is configuring different models for different tasks (what aider calls "architect mode"). I use ollama for this, and I'm getting useful stuff out of qwq:32b as the architect, qwen2.5-coder:32b as the edit model, and dolphin3:8b as the weak model (which gets used for things like commit messages). Now what that means is that performance swapping these models in and out of the card starts to matter, because they don't all go into VRAM at once; but also using a reasoning model means that you need straight-line tokens per second as well, plus well-tuned context length so as not to starve the architect. I haven't investigated whether a speculative decoding setup would actually help here, I've not come across anyone doing that with a reasoner before now but presumably it would work. It would be good to see a benchmark based on practical aider workflows. I'm not aware of one but it should be a good all-round stress test of a lot of different performance boundaries.