4 ms·
That's great. Personally, I'd interested in Qwen3.6-27B and deepseek V4 flash (or pro), with contexts above 60k. They seem to be popular and have good coding pe
by russianGuy83829 3mo ago
That's great. Personally, I'd interested in Qwen3.6-27B and deepseek V4 flash (or pro), with contexts above 60k. They seem to be popular and have good coding performance. I'd appreciate numbers on a single or two GPUs where a quantized version fits reasonably into the VRAM (Qwen in 16 or 24GB). 4 older GPUs approach a used 3090 in price, and the 3090 has better support for speedups like MTP. So cheaper but slower looks like a reasonable target to me.
- eso_logic 3mo agoNo problem. Varying context size is a common request I've been getting as well. Personally I'm looking forward to seeing how much we can cram into the ancient K80's 24GB of VRAM :0
- russianGuy83829 3mo agoThank you, looking forward to it. I just saw this simple patch to enable MTP (potentially 2x performance) on older GPUs (Kepler etc), so maybe it will work for you https://github.com/ggml-org/llama.cpp/pull/25680 https://github.com/ggml-org/llama.cpp/pull/25680 Also, for Qwen, the 4 bit _XL quantization seems to have a good balance of performance to size.
- NortySpock 3mo agoSimilar interest here, possibly including if qwen 3.6, Gemma4 or DiffusionGemma (with the largest quants that will fit in a single card) will offer, say, 50 tokens-per-second (fast enough for interactive human-in-the-loop code research, print-f iterations on code to debug things, etc; or let the LLM churn on a problem for a minute while I step out to handle something else), context of up to 200k preferred. Also if nothing else the below project lets you use an NVidia graphics card as low-latency swap, which has been nice as a buffer as RAM prices remain high and leaves me eyeing that 24GB card you mentioned as an alternative... https://github.com/c0deJedi/nbd-vram https://github.com/c0deJedi/nbd-vram