4 ms·
This is getting very close to fit a single 3090 with 24gb VRAM :)
by vladgur 6mo ago
This is getting very close to fit a single 3090 with 24gb VRAM :)
- originalvichy 6mo agoYup! Smaller quants will fit within 24GB but they might sacrifice context length. I’m excited to try out the MLX version to see if 32GB of memory from a Pro M-series Mac can get some acceptable tok/s with longer context. HuggingFace has uploaded some MLX versions already.
- ycui1986 6mo ago32GB RAM on mac also need to host OS, software, and other stuff. There may not even be 24GB VRAM left for the model.
- donmcronald 6mo agoI have an Mini M4 Pro with 64GB of 273GB/s memory bandwidth and it's borderline with 3.5-27B. I assume this one is the same. I don't know a ton, but I think it's the memory bandwidth that limits it. It's similar on a DGX Spark I have access to (almost the same memory bandwidth). It's been a while since I tried it, but I think I was getting around 12-15 tokens per second an that feels slow when you're used to the big commercial models. Whenever I actually want to do stuff with the open source models, I always find myself falling back to OpenRouter. I tried Intel/Qwen3.6-35B-A3B-int4-AutoRound on a DGX Spark a couple days ago and that felt usable speed wise. I don't know about quality, but that's like running a 3B parameter model. 27B is a lot slower. I'm not sure if I "get" the local AI stuff everyone is selling. I love the idea of it, but what's the point of 128GB of shared memory on a DGX Spark if I can only run a 20-30GB model before the slow speed makes it unusable?
- verdverm 6mo agoThere are a number of DGX benchmarks for these recent gemma-4 / qwen-3.6 models on the nvidia forum, ex: https://forums.developer.nvidia.com/t/qwen-qwen3-6-35b-a3b-and-fp8-has-landed/366822/5 https://forums.developer.nvidia.com/t/qwen-qwen3-6-35b-a3b-a...
- girvo 5mo agoTbf the Sparks usefulness isn’t for inference IMO. Its memory bandwidth is too low for that. But on the other hand, running Qwen 3.5 122B A10B locally on it using ~110GB of memory and getting 50tk/s generation and quite excellent prefill… I couldn’t do that on many other machines at this price point For me this has been awesome to learn CUDA on, fine tuning models (until I get it close to what I want then it’s off to H100 or something clusters) and a bit of inference on the side
- GaggiX 6mo agoAt 4-bit quantization it should already fit quite nicely.
- Aurornis 6mo agoUnfortunately not with a reasonable context length.
- kkzz99 6mo agoIt really depends on what you think a reasonable context length is, but I can get 50k-60k on a 4090.
- GaggiX 6mo agoThe model uses Gated DeltaNet and Gated Attention so the memory usage of the KV cache is very low, even at BF16 precision.
- regularfry 6mo agoI've got 139k context with the UD-Q4_K_XL on a 4090, q8_0 ctk/v. Could probably squeeze a little more but that's enough for me for the moment.
- corysama 6mo agoHey, buddy! Can I bum a command line arg list off ya?
- skiing_crawling 6mo agoI used to run qwen3.5 27b Q4_k_M on a single 3090 with these llama-server flags successfully: `-ngl 99 -c 262144 -fa on --cache-type-k q4_0 --cache-type-v q4_0`
- chr15m 5mo agoWith CPU offloading of e.g. 25% on that hardware it is still fast enough for a lot of things.