3 ms·
What is the min VRAM this can run on given it is MOE?
by tmaly 6mo ago
What is the min VRAM this can run on given it is MOE?
- mncharity 6mo agoFwiw, with its predecessor's Qwen3.5-35B-A3B-Q6_K.gguf, on a laptop's 6 GB VRAM and 32 GB RAM, with default llama.cpp settings, I get 20 t/s generation.
- rubiquity 6mo agoHave you tried running llama.cpp with Unified Memory Access[1] so your iGPU can seamlessly grab some of the RAM? The environment variable is prefixed with CUDA but this is not CUDA specific. It made a pretty significant difference (> 40% tg/s) on my Ryzen 7840U laptop. 1 - https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#unified-memory https://github.com/ggml-org/llama.cpp/blob/master/docs/build...
- zozbot234 6mo agoYour link seems to be describing a runtime environment variable, it doesn't need a separate build from source. I'm not sure though (1) why this info is in build.md which should be specific to the building process, rather than some separate documentation; and (2) if this really isn't CUDA-specific, why the canonical GGML variable name isn't GGML_ENABLE_UNIFIED_MEMORY , with the _CUDA_ variant treated as a legacy alias. AIUI, both of these should be addressed with pull requests for llama.cpp and/or the ggml library itself.
- rubiquity 6mo agoYou are right that it is an environment variable, and that's how I have it set in my nix config. Thanks for correcting that. Unfortunately llama.cpp is somewhat notorious for having lackluster docs. Most of the CLI tools don't even tell you what they are for.
- mncharity 6mo agoHmm. Perhaps there's a niche for a "The Missing Guide to llama.cpp"? Getting started, I did things like wrapping llama-cli in a pty... and only later noticing a --simple-io argument. I wonder if "living documents" are a thing yet, where LLMs keep an eye on repo and fora, and update a doc autonomously.
- mncharity 6mo agoI hadn't tried that, thanks! I found simply defining GGML_CUDA_ENABLE_UNIFIED_MEMORY, whether 1, 0, or "", was a 10x hit to 2 t/s. Perhaps because the laptop's RAM is already so over-committed there. But with the much smaller 4B Qwen3.5-4B-Q8_0.gguf, it doubled performance from 20 to 40+ t/s! Tnx! (an old Quadro RTX 3000 rather than an iGPU)
- tmaly 6mo agoThat is pretty solid, I have a 2070 with 8GB VRAM and 64GB RAM, but I haven't run too much. I regret not getting a 3090 back when I built this machine.
- mncharity 6mo agoNod. Mine was VR dev leftovers. Fwiw, running 6ish prompts in parallel, roughly doubles my aggregate t/s (but requires cooling kludgery). If one's goal is not local, but rather real-time or consistent or transparent or scalable, there's AWS.