3 ms·
tough ask, but since we're here: has anyone done this with 16GB of VRAM? I've been getting projects finished with LM Studio, but it definitely could stand to be
by catapart 4mo ago
tough ask, but since we're here: has anyone done this with 16GB of VRAM? I've been getting projects finished with LM Studio, but it definitely could stand to be more efficient. lots of time wasted with trying to get models to understand a problem with so few tokens.
- Rzor 4mo agoRX 9060 XT 16GB here on google/gemma-4-26b-a4b-qat using LM Studio. Context 65k, 23 layers on the GPU, 7 on the CPU, model in memory, mmapped. I'm getting 23-33 tks. Started experimenting 3 days ago (with gemma-4-e4b), don't know what half those settings mean, but 26B, even quantified, feels significantly better at a few small projects I asked it to create ("create a image converter using ffmpeg in bash", "create a canvas animation with real physics, no libraries"[1]). It's faster than I can read, but it feels slow as hell. I think 40-50 tks is probably much more comfortable and I hope I can reach that when trying this on llamacpp soon enough. [0] - https://pastes.io/9gaARxE8 https://pastes.io/9gaARxE8 [1] - https://jsfiddle.net/pou4nbh9/1/ https://jsfiddle.net/pou4nbh9/1/ Model: https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gg...