5 ms·
4 bit quants should require 85GB VRAM, so this will fit nicely on 4x 24G consumer GPUs, plus some leftover for KV cache optimization.
by deoxykev 2y ago
4 bit quants should require 85GB VRAM, so this will fit nicely on 4x 24G consumer GPUs, plus some leftover for KV cache optimization.
- hedgehog 2y agoI've found the 2 bit quant of Mixtral 8x7B is usable for some purposes with an 8GB GPU. I'm curious how this new model will work in similar cheap 8-16GB GPU configurations.
- cjbprime 2y agoWouldn't expect that to work at all.
- hedgehog 2y agoOllama (which wraps llama.cpp) supports splitting a model across devices so you get some acceleration even on models too big to fit entirely in GPU memory.
- reissbaker 2y ago16GB will be way too small unfortunately — this has over 3x the param count, so at best you're looking at a 24GB card with extreme 2bit quantization. Really though if you're just looking to run models personally and not finetune (which requires monstrous amounts of VRAM), Macs are the way to go for this kind of mega model: Macs have unified memory between the GPU and CPU, and you can buy them with a lot of RAM. It'll be cheaper than trying to buy enough GPU VRAM. A Mac Studio with 192GB unified RAM is under $6k — two A6000s will run you over $9k and still only give you 96GB VRAM (and God help you if you try to build the equivalent system out of 4090s or A100s/H100s). Or just rent the GPU time as needed from cloud providers like RunPod, although that may or may not be what you're looking for.
- dannyw 2y agoYou can QLoRA decent models on 24GB VRAM. There’s also optimised kernels like Unsloth that are really VRAM efficient and good for hobbyists.
- reissbaker 2y agoYes, but I still don't think you'll be able to run Mixtral 8x22b with 16GB VRAM, or QLoRA it, even with Unsloth. It's much bigger than the original Mixtral.
- timschmidt 2y agoReasonably priced Epyc systems with up to 12 memory channels and support for several TB of system memory are now available. Used datacenter hardware is even less expensive. They are on par with the memory bandwidth available to any one of the CPU, GPU, or NPU in the highest end Macs, but capable of driving MUCH more memory. And much simpler to run Linux or Windows on.
- hmottestad 2y agoDo you have any feel for the performance compared to the M3 Max?
- Manabu-eo 2y agoLLM inference is mostly memory bound. An 12-channel Epyc Genoa with 4800MT/s DDR5 ram clocks at 460.8 GB/sec. It's more than the 400GB/s of the M3 Max, only part of that accessible for the CPU. And on the Epyc System you can plug much more memory for when you need larger memory and PCI-E gpus, for when you need less faster memory. Threadripper PRO is only 8-channel, but with memory overclocking it might reach numbers similar to those too.
- hedgehog 2y agoI'm curious how the newer consumer Ryzens might fare. With LPDDR5X they have >100 GB/s memory bandwidth and the GPUs have been improved quite a bit (16 TFLOPS FP16 nominal in the 780M). There are likely all kinds of software problems but setting that aside the perf/$ and perf/watt might be decent.
- cjbprime 2y agoConsumer Ryzens only have two-channel memory controllers. Two dual-rank (double sided) DIMMs per channel, which you would need to use to get enough RAM for LLMs, drops the memory bandwidth dramatically -- almost all the way back down to DDR4 speeds.
- aydyn 2y agoAFAIK, 2-bit quant leads to too much loss of performance, such that you're better off using a different smaller model altogether. See here: https://www.reddit.com/r/LocalLLaMA/comments/18ituzh/mixtral_update_on_perplexity_testing_adding_an/ https://www.reddit.com/r/LocalLLaMA/comments/18ituzh/mixtral...
- qeternity 2y ago4bit should take up less than this, there are quite a few shared parameters between experts. But unless you’re running bs=1 it will be painful vs 8x GPU as you’re almost certain to be activating most/all of the experts in a batch.