7 ms·
LLM inference is GPU bound and VRAM bound. Given quantization however, 2x3090 or 4090 (48GB) is enough VRAM to load 65B quantized llama derivatives (30-40GB). W
by crasm 3y ago
LLM inference is GPU bound and VRAM bound. Given quantization however, 2x3090 or 4090 (48GB) is enough VRAM to load 65B quantized llama derivatives (30-40GB). With exllama you can get maybe 20-30 tok/s with dual 4090s on ~1200W. With an M2 Ultra maybe 10 tok/s on ~178W.
(I don't have these devices but have been researching LLM inference performance in the interest of buying a machine for them.)
Advantage goes to M2 Ultra if you ever might need more than 48GB VRAM. I think it's unlikely nvidia is going to release a significantly higher VRAM consumer card anytime soon, since that's what their A100s are for.
- mort96 3y agoIf only the big ML stuff didn't all use CUDA. Guessing that world won't transition to Metal compute any time soon.