3 ms·
I am not too familiar with LLMs and GPUs (Not a gamer either). But want to learn. Could you please expand on what else would be capable of running such models
by reddit_clone 3y ago
I am not too familiar with LLMs and GPUs (Not a gamer either). But want to learn.
Could you please expand on what else would be capable of running such models locally?
How about a linux laptop/desktop with specific hardware configuration?
- MeImCounting 3y agoIt pretty much comes down to 2 factors which is memory bandwidth and compute. You need a high enough memory bandwidth to be able to "feed" the compute and you need beefy enough compute to be able to keep up with the data that is being fed in by the memory. In theory a single Nvidia 4090 would be able to run a 70b model with quantization at "useable" speeds. The reason mac hardware is so capable in AI is because of the unified architecture meaning the memory is shared across the GPU and CPU. There are other factors but it essentially comes down to tokens per second advantages. You could run one of these models on an old GPU with low memory bandwidth just fine but your tokens per second would be far too slow for what most people consider "useable" and the quantization necessary might star noticeably effecting the quality.
- int_19h 3y agoA single RTX 4090 can run at most 34b models with 4-bit quantization. You'd need 2-bit for 70b, and at that point quality plummets. Compute is actually not that big of a deal once generation is ongoing, compared to memory bandwidth. But the initial prompt processing can easily be an order of magnitude slower on CPU, so for large prompts (which would be the case for code completion), acceleration is necessary.
- MeImCounting 3y agoThats a good point. For example both the RTX 4090 and the RTX 6000 Ada Generation use the AD102 chip. The RTX 6000 Ada though, would be able to run 70b models due to the larger memory pool despite having the same memory interface width.
- reddit_clone 3y agoWould the not-yet-released M3 Mac Mini with upgraded RAM be enough to do some beginner level LLM work?
- int_19h 3y agoThose supposedly will go up to 24Gb RAM, which is only enough for inference for ~34b quantized models. At which point you might as well get an RTX 3090 for a PC that you probably already have, and get a better bang for the buck. It also really depends on what you consider "beginner level". Fine-tuning is really easy these days and many people do it, but you really want CUDA for that.
- tinco 3y agoThere are no other laptops on the market with as much VRAM as the 64GB MBP's have as far as I know. You could make a Linux desktop computer with two 3090's, linked together giving 48GB of VRAM. Wich apparently can run a 4-bit quantized 6k context 70B llama model. People are recommending Macbooks because they're a relatively cheap and easy way to get a very large amount of RAM hooked up to your accelerator. Note that these are quantized versions of the model, so they're not as good as the original 70B model, though people claim their performance is really close to original performance. To run without quantization you'd need about 140GB of VRAM. Which would only be possible with an NVidia H100 (don't know the price) or two A100's (at $18,000 each).