19 ms·
It pretty much comes down to 2 factors which is memory bandwidth and compute. You need a high enough memory bandwidth to be able to "feed" the compute and you n
by MeImCounting 3y ago
It pretty much comes down to 2 factors which is memory bandwidth and compute. You need a high enough memory bandwidth to be able to "feed" the compute and you need beefy enough compute to be able to keep up with the data that is being fed in by the memory. In theory a single Nvidia 4090 would be able to run a 70b model with quantization at "useable" speeds. The reason mac hardware is so capable in AI is because of the unified architecture meaning the memory is shared across the GPU and CPU. There are other factors but it essentially comes down to tokens per second advantages. You could run one of these models on an old GPU with low memory bandwidth just fine but your tokens per second would be far too slow for what most people consider "useable" and the quantization necessary might star noticeably effecting the quality.
- int_19h 3y agoA single RTX 4090 can run at most 34b models with 4-bit quantization. You'd need 2-bit for 70b, and at that point quality plummets. Compute is actually not that big of a deal once generation is ongoing, compared to memory bandwidth. But the initial prompt processing can easily be an order of magnitude slower on CPU, so for large prompts (which would be the case for code completion), acceleration is necessary.
- MeImCounting 3y agoThats a good point. For example both the RTX 4090 and the RTX 6000 Ada Generation use the AD102 chip. The RTX 6000 Ada though, would be able to run 70b models due to the larger memory pool despite having the same memory interface width.
- reddit_clone 3y agoWould the not-yet-released M3 Mac Mini with upgraded RAM be enough to do some beginner level LLM work?
- int_19h 3y agoThose supposedly will go up to 24Gb RAM, which is only enough for inference for ~34b quantized models. At which point you might as well get an RTX 3090 for a PC that you probably already have, and get a better bang for the buck. It also really depends on what you consider "beginner level". Fine-tuning is really easy these days and many people do it, but you really want CUDA for that.