3 ms·
Looking for advice from someone who knows about the space - Suppose I'm willing to go out and buy a card for this purpose, what's a modestly priced graphics car
by morcus 2y ago
Looking for advice from someone who knows about the space - Suppose I'm willing to go out and buy a card for this purpose, what's a modestly priced graphics card with which I can get somewhat usable results running local LLM?
- estreeper 2y agoThis is a tough question to answer, because it depends a lot on what you want to do! One way to approach it may be to look at what models you want to run and check the amount of VRAM they need. A back-of-the-napkin method taken from here[0] is: VRAM (GB) = 1.2 * number of parameters (in billions) * bits per parameter / 8 The 1.2 is just an estimation factor to account for the VRAM needed for things that aren't model parameters. Because quantization is often nearly free in terms of output quality, you should usually look for quantized versions. For example, Llama 3.2 uses 16-bit parameters but has a 4-bit quantized version, and looking at the formula above you can see that will allow you to run a 4x larger model. Having enough VRAM will allow you to run a model, but performance is dependent on a lot of other factors. For a much deeper dive into how all of this works along with price/dollar recommendations (though from last year!), Tim Dettmers wrote this excellent article: https://timdettmers.com/2023/01/30/which-gpu-for-deep-learning/ https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni... Worth mentioning for the benefit of those who don't want to buy a GPU: there are also models which have been converted to run on CPU. [0] https://blog.runpod.io/understanding-vram-and-how-much-your-llm-needs/ https://blog.runpod.io/understanding-vram-and-how-much-your-...
- morcus 2y agoThank you for the background!
- loudmax 2y agoThe bottleneck for running LLMs on consumer grade equipment is the amount of VRAM your GPU has. VRAM is RAM that's physically built into the unit and it has much higher memory bandwidth than regular system RAM. Obviously, newer GPUs will run faster than older GPUs, but you need more VRAM to be able to run larger models. A small LLM that fits into an RTX 4060's 8GB of VRAM will run faster there than it would on an older RTX 3090. But the 3090 has 24GB of VRAM, so it can run larger LLMs that the 4060 simply can't handle. Llama.cpp can split your LLM onto multiple GPUs, or split part of it onto the CPU using system RAM, though that last option is much much slower. The more of the model you can fit into VRAM, the better. The Apple M-series macbooks have unified memory, so the GPU has higher bandwidth access to system RAM than would be available over a PCIe card. They're not as powerful as Nvidia GPUs, but they're a reasonable option for running larger LLMs. It's also worth considering AMD and Intel GPUs, but most of the development in the ML space is happening on Nvidia's CUDA architecture, so bleeding edge stuff tends to be Nvidia first and other architectures later, if at all.
- morcus 2y agoThank you for the response!