5 ms·
for those who are already in the field and doing these things - if I wanted to start running my own local LLM.. should I find an Nvidia 5080 GPU for my current
by iamtheworstdev 1y ago
for those who are already in the field and doing these things - if I wanted to start running my own local LLM.. should I find an Nvidia 5080 GPU for my current desktop or is it worth trying one of these Framework AMD desktops?
- wmf 1y agoIf you think the future is small models (27B) get Nvidia; if you think larger models (70-120B) are worth it then you need AMD or Apple.
- yencabulator 1y agoI wonder how much MoE will disrupt this. qwen3:30b-a3b is pretty good even on pure CPU, but a lot smarter than a 3B parameter model. If the CPU-GPU bottleneck isn't too tight, a large model might be able to sustainably cache the currently active experts in GPU RAM.
- hengheng 1y agoThe recent qwen3 models run fine on CPU + GPU, and so does gpt-oss. LM Studio and Ollama are turnkey solutions where the user has to know nothing about memory management. But finding benchmarks for these hybrid setups is astonishingly difficult. I keep thinking that the bottleneck has to be CPU RAM, and for a large model the difference would be minor. For example with an 100 GByte model such as quantised gpt-oss-120B, I imagine that going from 10G to 24G would scale up my tk/s like 1/90 -> 1/76, so 20% advantage? But I can't find much on the high-level scaling math. People seem to either create calculators that oversimplify, or they seem too deep into the weeds. I'd like a new anandtech please.
- whizzter 1y agoDoesn't matter, people will always find ways to eat RAM despite finding more clever ways to do things.
- yencabulator 1y agoMoE eats the same amount of RAM but accesses less of it.
- sliken 1y agoMore accurately, same amount of ram but accessed in a cache friendly manner with greater locality.
- loudmax 1y agoThe short answer is that the best value is a used RTX 3090 (the long answer being, naturally, it depends). Most of the time, the bottleneck for running LLMs on consumer grade equipment is memory and memory bandwidth. A 3090 has 24GB of VRAM, while a 5080 only has 16GB of VRAM. For models that can fit inside 16GB of VRAM, the 5080 will certainly be faster than the 3090, but the 3090 can run models that simply won't fit on a 5080. You can offload part of the model onto the CPU and system RAM, but running a model on a desktop CPU is an enormous drag, even when only partially offloaded. Obviously an RTX 5090 with 32GB of VRAM is even better, but they cost around $2000, if you can find one. What's interesting about this Strix Halo system is that it has 128GB of RAM that is accessible (or mostly accessible) to the CPU/GPU/APU. This means that you can run much larger models on this system than you possibly could on a 3090, or even a 5090. The performance tests tend to show that the Strix Halo's memory bandwidth is a significant bottleneck though. This system might be the most affordable way of running 100GB+ models, but it won't be fast.
- cpburns2009 1y agoJust a point of clarification. I believe the 128GB Strix Halo can only allocate up to 96GB of RAM to the GPU.
- geerlingguy 1y ago108 GB or so under Linux. The BIOS allows pre-allocating 96 GB max, and I'm not sure if that's the maximum for Windows, but under Linux, you can use `amdttm.pages_limit` and `amdttm.page_pool_size` [1] [1] https://www.jeffgeerling.com/blog/2025/increasing-vram-allocation-on-amd-ai-apus-under-linux?fjs https://www.jeffgeerling.com/blog/2025/increasing-vram-alloc...
- amstan 1y agoI have been doing a couple of tests with pytorch allocations, it let me go as high as 120GB [1] (assuming the allocations were small enough) without crashing. The main limitation was mostly remaining system memory: htpc@htpc:~% free -h total used free shared buff/cache available Mem: 125Gi 123Gi 920Mi 66Mi 1.6Gi 1.4Gi Swap: 19Gi 4.0Ki 19Gi [1] https://bpa.st/LZZQ https://bpa.st/LZZQ