4 ms·
What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?
by kcb 6mo ago
What benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?
- blizdiddy 6mo agoIs it so hard to project out a couple product cycles? Computers get better. We’ve gone from $50k workstation to commodity hardware before several times
- kcb 6mo agoSubscription services get all the same benefits from computer hardware getting better. But actually due to scale, batching, resource utilization, they'll always be able to take more advantage of that.
- deminature 6mo agoIntel has just released a high VRAM card which allows you to have 128GB of VRAM for $4k. The prices are dropping rapidly. The local models aren't adapted to work on this setup yet, so performance is disappointing. But highly capable local models are becoming increasingly realistic. https://www.youtube.com/watch?v=RcIWhm16ouQ https://www.youtube.com/watch?v=RcIWhm16ouQ
- kcb 6mo agoThat's 4 32GB GPUs with 600GB/s bw each. This model is not running on that scale GPUs. I think something like 96GB RTX PRO 6000 Blackwells would be the minimum to run a model of this size with performance in the range of subscription models.
- acchow 6mo ago> I think something like 96GB RTX PRO 6000 Blackwells would be the minimum to run a model of this size with performance in the range of subscription models. GLM 5.1 has 754B parameters tho. And you still need RAM for context too. You'll want much more than 96GB ram.
- marcus_holmes 6mo agoWhy would anyone need more than 640Kb of memory?
- kcb 6mo agoExactly the point though. In the 640KB days there was no subscription to ever increasing compute resources as an alternative.
- marcus_holmes 6mo agoWell, there kinda was - most computing then was done on mainframes. Personal / Micro computers were seen as a hobby or toy that didn't need any "serious" amounts of memory. And then they ate the world and mainframes became sidelined into a specific niche only used by large institutions because legacy. I can totally see the same happening here; on-device LLMs are a toy, and then they eat the world and everyone has their own personal LLM running on their own device and the cloud LLMs are a niche used by large institutions.
- kcb 6mo agoThe difference is computers post text terminal are latency and throughput dependent to the user. LLMs are not particularly.
- marcus_holmes 6mo agoSorry, I don't understand that comment. Can you clarify, please?
- kcb 6mo agoMy point is LLMs aren't more usable if the hardware is in your room versus a few states away. Personal computers still to this day aren't great when the hardware is fully remote.
- 6mo ago
- CamperBob2 6mo agoIt will run exactly the same tomorrow, and the next day, and the day after that, and 10 years from now. It will be just as smart as the day you downloaded the weights. It won't stop working, exhaust your token quota, or get any worse. That's a valuable guarantee. So valuable, in fact, that you won't get it from Anthropic, OpenAI, or Google at any price.
- kcb 6mo agoThat's why we all still use our e machines its never obsolete PCs. Works just the same it did 20 years ago, though probably not because I've never heard of hardware that's guaranteed not to fail.
- fwipsy 6mo agoAgree directionally but you don't need $50k. $5k is plenty, $2-3k arguably the sweet spot.
- unlikelytomato 6mo agoas a local LLM novice, do you have any recommended reading to bootstrap me on selecting hardware? It has been quite confusing bring a latecomer to this game. Googling yields me a lot of outdated info.
- fwipsy 6mo agoFirst answer: If you haven't, give it a shot on whatever you already have. MoE models like Qwen3 and GPT-OSS are good on low-end hardware. My RTX 4060 can run qwen3:30b at a comfortable reading pace even though 2/3 of it spills over into system RAM. Even on an 8-year-old tiny PC with 32gb it's still usable. Second answer: ask an AI, but prices have risen dramatically since their training cutoff, so be sure to get them to check current prices. Third answer: I'm not an expert by a long shot, but I like building my own PCs. If I were to upgrade, I would buy one of these: Framework desktop with 128gb for $3k or mainboard-only for $2700 (could just swap it into my gaming PC.) Or any other Strix Halo (ryzen AI 385 and above) mini PC with 64/96/128gb; more is better of course. Most integrated GPUs are constrained by memory bandwidth. Strix Halo has a wider memory bus and so it's a good way to get lots of high-bandwidth shared system/video RAM for relatively cheap. 380=40%; 385=80%; 395=100% GPU power. I was also considering doing a much hackier build with 2x Tesla P100s (16gb HBM2 each for about $90 each) in a precision 5820 (cheap with lots of space and power for GPUs.) Total about $500 for 32gb HBM2+32gb system RAM but it's all 10-year-old used parts, need to DIY fan setup for the GPUs, and software support is very spotty. Definitely a tinker project; here there be dragons.
- terbo 6mo agoAgree on the framework, last week you could get a strix halo for $2700 shipped now it's over $3500, find a deal on a NVME and the framework with the noctua is probably going to be the quietest, some of them are pretty loud and hot. I run qwen 122b with Claude code and nanoclaw, it's pretty decent but this stuff is nowhere prime time ready, but super fun to tinker with. I have to keep updating drivers and see speed increases and stability being worked on. I can even run much larger models with llama.cpp (--fit on) like qwen 397b and I suppose any larger model like GLM, it's slow but smart.