5 ms·
So close! My machine with 192GB RAM + RTX 3090 24GB can almost run this. It says it needs 24GB of VRAM and 256GB of RAM for MoE offloading. https://unsloth.ai/
by xrd 4mo ago
So close! My machine with 192GB RAM + RTX 3090 24GB can almost run this. It says it needs 24GB of VRAM and 256GB of RAM for MoE offloading.
https://unsloth.ai/docs/models/glm-5.2#usage-guide https://unsloth.ai/docs/models/glm-5.2#usage-guide
In a prior thread, someone said it would take $500k in hardware:
https://news.ycombinator.com/item?id=48629970 https://news.ycombinator.com/item?id=48629970
- mgambati 4mo agoWith 2 wouldn’t have good results. Ideal range for coding is at least Q8.
- kibibu 4mo agoAccording to this very article, 4-bit dynamic is essentially lossless
- Aurornis 4mo agoWatch out. Those claims are often made based on KL-divergence over some arbitrary corpus, not performance in the real world or benchmarks. I’ve found that I need to go a couple steps past whatever quantizations are good enough in the KL-divergence testing to get good performance in real tasks with long context. So when Q4 is claimed to be lossless I end up with Q5 or Q6 for actual long-context tasks.
- cheema33 4mo agoI have the RAM, but not the VRAM. What kind of speed/tps could you expect from a 3090 with 24GBs of RAM? I am somewhat tempted to pick a GPU with 24GBs of RAM.
- phamilton 4mo agoGeneration is basically just memory bandwidth math. Each token has to read all the active weights. I think that's around 40B parameters active. At a 4-bit quant that's 20GB. With 100GB/s (replace with whatever your bandwidth is) and you get 5 tokens per second.
- ekidd 4mo agoA GPU with 24GBs of RAM is mostly useful for running a very carefully squeezed Qwen3.6 27B (4-bit Unsloth quants, 8-bit K/V cache, possibly MTP, 128k context). This is a fun little model that's smart enough to do debugging, refactoring, and implementing "clean" specs that don't force it to make complicated design choices. I've seen it rip through a 9-year-old Terraform AWS config, and (without using the network) correctly identify nearly everything that would need to be upgraded or migrated for modern AWS. But if I give it some poorly conceived spec with lurking design headaches, then it goes on an endless thinking binge and ultimately fails. Speed-wise, I don't have numbers, but it feels subjectively faster than Opus in Claude Code. YMMV. Once you go above "a used 3090 at a decentish price", then I strongly recommend renting cloud GPUs or at least testing models using paid APIs. This allows testing your use case before spending piles of money.
- elliotbnvl 4mo ago$500k is a vast overestimation. For massive concurrency at FP8 or even BF16 maybe. NVFP4 at reasonable speeds (~120 tok/s) and concurrency is possible at a $80/90k figure with today's prices, maybe even less. That buys you 6 RTX 6000 PRO Blackwells, a decent CPU and motherboard, power supply. 576gb of VRAM. You could do it for under $50k if you're OK with 40 tok/s decode, ~1200 tok/s prefill.
- __m 4mo agoHow fast will the hardware become outdated? Are there big improvements expected in the next 3 years?
- easygenes 4mo agoM5 Ultra will ship before end of year, likely. Though with current RAM shortage, likely max spec will be 256GB and in short supply. In late 2027 or early 2028, Nvidia will release Vera Rubin DGX Spark, likely with double or better the performance of current Blackwell, though unclear if memory capacity will go up much from current 128GB. Two to four of those will run models like this decently. In 2028 we should expect Vera Rubin RTX discrete lineup, including the replacement to the RTX PRO 6000. Likely memory spec will be minimum 128GB. Good chance of up to 200GB. Two to four of those will run NVFP4 models in this class very well.
- deleted 4mo ago[deleted]
- jiqiren 4mo agoI hope all this speculation comes true. Right now this ram crunch is ridiculous.
- hajile 4mo agoIt might be M6 Ultra and I think the real reason for stopping selling top-tier units was to avoid mid-generation price hikes and increasing demand for the more expensive next-gen systems that I assume will come with 512gb (maybe 1TB) of RAM and a massive markup to match.
- uberex 4mo agoFunny I casually asked Gemini and it said 500k for unquantized with decent throughput.
- stymaar 4mo agoThis is why you shouldn't believe uncritically an answer from an LLM (neither should you do for any answer from a human either though).
- andy_ppp 4mo agoBut I did my research online and the sun cycle is every 11 years and something something global warming is a hoax every single year now.
- j45 4mo agoLLMs aren't discrete calcluators or estimators of things unless framed and guided to do so.
- uberex 4mo agoGood job I didn't use a vanilla LLM without tool use harness then.
- colinsane 4mo agoi asked gemini and it replied with "Error: 400 Your prompt was blocked by safety filters. Please revise and try again."
- digitaltrees 4mo agoI asked and it said “403 forbidden - careful peon attempts to bypass the late stage capitalism api with your monetary offerings in exchange for you daily tokens will get you perma banned right to jail”.
- matheusmoreira 4mo agoSafety from competition!
- ijidak 4mo agoCrossing my fingers that this boom jumpstarts 90's like improvements in computing hardware. I feel like part of the reason for the relative stagnation in hardware over the last twenty years was simply the lack of use cases to justify hardware refreshes by businesses. Most of the money and energy went to mobile for the last fifteen years. Affordable local inference might be the gravy train the server, desktop, and laptop manufacturers need to get back in gear.
- linzhangrun 4mo agoPhysical limitation of the manufacturing process may be more significant factor, starting from the TSMC 10nm ten years ago
- gruez 4mo ago>I feel like part of the reason for the relative stagnation in hardware over the last twenty years was simply the lack of use cases to justify hardware refreshes by businesses. No, we're running into limits of moore's law, and it's showing in prices for new nodes, where they're getting denser but not cheaper.
- horsawlarway 4mo agoIt's true we hit limits, but I feel like a lot of it was "limits" in the sense that the tradeoff stopped being worth the cost, so we optimized in other areas. So we hit limits on clock speed in the early 2000s (ex - the 4ghz wall) but it also turned out that mobile as the driver for sales meant no one really cared much about clock speed compared to performance/watt. Clock speed mattered, but only relative to how many watts it took to get it (and above 4ghz... too many watts). But we've seen a 15x improvement over the last 20 years. Performance/Watt is WAY up. My guess is that LLMs are going to drive another "improvement cycle" in areas that we didn't care much about before. I've built about 10 personal desktop machines (1 every ~4 years) and I can honestly say that I didn't care much about memory bandwidth prior to 2021. In the same way that I didn't care much about how many watts my pentium 4 was using in 2005. But now... now I care a lot about memory bandwidth. I care about memory speeds and total system ram in a manner I really, really didn't before. So I think we're going to see a big shift to machines built on unified ram with a crazy focus on squeezing memory bandwidth and total ram capacity as far as we can. My bet is that we'll get a similar 10-15x improvement by 2040 in unified system ram designs. I fully expect to see 2tb unified ram desktops and 200gb unified ram phones be relatively common on a 20 year timeline, assuming we see similar levels of geopolitical stability (ex - world war 3 throws a wrench into things).
- bbor 4mo agoI’m kinda lost here… do y’all really have machines in your houses with hundreds of gigs of RAM?? Am I just behind the times? The page advertises the 8-bit quant as taking ~800GB, which seems like it would require at least 3 consumer motherboards fully stacked w/ 4x64GB cards each. Maybe “locally” has slowly come to imply “…on your homelab”?
- cpburns2009 4mo agoRAM wasn't expensive even a year ago. I maxed out a used Dell Precision T5610 with 128 GB DDR3 for $250 in 2021.
- numpad0 4mo agoDRAM prices at mid-2025 rates were ~$2.5/GB for DDR5, and ~$1.5/GB for DDR4. "Hundreds of gigs" of RAM used to be under $500. 128GB of cheapest RAM used to be like $200. It seemed to go over heads for a lot of people that you could get hypothetical future machines on CS/CE textbooks were attainable for that little, for some reason - there seemed to be some fixation on the idea that 16GB is all you need.
- Gracana 4mo agoYou don't have to have a server, workstation motherboards support lots of memory channels. I was lucky to buy a lot of RAM before prices skyrocketed. I knew I wanted to play with this stuff, so I spent what felt like a lot of money at the time to buy 8x96GB DDR5-6400 RDIMMs. Now the same RAM costs at least 6x more.
- woodrowbarlow 4mo ago[dead]
- nijave 4mo agoI got a 2U rackmount with 192Gi DDR4 for $1.1k USD in 2023. Around 1.5 yrs ago, server RAM could be had pretty cheap--especially slower LRDIMMs (I wanna say 512Gi DDR4 was <$500 USD). I checked a couple old ServeTheHome threads and seeing maybe around $50/32GB RDIMM although thought it was cheaper than that for a little while