5 ms·
>$40k gets you almost-Opus GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference (so it's closer to $400k than $40k). They suggest
by kgeist 3mo ago
>$40k gets you almost-Opus
GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference (so it's closer to $400k than $40k).
They suggest using this modified model:
>A REAP-pruned (≈22% of experts removed), Int8-mix NVFP4 quantized version of GLM-5.2, ≈594B parameters.
I wonder how it behaves in practice outside of benchmarks. Qwen3.6, even at 6-bit quantization, often gets stuck in loops while reasoning. And here they've also removed some experts. I mean, sometimes an 8-bit or 16-bit small model can be smarter than a lobotomized large model. I heard the consensus is you shouldn't go below 8 bit for coding.
Also, it's not clear what is left of the available context when you try to fit a lobotomized model into 4 RTX 6000s. Anything below 100k is barely usable because it often hits compaction before it's able to gather the necessary context
P.S. found in the repos, 240k context
- amelius 3mo agoHow does this work with scaling? I assume you can then somehow run several hundreds of prompts concurrently?
- CamperBob2 3mo agoYou can get 1M context with the lukealonso NVFP4 quant on 8x RTX6000s, which remains coherent and useful through at least 400k. No real need to run 8x H200s unless you just want to. Or unless you need to serve many concurrent users or agents on a regular basis.
- rsync 3mo ago"GLM 5.2 is "almost Opus," and it needs at least 8xH200s for comfortable inference ..." What is the behavior if one were to run GLM 5.2 with only a single H200 ? Would it fail to run at all, or would it just run so slowly as to be unusable ? I would like to prove out the build, and concept, of a SOTA model locally, but then backfill the rest of the GPUs in 18-24 months when they cost significantly less ...
- BoorishBears 3mo ago> in 18-24 months when they cost significantly less ... going to need you to sit down for this one...
- sanderjd 3mo agoSay more. My expectation is that the current gen of gpus will start being replaced by the next gen, and then it may be possible to get used ones that are still within their useful life at lower prices. My expectation is also that memory vendors are likely to increase production, which will drive those prices down eventually. Maybe not over the next 18-24 months though.
- scheme271 3mo agoGiven the prices and shortages, I'd think people would keep and use the current gen stuff till it drops dead. It may not be as good but it's paid for and given the prices for next gen stuff, it's probably worth using for another cycle or two.
- BoorishBears 3mo agoThe only thing that diminishes the value of a GPU right now is unsupported features with outsized value during inference and/or training (like FP4 support) and it takes time for those features to actually take off And labs are fully leaning into pricing for intelligence, so their margins are improving very quickly (which allows them to pay even more for existing compute) I'd be shocked if current prices aren't the bottom for the next 18-24 months.
- bradfa 3mo agoMany newer Chinese lab models are releasing with int4 native weights. Latest NVIDIA generation GPUs have a hard time with this and can actually be slower than previous generations. This may make Blackwell depreciate faster than other recent generations.
- BoorishBears 3mo ago
- Der_Einzige 3mo agoLooping, like most other phenomenons related to LLMs, is a sampling problem and can be easily solved with the DRY penalty. It’s in llamacpp. The same guy who wrote heretic invented the SOTA antilooping and diversification strategies.