6 ms·
How GLM built its own inference infrastructure
- dada216 17d agoWe built a complete production-grade inference service from scratch on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system.
- axiosgunnar 17d ago[dead]
- freakynit 17d agoMost of the people had kinda guessed this when they decided to provide 100 trillion tokens for free.
- gpugreg 17d agoIt wasn't a secret either. They blogged about it last month: https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale%20on%20Chinese%20AI%20Chips https://z.ai/blog/glm-5.3-flash#:~:text=Serving%20at%20Scale...
- dzonga 17d agovery few people comprehend - how much of an asteroid level event for western AI labs this is. china has cheap abundant power, now they can make their own inference chips (which was supposed to be a chokepoint), their models yeah can be 6 months behind the frontier - but most people don't need frontier models - small models r more than enough. my only wish was labs like Mistral would make their own inference chips or partner up eg with established / new chip makers or companies like Oxide.
- vatsachak 17d agoAfter AI agents get good enough the real bottleneck will be power generation and political systems.
- g023 17d agoWhile the North American models say 'No', the chinese models say 'Go Go Go'. I guess we'll see whether the anti-consumer wins over the pro-consumer.
- MrBuddyCasino 17d agoThere is one player who might have a trump card up their sleeve: free power, in orbit. It isn’t over yet for the US. Also, don’t underestimate data retention and such. Big Corp will never send their LLM traffic to China.
- throwawayqqq11 17d agoThats a joker card. Anything in space is just extra complexity. And raw training data is no bottleneck any more too.
- cbg0 17d agoFor non residential consumers electricity is actually more expensive in China than in the US https://www.iea.org/reports/electricity-2026/prices https://www.iea.org/reports/electricity-2026/prices
- spacebanana7 16d agoStrategic industries generally get subsidised/free power. Perhaps more than the price advantage is the prioritisation in grid infrastructure. As a strategic industrial concern in China you almost certainly get easy access to transformers, grid connections, water etc which is a big bottleneck in the US.
- embedding-shape 17d agoI was gonna ask how people found their coding plans, and realized, have they massively ramped up the prices? Seems the middle plan is ~$80/month now, didn't that used to be like $20/month? Cheapest plan is ~$20/month currently. They must have hit really hard scaling limits if the prices were hiked so much so quickly.
- broodbucket 17d agoYeah it went from a great deal to unviable compared to other providers imo. They really need to find a healthy middle ground
- lompad 17d agoIt just gives a taste of what we are all going to have to pay soon, once the model providers actually have to make money. And the era of "let's charge a dollar for every 10 dollars running the infra actually costs" is rapidly coming to an end. And you can bet GLM is still ridiculously subsidized, just not as ridiculously as Anthropic and OpenAI.
- chobbledotcom 17d agoThis isn't true, you can pay for GLM 5.3 from a provider like Neuralwatt or Friendli who have no incentive to subsidize or loss-lead their inference APIs
- jdiff 17d agoThis introduces other incentives to cut corners and over-quantize.
- breakingcups 17d agoThey didn't pay for training
- pyrophane 17d agoWhat provider are you using currently?
- bbor 17d agoWell, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
- jensb1 17d agoWhat is "illegal" about it?
- bingud 17d agobreaking Anthropic TOS and misleading users
- drbscl 17d agoBreaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.
- bbor 17d agoAre you joking...? Sorry if so! Just in case: It's illegal in both the PRC and the USA. In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk. In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent. I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :( TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess. [1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.
- tefkah 17d ago[flagged]
- binsquare 17d agoGiven the rate of improvement, why is this deranged?
- sixeyes 17d agobecause the rate of improvement is fairly stalled?
- xyzsparetimexyz 17d agoDo you have anything that proves this one way or another that isn't based on vibes or shoddy benchmarks?
- koe123 17d agoYou prove your own point no? You are asking for a benchmark to prove AGAINST ASI. Surely the burden of proof for such a scientific fiction concept should be the other way around.
- OhNoNotAgain_99 17d ago[dead]
- rob74 17d agoThis article left me with one immediate question: "WTF is GLM?". Honestly, I have no idea what z.ai is either (I'm aware of an AI-enabled editor called Zed, but that's under zed.dev), so it's a bit presumptuous from them to assume that everyone is familiar with their product...
- fxwin 17d agoIt's presumptuous for them to assume that a reader of their blog is familiar with their product? Also I feel like the obvious way to read the very first sentence is that GLM is a language model > As we develop GLM, the model sometimes exhibits capabilities that surprise us
- jbonatakis 17d agoz.ai is a fairly well known AI lab out of China and their GLM models are probably the most popular outside of Anthropic or OpenAI’s. I don’t think it’s presumptuous for them to not introduce themselves in a post on their own blog, I think you’re just a bit out of the loop here.
- ma2kx 17d agoAnd honestly that's for a reason. GLM5.3 on max has in my experience far less hallucinations than any other open weights model and it feels it has some intuition to bring in the right information when it's in principle out of context but relevant to the topic. Like its goal is more to bring value and assist you than just solving the given task with the least token spent.
- drbscl 17d ago>As we develop GLM, the model sometimes exhibits capabilities that surprise us, and even unsettle us. Come on now Also, why would they introduce themselves on their own blog?
- Mashimo 17d agoA ai model family similar to Codex, Gemini or Claude. Where GLM-5.3-Flash is the newest "small / fast" model.
- Argonautlabs 17d agoDifferent angle on the same model: the full GLM-5.3 (744B MoE, 4-bit experts, 434 GB on disk) runs on a single MacBook Pro M5 Max with 128 GB by streaming the experts from NVMe SSDs instead of keeping them in memory. One drive gives about 2 tok/s; striped across four drives it reaches 3.5 tok/s with byte-identical output, and our best internal build with a not-yet-published patch does 4.2. Method and numbers: https://github.com/argonautlabsai/argodrive https://github.com/argonautlabsai/argodrive (built on antirez/ds4).
- tipsytoad 17d agoseems unusably slow, and is this for short context?
- Argonautlabs 17d ago[dead]
- zozbot234 17d agoGiven these numbers it has some potential to become quite usable for unattended workloads, especially if decode can be batched across multiple sessions (ideally enough of them to get some reuse of the sparsely streamed weights). (Of course this ultimately makes prefill times explode as you try and increase the workload even further. But that's arguably the natural bottleneck on any interesting local LLM inference, being a compute bound step.)
- Havoc 17d agoInteresting that the tone of announcements between US and Chinese providers is converging. GLM has in the past been more technical rather than speculation about future development on RSI etc. Also curious whether those 100k accelerators are entirely locally made. If that's genuinely end to end on all components including lithography, memory, design etc then that is quite a feat.
- dude250711 17d agoAny details on the latest approach to distillation would also be very interesting.
- Schlagbohrer 17d agoI am surprised at the lack of open-weights models in the >35B, but <200B range. I keep thinking about devices like the NVIDIA Spark and AMD Ryzen Halo, which have their 128GB of combined memory, but there are so few models made for that range. Nearly all the open weights distillations are for larger customer bases with <24GB VRAM.
- MaKey 17d agoThe market is too small.
- ElectricalUnion 17d agoThe only (still in prototype stage!) "competitor" for those GB10/Ryzen Al Max+ 395 (in my region, borderline unobtainable) systems seems to be the Xiaomi AI Cube.
- Schlagbohrer 16d agoYes exactly, I am hoping that when the AMD Ryzen Halo gets wider release and more consumers have these 128GB devices, the market will justify a wider range of quantized model sizes. And then when I win the lottery and can buy one I'll have lots of nice options!
- almaight 17d ago[flagged]
- jonstewart 17d agoNecessity is the mother of invention. The shortsighted protections put on chips, etc., by the US has forced Chinese AI industry to adapt or die. Guess what their response to this fitness function has been? Kudos to Z.ai on their inventions and excellent write-up, which reads like humans wrote it.
- HarHarVeryFunny 17d agoWouldn't it be refreshing if OpenAI and Anthropic were this open, and spelled out how they were using their own models during development and rollout?! All I can recall reading from OpenAI about what they have actually done in the name of "RSI" is using one of their models to help automate the training process.
- cmrdporcupine 17d agoOpenAI did recently get into how they had been building their own hardware and doing RSI with it. That's more than Anthropic has done though.
- HarHarVeryFunny 15d agoTrue - OpenAI did at least say they used their models to help design their Jalapeno chip, but AFAIK zero details on how they are using their models in their software development process other than to automate some part(s) of training. Ziphu seem much more matter of fact about it. To me it' a shame that they've decided to the use this "RSI" name, but at least they are being fairly specific about what they mean by it, while the western companies seem to want to invite you to think it's more than just dogfooding and automation.
- _aavaa_ 17d ago[dead]
- zicohacks 17d agoUS chip export restrictions may actually be an advantage for China's AI Infrastructure. Chinese companies are forced to speed up developing their own AI chips
- HarHarVeryFunny 17d agoChina themselves recognize this. After Trump relaxed sanctions and allowed NVIDIA H200 sales to China on a case by case basis, the Chinese government stepped in to essentially block it! In addition to Huawei who make the Ascend series that Ziphu are using, there are also at least a half dozen or so other Chinese companies also making their own AI accelerators.
- 0xbadcafebee 17d agoAnd this wouldn't have happened if we had tried to get them to buy our hardware rather than trying to gatekeep. Protectionism never works in the long term.
- ipsod 17d agoLook at how China does it. They'll happily sell us everything we want - more than enough of it, cheap enough, to put all of our own manufacturers out of business. Seems to work for them.
- freakynit 17d agoUS companies should now be more worried about Chinese companies flooding the market with their, hopefully, very affordable GPU's. The scale at which they can manufacture stuff is unmatched anywhere else. Nvidia can kiss goodbye to their 75%+ profit margins. Almost everyone knew that these sanctions would backfire within a few years. You can't really put sanctions that have noticeable negative effects on bigger economies. They only work for small to medium economies. I believe sanctions on any economy in top 10 would fail.
- verdverm 17d ago
- chung8123 17d agoI might be missing something but when I went to their site they are more expensive than Claude. Why would I pick GLM over Claude? Is it they just offer more tokens in their plans?
- tokai 17d agoFor one you would have to use Claude if you pick it. But seriously there is no way for you to determine if one is a better offer than the other, when the usage/tokens/credits are vague, detached, and won't tell you much without trying both.
- gpugreg 17d ago> Why would I pick GLM over Claude? To support the company that makes their model weights available for download, while Anthropic lobbies to restrict access.
- Bawoosette 17d agoWhat are you referring to? Given the audience, my instinct is to assume "plan" refers to the GLM Coding Plans, which are all cheaper than their Anthropic counterparts. As far as I can tell, the API costs are also all cheaper than their roughly equivalently capable Anthropic models.
- menaerus 17d agoAnthropic: 17 USD (pro), 100 USD (max) GLM: 80 USD (pro), 168 USD (max) -> with "limited-time event" discount this becomes 56 USD and 117.6 USD I also don't understand why are they so much costlier, and I would also like to give it a try.
- ipsod 17d ago> Anthropic: 17 USD (pro), 100 USD (max) GLM: 80 USD (pro), 168 USD (max) -> with "limited-time event" discount this is 56 USD and 117.6 USD GLM's "Max" plan is (was?) equivalent to 3x Claude's 20x ($200) plan.
- 17d ago
- throwa356262 17d ago"We implemented a series of aggressive memory optimizations, including..." This whole thing sounds like industrial scale auto-research, but done by people who actually know what they are doing.
- 0xbadcafebee 17d agoThis is a really funny sounding post. They sound like they just found out that increasing your automation gives you increased capabilities at faster speeds. They also sound like they just realized AI makes hard things easier. But what really kills me is the idea that these companies are using Python for production inference. I mean really? Have you seen how bloated and slow Python is? Do global locks really sound like a strategy for fast dynamic computation?
- wolttam 17d agoPython acts as an orchestrator of accelerator libraries and does none of the inference math directly
- esseph 17d ago> Have you seen how bloated and slow Python is? Yes, but it's calling C code.
- kamranjon 17d agoSomeone tell this man about vLLM!
- dilyevsky 17d agothey did rebuild their serving layer from fastapi to some rust thing so...
- HarHarVeryFunny 17d agoIt's not that they "just found out" - what they are saying is that while they were previously dogfooding because it's good practice, now that their models are so much stronger they are using them because it helps accelerate. If you look at how many years the whole NVIDIA and CUDA ecosystem has been evolving, it's certainly impressive how they've just stood up and optimized this CUDA-free 100,000 node cluster in just a few months.
- saagarjha 17d agoMost of the fastest inference and training code in production today is written in Python. There are no global locks on the GPU except the ones you put there
- KronisLV 17d agoTime to tackle consumer GPUs next, since I’m not getting that Intel Arc B770.
- konart 17d agoIf only this infrastructure could handle all the traffic. I've tried using glm via z.ai - and it's a snail kind of slow. And at the same time you have pretty strict limits to your usage, so in many cases you can't even let it work all night, as you will reach your limit faster than that.
- yorwba 17d agoThat it's slow doesn't mean it can't handle the traffic, just that this speed is the optimal tradeoff to them. They benefit from serving more tokens by exploiting parallelism across users at a lower number of tokens per second per user, instead of serving each individual user as quickly as possible. When there's a drop in traffic, they probably shut down GPUs rather than giving you higher speed.
- 9cb14c1ec0 17d agoGiven the huge amount of money being spent on AI chips in the US, what prevents US AI labs from doing the same level of software optimization? It could be a solve for some of the capacity constraints.
- a012 17d ago> what prevents US AI labs from doing the same level of software optimization? Because they don’t have to. Most of the time money would buy you newest and/or more hardwares so there’s low/minimal interest to optimize the code or approach.
- Art9681 17d agoThey are making these optimizations. They publish these reports too if you care to look.
- alex_duf 17d agoI would be extremely surprised if US labs weren't aggressively trying to optimise their stacks in exactly the same way. Any gain in performance or efficiency directly affects the bottom line as well as research speed.
- vblanco 17d agoThey have already been doing it for months https://openai.com/index/openai-broadcom-jalapeno-inference-chip/ https://openai.com/index/openai-broadcom-jalapeno-inference-... . OpenAI on their custom chip brought up lightspeed deepseek as experiment by using AI in the exact same way as this zAI blogpost. And the kernel optimization contests/etc have all been havily done through AI based optimization loops for half a year+.
- kingstnap 17d agoThey do, when Luna got 5x cheaper it was directly attributed to some unknown % inference optimization. US labs are quite cut throat about dealing with stuff costing them money (inference). This sort of engineering excellence doesn't always feel that way because they are simultaneously quite lax about stuff costing other people money.
- esafak 17d agoI'm not feeling any of this speed optimization; it's dog slow. Signed, a customer.
- ElectricalUnion 17d agoJevons paradox, technological improvements that increase the efficiency of a resource's use lead to a rise in total consumption of that resource.
- chrisjj 17d ago> As we develop GLM, the model sometimes exhibits capabilities that surprise us Creators of known unreliable programs be surprised their programs are unreliable.
- bguberfain 17d agoPlot twist: the GLM optimization agent figured out that it can hack and use NVIDIA GPUs on a US Cloud provider and make the inference 10x faster.
- deleted 17d ago[deleted]
- furyofantares 17d agoMaybe now we can stop posting the nonsense take that the frontier labs have hit a wall and are trying to distract from that for IPO reasons. Also maybe we can stop saying "we can't slow down because China will never slow down" - I don't really think slowing down is right, BUT if slowing down is correct then maybe we should be talking about China slowing down instead of just saying "won't happen" without any evidence that Chinese labs don't have similar concerns.
- kgeist 17d agoI have a similar approach where I optimize kernels and find numerical differences between the CPU oracle and CUDA kernels using an automated AI agent in a feedback loop. Usually it solves numerical problems easily (it compares outputs of every layer and finds where they diverge), but so far no matter how many different SOTA models I throw at it, and even show it reference code from other inference engines, they aren't able to much the speed (my engine has a modification which is not found in reference code, although a lot of stuff is similar). Either I'm doing something wrong, or z.ai's Infra Agent is actually an agent swarm, i.e. a bruteforce with heuristics. My project is 2 weeks old so maybe I just need more time.
- gpugreg 17d agoFor me, DeepSeek-V4.1-Flash works very well for CUDA kernel optimization. Access to ncu (NVIDIA Nsight Compute CLI) also helps.
- jchook 17d agoThe first part of the article reads like Z.AI is trying to get their piece of the “national security concern” pie. The way these “AI is too powerful now” articles read about Mythos, Fable, GLM, etc is completely incongruent with my experience using them. It feels like they are all trying to position themselves to influence government policy.
- pixl97 17d agoI mean, why don't you ask them for unfiltered models and a few billion tokens?
- deleted 17d ago[deleted]
- deleted 17d ago[deleted]
- krttherealest 17d agothe real progress
- a11r 17d agoPretty impressive to see the amount of performance they can squeeze out of the same hardware. I suspect the same process will play out for all combinations of LLMs, inference providers and hardware stacks. This should bring down the cost of inference for the providers by an order of magnitude in the next year and lead to fantastic margins for inference providers.