6 ms·
GLM-5.3 Artificial Analysis Benchmarks
- Escapade5160 1mo agoSol is an underappreciated model. Dropped Claude today and went to codex. None of that god awful prose Claude used for me any longer.
- swingboy 1mo agoDoes Artificial Analysis use OpenRouter for model access to do their benchmarks?
- colingauvin 1mo ago...do I take out a double mortgage to buy a 4 Spark cluster?
- lisplist 1mo ago$20k is personal loan territory, not a second mortgage lol
- colingauvin 1mo agoNot with my credit!
- nvme0n1p1 1mo agoNo, you use openrouter and spend 10% as much as using a proprietary model.
- jtbaker 1mo agoQwen3.8 27B doing a lot of lifting right now, and people seem to run it pretty well on 1-2x 3090 setups...
- killingtime74 1mo ago20k is credit card territory
- markasoftware 1mo agoVery impressive score for the size, though token use is higher than k3 and far higher than proprietary models, and its price to performance isn't all that far ahead of k3 as a result
- Havoc 1mo ago>token use is higher than k3 and far higher than proprietary models GLM sets effort to max by default historically.
- markasoftware 1mo agoAa also benchmarked k3 at max
- Zaheer 1mo agoIs it worth using these models if I have a claude code subscription already? The appeal of lower cost is nice but I haven't gotten over the switching cost yet.
- colingauvin 1mo agoAt least by API usage, they aren't yet lower cost than subscriptions. Not sure about GLM's subscription plans though.
- glub 1mo agoGLM subscription is better than API, but significantly worse than Codex, even when used outside peak hours.
- culi 1mo agoUse a unified proxy that lets you switch between models seamlessly. We are far from an equilibrium in this market and you will continue to have FOMO no matter who you pick if you go all in on one company
- notatoad 1mo agono, at subscription prices claude is a better value than GLM. They're only a better value if you're paying API rates
- scotttrinh 1mo agoI like to compare models with a similar score on cost per task and output tokens per task since those measure two things I'm interested in: cost efficiency and token efficiency. Here's how GLM-5.3 compares to other models in a similar score and against GLM-5.2 to save a few clicks for others who care about these metrics: Model Score Cost / Task Output Tokens / Task ------------------------------------------------------------------------- GLM-5.3 (max) 59.5 $0.68 41,107 GLM-5.2 (max) 53.0 $0.56 32,200 Claude Opus 5 (high) 61.5 $1.52 21,353 GPT-5.6 Sol (max) 60.9 $1.23 16,879 Grok 4.6 (high) 60.9 $0.84 21,735 Kimi K3 (max) 59.7 $0.84 25,474 GPT-5.6 Sol (xhigh) 59.0 $0.87 11,098 Claude Opus 5 (medium) 58.6 $0.98 12,459 Qwen3.8 Max 58.1 $1.13 38,287 Qwen3.8 2.4T A95B 57.7 $0.95 32,472 Claude Opus 4.8 (max) 57.3 $1.65 33,557 GPT-5.6 Sol (high) 57.3 $0.52 7,545 Muse Spark 1.2 (xhigh) 56.8 $0.40 30,430 GPT-5.6 Terra (max) 56.6 $0.51 20,838 GPT-5.5 (xhigh) 56.3 $0.69 16,893 Gemini 3.7 Flash (high) 56.0 $0.40 36,847 Edited for accuracy and more models.
- sourcecodeplz 1mo agoMuse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.
- sscaryterry 1mo agoI found the sweetspot here: GPT-5.6 Sol (high) 57.3 $0.52 7,545 (Edit: TLDR; It gets on with it, makes the same mistakes you would, without overthinking and overengineering, most of the time)
- glub 1mo ago
- BinRoo 1mo agoBeware of the benchmarks listed. SciCode and EnterpriseOps for instance: https://shukla.io/blog/2026-08/gym.html https://shukla.io/blog/2026-08/gym.html
- Onavo 1mo agoThe Chinese models also like to cut corners on stuff like science. Their scores on stuff like biotech and scientific knowledge is far from ChatGPT unfortunately. (Claude is pretty good but it just refuses all prompts).
- colingauvin 1mo agoTied for #1 by agentic index (with Opus 5).
- fenestella 1mo ago[flagged]
- glub 1mo agoI've tested GLM 5.3 on the release day and Artificial Analysis is spot on. It's a really good model. But my main takeaway was something else. I've used closed weight models for long enough that I've forgotten how good it feels to see reasoning tokens. With GPT/Claude, you kind of hope that intent was captured well, that agent had all the information, all the tools it needed, because you won't see "hmmm it seems like nix flake isn't available here and I shouldn't install something globally" until it slopped out millions of tokens and wasted hundreds of dollars for 8 hours. With GLM and the likes, you just stop the disease right where it begins.
- Havoc 1mo agoYes, not necessary often but being able to stop something that is going off the rails is super useful. Especially if the root cause is prompt ambiguity - inject a clarification & it recovers
- glub 1mo agoIt's also starting to go beyond reasoning and it's becoming much more problematic. Reasoning is one thing, but codex, for example now encrypts agent-to-agent messages as well, and compaction. I've no idea what subagents are instructed to do, or what they reported back in native codex. The only thing that's keeping me is the value $200 subscription provides. If that value disappears, I see no reason why not to switch to something that isn't a black box.
- tw1984 1mo agoWith GPT/Claude, hiding those from users to waste their tokens is a feature, not a limitation.
- aitchnyu 1mo agoGenerally, are closed sourced models hiding their traces? I was making an agent to develop and deploy apps and fed the traces to dispel time-consuming detours and made it a few times faster.
- scosman 1mo agoAnd reminder: it's less than a quarter the size of Kimi K3!
- kzrdude 1mo agoThis comment by the Z AI lead is relevant, about parameter vs data scaling: https://news.ycombinator.com/item?id=49357405 https://news.ycombinator.com/item?id=49357405
- AnodicElegy 1mo agoI understand that running these benchmarks can get expensive, but it would be really nice to see AA include more benchmarks of models at reasoning settings other than the maximum, at least for the biggest releases. They have that nice graph of cost vs. composite benchmark score with the Pareto frontier line, but who knows if those are actually the optimal choices? There are already a few non-max-reasoning models on the Pareto line, among the few that were tested.
- apitman 1mo agoYou can turn on various levels of some of many of the models in the UI
- AnodicElegy 1mo agoYes, they have multiple levels of Claude, GPT, Gemini, and Kimi, but not the other top models (I would put GLM, Qwen, Muse, Grok, and Deepseek in that bucket).
- qqt 1mo ago[flagged]
- yipinwong 1mo agoStill yet, I cannot justify switching from dirt-cheap Luna model, which is pretty damn "intelligent" and works well for my flow
- gdorsi 1mo agoTo me one of the biggest limitations of GLM is the lack of multi-modality. For web dev is just a must to have, and offloading that part to a secondary model doesn't work really well in my experience.
- ushiro35 1mo ago[flagged]
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- deleted 1mo ago[deleted]
- lluisantoni 1mo ago[flagged]
- sumedh 1mo agoI ran the same Mac SVG drawing prompt through GLM 5.2 and 5.3 across every reasoning effort level, and 5.3 showed improved performance https://sumedh.info/models/glm-5-3 https://sumedh.info/models/glm-5-3