5 ms·
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost
- zero0529 1mo agoHaving used GLM-5.3 I honestly don't think it is better than 5.1. It is slower and the result is often overengineered, it is if it overthinks everything.
- ed-is-ai 1mo ago[flagged]
- CamperBob2 1mo agoThe GLM-5.3 weights are not yet open, and they've said that the delay is due to the need to nerf them for "safety." So I have a feeling a lot of these early claims are not going to pan out in the long run.
- sehw 1mo agoopen-source when?
- ed-is-ai 1mo agohttps://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench
- scottfits 1mo agomy prediction is even if open source Chinese models are 90% as good (or even a bit better, which I don’t really believe because of benchmark hacking) enterprises will still pay for Claude / ChatGPT and the harness, integrations, and peace of mind versus using some Chinese cloud.
- sschueller 1mo agoEnterprise's peace of mind is being able to use the model and not have the US government decide on a whim to block access. Additionally I may want to run attack simulations which requires the removal of safeguards. My only option is to use an open model I can run on my own hardware.
- eightysixfour 1mo agoYou think the US can’t, on a whim, decide US companies can’t use Chinese models? They already showed exactly how they would do it - designate it a supply chain risk and say anyone using it can’t be a provider to the government.
- sschueller 1mo agoHow exactly is that going to be enforced? Anyway I am not actually talking about US based companies.
- eightysixfour 1mo ago> How exactly is that going to be enforced? How is anything enforced on B2G agreements? Contractually & legally, which turns into internal policy, which shuffles the risk on to the rogue dev deciding to use GLM instead of the mandated Grok subscription. This is literally the playbook they ran for Claude. I know folks who work for government contractors who were immediately going through the evals to get rid of Anthropic because it became a risk for them.
- hyperpape 1mo ago
- jchw 1mo agoAlmost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated? I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM. "Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5... This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.
- rfgplk 1mo ago> Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated? The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.
- finaard 1mo agoInteresting - when Kimi K2.6 came out I switched over from Anthropic models, with at that time comparable to better results for me. I was using Anthropic via API, heavier months were roughly $400 worth of Anthropic tokens - I can get the same thing done via a $100 ollama subscription.
- sambusa_123 1mo agoJust read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/ https://reinvently.co.uk/tools/ed-o-meter/tests/ Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...
- nylonstrung 1mo agoI think most of them would have actually failed, it's only recently that models were any good at using tool calls and harnesses after they started post-training for that
- zuzululu 1mo ago
- ac29 1mo agoNot sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last place
- svachalek 1mo agoFable got heavily beaten down by its refusals, which is not too surprising; although a couple of problems got refused for reasons I can't even imagine and the page doesn't quote the refusal. Some of the other failures like the colicky baby one are also probably soft refusals, it's not clear what the grading criteria are but I'm guessing it got docked for not going anywhere near a possible diagnosis.
- rfgplk 1mo agoThey loosened the refusals recently. They aren't anywhere close to the Sol refusals.
- poincareball 1mo agoI've been not just unimpressed by Fable, but actively find it to generate negative value. It hallucinates more, and in more destructive ways, than other models I've worked with and generates truly atrocious jargon and bizarre inhuman explanations that end up cluttering things. The code it writes is terrible too. Overly complex with a lot of technical debt.
- rfgplk 1mo ago> I've been not just unimpressed by Fable, but actively find it to generate negative value. Fable is good in a few very specific domains (graphics programming) but otherwise it's an overhyped model. Far too expensive too. Opus 5 is outright better in every metric.
- poincareball 1mo agoI honestly find Opus 5 a downgrade too. It's not as bad as fable, but nearly so, and is very argumentative. It'll straight up ignore directives and design decisions.
- svachalek 1mo agoNice to see the TTFT chart, wish aggregators like OpenRouter would track this. Matches my experience, the Deepseek models while fast overall can have a horrendous wait before they start responding, and Claude models are superbly responsive. It's particularly annoying that models like flash and luna, where you've explicitly chosen speed over quality, can still stall out before they even get started.
- iamcoder18 1mo agoThere's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.
- zuzululu 1mo agobenchmark is saturated but your claim that gemini-3.6-flash is better than 5.6-sol is also not trustworthy or accurate. maybe if you mentioned 3.7-flash it might have been slightly more believable (but still false).
- MostlyStable 1mo agoHe was remarking on how unbelievable it was. He sentence is a little bit hard to parse but he's saying that both 5.5 being better than 5.6 AND Gemini 3.6 scoring better than both are signs that this benchmark is not very useful. The reason for both of those things is, as you point out, the benchmark is very obviously saturated
- zuzululu 1mo agoah fair enough , 5.5 being better than 5.6 is just plain wrong and yeah makes everyone skeptical
- Klaster_1 1mo agoFor the last week, I've been heavily immersed reverse engineering a device with help of GLM-5.3 and it surpassed all my expectations - I actually managed to achieve very way more than I thought I would. I never worked on such low level stuff, it would have taken me months to learn ARM assembly and how to find for and write exploits. Initially, I attempted this with Claude, but it blocked me on the very first message, so I got a refund and decided to try z.ai. The only downsides are that it's maybe a bit slower than my day job Opus and I had to pay ~200 EUR for a monthly plan in order not to bump into weekly limits in a couple of days. If this level of capability cost maybe 50 EUR, I'd strongly consider getting a long time subscription.
- matheusmoreira 1mo ago> z.ai > The only downsides are The catch's in their revolting terms of service.
- HighGoldstein 1mo agoCan you elaborate?
- 3asgf 1mo agoThey are pretty broad and vague: https://chat.z.ai/legal-agreement/terms-of-service https://chat.z.ai/legal-agreement/terms-of-service I don't even know if I could use code generated by them, because they claim the copyright. Better don't travel to Singapore (wise anyway because someone could slip drugs into your suitcase) or China if you use them.
- matheusmoreira 1mo agoObnoxiously broad and perpetual license over inputs and outputs. Yes, Z.ai demands an unconditional, irrevocable, transferable, sublicensable, perpetual, worldwide license to use, modify, reproduce, adapt, publish, perform, distribute, and create derivative works from your prompts and outputs. The same license extends to your username and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Prohibitions on "disturbing" or "inappropriate" content, whatever that is. Professional use prohibitions. Discussing Z.ai is prohibited to the point even my posting this comment is against their terms of service. And they can of course ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan you won't ever see that money ever again. Even the US companies aren't this bad.
- _joel 1mo agoSorry, just can't read that page, too AI spammy
- gosolozero 1mo ago2 articles on how G 5.3 is the best in the top 10? Seems like a bit of astroturfing going on
- ed-is-ai 1mo agoI had to google astroturfing. I liked the due diligence... I am a hacker news infant, so all my history is about this. Give it a proper look. The results all there, and code if you want to run the evals (or make your own - it's very easy) https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench You might have read about Ox Alpha aka glm5.3-flash, well I updated to include that
- CMay 1mo ago> if you run one model, run glm-5.3 That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing. Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case. Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.
- ed-is-ai 1mo agoThe point I'm making is that most models are good enough for most tasks, so choose on speed/cost. Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark. My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench Encourage everyone to eval like the devil
- silverwind 1mo agoThose benchmarks don't tell much, they only check if a problem was solved, not how. Also there's surely a lot of benchmaxxing going on in the model training.
- gertlabs 1mo agoOne problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class. We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier. Data at https://gertlabs.com/rankings https://gertlabs.com/rankings
- angoragoats 1mo agoThe writing in the first three sentences is so bad that I closed the page. Please, bloggers, write with your own voice. Don’t let an LLM do it for you.
- ed-is-ai 1mo agoNoted. Feedback received
- hellohello2 1mo agoThis whole thing immediately reads as Claude generated, making it hard to take seriously. Why do these results contradict existing serious attempts at benchmarking LLMs? Namely: https://artificialanalysis.ai/ https://artificialanalysis.ai/ https://arena.ai/leaderboard/agent https://arena.ai/leaderboard/agent
- koe123 1mo agoHaha theres even an em-dash in the title
- ckocagil 1mo agoBecause it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.
- irishcoffee 1mo agoThe whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.
- hellohello2 1mo agoYes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.
- Tepix 1mo ago
- visiondude 1mo agoso the coding tasks illicit refusal by anthropic classifiers and instead of updating tasks to do similar things that don’t trigger classifiers their choice is to count those as fails? feels wrong given their stated task list.
- Youden 1mo agoI think that's reasonable. The goal of the benchmark is to determine how the model performs on real-world tasks. If real-world tasks trigger Anthropic's classifiers, that's a failure on Anthropic's part. I feel that choosing tasks that don't trip the classifier would also be a form of bias towards Anthropic. In case it's not clear, the coding tasks are really benign things; there's nothing security-, health- or biology- related in there. The classifier being tripped is definitely unreasonable.
- ed-is-ai 1mo agoIndeed; I find it really interesting there is so much disbelief and outright hostility to my post, in these comments. I am an AI Realists rather than AI hype-artist, say it as it is. For me, the evidence is clear that for our own use cases OpenAI, Anthropic models should not be the default choice any longer. Incidentally, if you want to look at this in more depth - my previous post about the eval framework got like zero response; that's where the results, methodology, code for running the evals, evals are all shared. Anyone can run this to verify it for themselves https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench
- spiderfarmer 1mo agoAnd now there's 0x Alpha, which is their (now free) new model.
- ed-is-ai 1mo agoThanks for point out - I updated the benchmark today to include. It is every bit as good as everyone says it is. The intelligence / $ is something else...
- dainiusse 1mo agoIs the beater in the room with us now?
- rfgplk 1mo agoIn their current form open weight models are simply not worth running. Literally the amortized cost of hardware + electricity you need to operate them is >> than the cost of paying for subscriptions.
- pcwelder 1mo agoThe whole thing (article, benchmark) is a slop soup. - Metaphor overload. We get it, it's just like car racing. Show some mercy on human readers. - X, not Y - A, never B - Tasteless em dashes - Hallucinated data, like model add date. The standings are so unbelievable that they border on laughable.
- ed-is-ai 1mo agopcwelder -https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench all the data is here if you care to look But thanks for the feedback
- couAUIA 1mo agohaiku 4.5 top 7, that benchmark is absolute crap
- gilesvangruisen 1mo agoAll of the answers. None of the understanding.
- hereme888 1mo agoChinese propaganda. Both current top articles on HN are shilling for GLM-5.3
- deleted 1mo ago[deleted]
- ed-is-ai 1mo agoIt's good to check the facts. Time will tell, I'm not the only one saying these things. But you look at the raw data and test to your hearts content if you care enough - all results and the code use to run it is above. And you can add your own tests if this isn't enough for you... Policy of total transparency https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench
- Arcuru 1mo agoAnecdotal, but from my personal usage I found GLM-5.3 was not as capable at performing autonomous tasks as Opus/Sol. Certainly competitive with the Sonnet/Terra level, but not with the Frontier. I got their lowest subscription tier and burned a week of quota on trialling it.
- jacobgold 1mo agoThe entirety of the coding evals are just 7 trivial coding tasks in Python? This is a joke.
- ed-is-ai 1mo agoThanks for the feedback. It was limited to 7 to keep things balanced. On the premise most people don't just code. The 7 were an example of my realworld use cases - the point of this is to encourage people to run their own benchmarks and not just take what they read as gospel All the code and results are here for anyone who wants to delve, see if they can reproduce. That is the point of discourse https://github.com/ed-is-ai/featherbench https://github.com/ed-is-ai/featherbench
- amazingamazing 1mo agoHopefully all of these models lead to manufacturing breakthrough so we can bring down prices of cards.
- deleted 1mo ago[deleted]
- solenoid0937 1mo agoIDK I use open models every day for personal projects, and closed models for work. Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me. I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.
- tamimio 1mo agoWhat’s the best model for planning and architecture design, rather than solving problems in codes or an issue?
- qwertox 1mo agoI can't wait for Mistral to host these models. For those who don't know yet, Mistral is pivoting to also offer Chinese models in their own cloud.
- ed-is-ai 1mo agoThat's pretty cool. All available already in OpenCode if you're a developer, see for yourself!
- deleted 1mo ago[deleted]
- Aeroi 1mo agoai;dr can't take any generated benchmark seriously. if you produce actual results, then produce actual copy to go with it.
- nylonstrung 1mo agoThe stealth model "Ox Alpha" has been crushing benchmarks and appears to be the next release in the GLM family
- pulkitsh1234 1mo ago>Short version: if you run one model, run glm-5.3 — 100% pass, a 9.3 rubric, $0.28 for the lap, about a fifth of gpt-5.5's cost ... I mean wtf is this prose ? what rubric ? what lap ? I can definitely say this is Opus 5.
- vatsachak 1mo ago> reinvently Slop
- ed-is-ai 1mo agoImpolite
- garn810 1mo agoAmerican people are fed propaganda that Chinese are bad people with bad tech etc In fact, it's often superior Looks at cars, BYD, Geely etc... Looks at neural networks benchmarks...