17 ms·
GLM 5.2 beats Claude in our benchmarks
- kordlessagain 3mo agoYou can launch GLM-5.2 in Opencode using Nemesis8: https://github.com/DeepBlueDynamics/nemesis8#nemesis-8 https://github.com/DeepBlueDynamics/nemesis8#nemesis-8 After installing, do a `n8 build` to build the image, then `n8 --danger --provider opencode interactive` to launch it in a container. Signup for GLM-5.2 here: https://z.ai https://z.ai
- sanid 3mo agoOne can also try https://neuralwatt.com https://neuralwatt.com using it in opencode. I think they give $5 trail credits to test with any of the open weight models.
- MaKey 3mo agoInitially, I was confused where to find their open weight model offering. It's here: https://portal.neuralwatt.com https://portal.neuralwatt.com
- generichuman 3mo agoYou can use GLM in OpenCode with a z.ai subscription by default as well. Also it'd be good if you mentioned you were involved with nemesis8.
- kordlessagain 3mo agoI think it would be good not to suggest someone run a new Chinese agent on their bare metal. When I posted the comment I was both the first commentor as well as the first person to upvote the submission. That matters. My name is ALSO on the open source repo that allows Opencode to be run in a container. That's transparency, maybe not here, but on a clickthrough to Github it is immediately obvioius.
- wadim 3mo ago> I think it would be good not to suggest someone run a new Chinese agent on their bare metal. Not sure a project nobody knows or uses is much better in this regard?
- kordlessagain 3mo ago[flagged]
- solenoid0937 3mo agoGLM export controls incoming? I predict Commerce will force OpenRouter, HuggingFace to take some open models down within the next few months. Not that it would make any sense.
- gruez 3mo ago>GLM export controls incoming? US imposing export restrictions on a model from China?
- manquer 3mo agoWhile unlikely , it is not without precedent , there are restrictions on ASML a Dutch company to sell EUV machines
- verdverm 3mo agoASML complies as an ally, why would China comply? The weights are already available and downloaded, is it going to be a crime to have them, run them, make them available? Constitutional rights still exist (I hope)
- solenoid0937 3mo ago> is it going to be a crime to have them, run them, make them available? Now you're getting it! Commerce will call it a munition and those harboring it as harboring illegal/foreign munitions. No business will take the hit, so they will quickly deplatform the models. No end user has the GPU capacity to use GLM 5.2 or similar models at full precision so the government will call the problem "mostly solved." But they might choose to "make examples" out of a few people using p2p software to download the weights if they choose to.
- verdverm 3mo agoOr we use the models to work on fixing vulns and stop over-blowing the doom scenarios. Gotta save the kids and kill the terrorists though! I'm for making software better instead of banning it based on what the rich and powerful claim. I suspect the real fear is that open weight models undermine the financials and token prices they thought were going to pay off their ludicrous spending because they have all raced and raised hardware prices.
- veselin 3mo agoHere, it appears they compare a single prompt "find IDOR", against a multi-agent system. However, one can also start far more sophisticated skills that spin up subagents and mostly do the same in Claude Code, Codex, OpenCode, Pi, etc. Which I guess makes what semgrep sells obsolete. Unless they have built a pareto-optimal point in terms of capabilities and token usage maybe?
- blazespin 3mo agoI think the point is less "how can we throw shade on the OP" and more "a harness can enable a lot of models to do very serious cybersec, glm 5.2 is one of them"
- s3p 3mo agoAre you replying to a response to the original comment? I looked but i didn't see anyone saying he's throwing shade.
- BikiniPrince 3mo agoYou have to forgive the GLM bot. It's not very good.
- danslo 3mo agoIt reads like an ad. Secondly these are "just" IDORs, arguably the easiest class of vulnerabilities. Thirdly it compares to GPT 5.5 and Opus 4.8. No, we don't have Mythos at home.
- vlian2088 3mo ago>Thirdly it compares to GPT 5.5 mythos is <10% ahead of gpt 5.5 on all benchmarks, which it gains by being several times the size of opus. had it been economical to provide, it would've been released to the public on day one instead of the marketing circus those effective altruism clowns had exhibited. admitting that it costs >1000% to run inference on a <10% better model would've been very damning.
- oa335 3mo ago> it costs >1000% to run inference do you have a source for this claim? i thought LLM providers earn high margins from inference (charged by token). is this no longer the case?
- 3836293648 3mo agoThis was just theorised. The leaked OpenAI financials suggest otherwise (because of shady naming of losses) The only ones who seem to profit are the ones running smaller Chinese models. Even NVIDIA seems to have to "reinvest" their profits into sponsoring companies to buy their cards now.
- vlian2088 3mo agoif a $6000000 cabinet can generate 10000/s tokens of Opus but only 1000/s tokens of Mythos, then Mythos costs 1000% to run no matter the markup. no one has a source, because no one knows closed model parameter counts. we have only heuristics which strongly indicate that Mythos is simply a big fucking model that any other lab could make an equivalent of.
- InsideOutSanta 3mo agoIn my experience, GLM 5.2 is extremely good at finding vulnerabilities, and more importantly, unlike Opus, I've never seen it refuse a command. It genuinely is a very strong model for finding and fixing vulnerabilities.
- admax88qqq 3mo ago> beats Claude in our Cyber Benchmarks Beats which model in Claude? Whenever a "benchmark" doesn't put precise model numbers in their headlines I am immediately skeptical. Either they don't know the difference (bad) or they are benchmarking against weaker models (misleading, also bad). It's like when studies say "AI is bad at X" and they used GPT-3.5 in current year.
- ls612 3mo agoOpus 4.8 according to TFA. Whether or not the safety guardrails were responsible for the difference is an open question but for a dev who wants to secure their software who doesn’t work at one of the blessed Glasswing companies it doesn’t really matter why, it matters what the best tool you actually have is.
- InsideOutSanta 3mo agoThey say "Claude Opus 4.8" in the first paragraph.
- crm9125 3mo agoWe're supposed to read the article? How are we supposed to stay skeptical of everything if we read anything!?
- simplyluke 3mo agoAnthropic's own models perform differently under the same version depending on how much they've decided to quietly downgrade them.
- himata4113 3mo agoThese numbers are seem pretty low compared to what I was able to achieve specifically around windows kernel, win32k<->win32u to be exact. It honestly wouldn't surprise me anymore if china started surpassing models that US makes public, at least in specific categories such as cyber. GLM 5.2 is already capable enough to assist in self-training which is similar to what we saw happen with frontier models and they appear to be getting there at a significantly lower cost than openai/anthropic.
- danmaz74 3mo agoIt will almost for sure surpass the models which Trump will allow US "allies" (which he just considers client states) to use. This, together with China's growing dominance in PV, rechargeable batteries, EV, could really be the nail in the coffin for the post WWII economic world order.
- himata4113 3mo agoHonestly, it's becoming increasily hard to disagree with such sentiment when china is preparing itself to lead in energy, manufacturing, research, chip production and so on while there's an entire group of people trying to put datacenters in space.
- woeirua 3mo agoYou are delusional if you think China is going to let Europe have access to Mythos level models for free.
- WithinReason 3mo ago> [...] beating Claude Code (32%) at roughly $0.17 per vulnerability found Claude Code is an agent harness, not an LLM. Claude is a brand (or group of LLMs), not an LLM.
- Onavo 3mo agoClaude code it's the only way to get access to the actual amortized cost of running a Claude-scale model. The consumer non-enterprise API is extremely expensive (with increasing marginal costs for the user and fat profit margins for Anthropic). If you want to approximate a State level attacker's cost where they can have the model on their own hardware, Claude Code is probably the best guess at the amortized cost.
- raincole 3mo agoYes, and the article author is fully aware of that. Thank you for pointing out this small mistake though.
- mkagenius 3mo agoIt looks like the author is specifically avoiding model's name, because results are really weird. Opus 4.8/4.7 scored 28% Opus 4.6 score 37% So the author thought as let's not get into that just write Claude.
- insiderphd 3mo agoHello! Author here (Katie) Ty for your comments, 4.6 and 4.7 both scored 28% on our benchmark, I just wanted to have 10 things in the list because I wanted a round number.
- andriy_koval 3mo agomany people think opus 4.6 was the best
- happycube 3mo agoNot weird at all, given the variance in Opus' quality over the last few months. wild guess - I wouldn't be surprised if Opus 4.6 was run quantized for a while, and 4.7/4.8 have QAT for that nerfed size.
- rode1974 3mo agoHopefully i get a macbook pro soon enough to run some small or medium sized LLMs
- paperterminal 3mo agoSame, but so much $$
- deleted 3mo ago[deleted]
- bArray 3mo agoApparently GLM 5.2 is 753B parameters [1], what kind of hardware are people using to run this locally? [1] https://huggingface.co/zai-org/GLM-5.2 https://huggingface.co/zai-org/GLM-5.2
- crocowhile 3mo agofollow antirez - https://x.com/antirez/status/2071173841175363905?s=20 https://x.com/antirez/status/2071173841175363905?s=20
- JamesSwift 3mo agoThats quantized
- nozzlegear 3mo agohttps://xcancel.com/antirez/status/2071173841175363905 https://xcancel.com/antirez/status/2071173841175363905
- anentropic 3mo agoIt's a nice technical achievement but looks unusably slow for actual work
- dakolli 3mo ago8 X RTX6000. It will run you around 80-100k to get started with a model at this size with decent tps.. Don't worry though, open source evangelists will tell you that these will be running on your phone in the next 3 years. For $100k you could run this model 24/7 through open router with 10 concurrent sessions at 50tps for a decade and have money left over for a vacation. There's no point in investing this type of money in local models unless you have a business where you're already paying for many employee's individual token usage.
- 8note 3mo agoyou can however, have fun with it. oil workers buy 100k trucks they do not-much with. why not a 100k in computer?
- theteapot 3mo ago> Constant: the IDOR dataset (the same real, open-source applications we've used in prior research) ... What we're they? Also, wouldn't one expect a more recently released coding agent (with a more recent knowledge cut off) to perform better because they have access to more knowledge about vulns in these OSS projects, and even possibly have knowledge of your own "prior research"?
- mkagenius 3mo agoOne would. But then the results are even weirder as opus 4.6 scored more than opus 4.8 by a huge margin
- BikiniPrince 3mo agoThis is a joke right? I wouldn't install this in a sandbox.
- mlnj 3mo agoWhy? Don't tell me you've never tried a non-US based model, ever. There's a number of US providers who also run it, if that is your preference.
- g42gregory 3mo agoIf only the "cybersecurity" crowd were focused on patching the vulnerabilities. Instead of shilling for the LLM providers.
- _factor 3mo agoThe robot figured out how to bump the lock. The obvious solution is to ban the robot.
- __MatrixMan__ 3mo agoBut if we patch all of the vulnerabilities, who will pay for our vulnerability scanner?
- pimeys 3mo agoI have taken another look on these open models after the fiasco of Fable and GPT 5.6 this weekend and... GLM-5.2 truly is a good workhorse model for daily programming. I consider myself a heavy user of LLMs and a seasoned developer. A typical session for me with GPT is usually over a hundred dollars... This weekend I programmed a matrix bot with encryption and a Rust agent with some tools. Because I need one and OpenClaw just felt... not what I wanted. Two days later and 20 dollars poorer I have what I need: a multimodal agent written in rust that has access to my homelab. Nothing felt off with GLM. It did what I wanted, was fast, had a decent not very annoying personality and was much cheaper than Opus or GPT. I used it unquantized through Fireworks, but there are multiple other providers too.
- dist-epoch 3mo agoAnthropic is saying other models were good at detecting vulnerabilities, where Mythos excelled was in creating functional exploits for them. This article only talks about detecting vulnerabilities, so it's unclear if it's a true Mythos equivalent.
- igregoryca 3mo agoIt seems "Mythos is really good at finding vulnerabilities" has been what people took away from the Project Glassing announcement, which makes sense. Unfortunately for Anthropic, most seem to have forgotten the best argument Anthropic had for holding Mythos back from the general public, "it's crazy good at crafting exploits". Then, without that context, the tinfoil hats came out.
- laybak 3mo agohow representative are Semgrep's benchmarks? everyone seems to have their own benchmark these days (guess it's good "content marketing") I'm honestly losing track
- csjh 3mo agoI found it to spiral into complete nonsense a few times when I tested it out, but it's possible that was a bug in the provider
- aussinholdn 3mo ago[dead]
- SwellJoe 3mo agoI added GLM 5.2 to my security bug hunting benchmark when it came out, and found it to be a good performer, but not the best open model. The benchmark tests whether models can find bugs Mythos found. The best open models in the initial benchmark were DeepSeek V4 Pro or MiMo 2.5 Pro. But it turned out MiMo got lucky, it's performed worse on almost every test I've done since, while DeepSeek has consistently been among the best performers and its extreme caching performance makes it cheaper than just about anything, including much smaller models. https://swelljoe.com/post/will-it-mythos/ https://swelljoe.com/post/will-it-mythos/ Also of note, I found giving models access to the open source semgrep as a tool makes some perform worse and none perform better, though it's plausible there's a way to wire it up in a harness that presents useful information to the model without the model having to know how to use it (my theory is that semgrep isn't heavily represented in the training data, so you're asking the model to do two things at once: figure out how to use semgrep and find security bugs, and both tasks suffer for the lack of focus...most small models, and some big models, can't do that well). Edit: But, also, more testing is ongoing. I suspect GLM 5.2 will also be a consistently strong performer. It seems to excel at most things I've tested on it.
- Barbing 3mo agoWe need a benchmark of independent community sourced benchmarks! …probably already is one
- SwellJoe 3mo agoI don't know how you'd judge benchmarks beyond "did it test and measure what it says it tests and measures". And, I guess there have been instances where the benchmark failed to do that, and the models could cheat in some way and it just tested the models ability to find the answer key. In the case of my benchmarks every model other than Claude models running in Claude Code never have network access and all information from after the bug was discovered has been removed from the repository the model can see. But, there are benchmarks for so many different kinds of ability, I don't know how to compare them directly against one another. Like, models that do well on terminal and agentic coding benchmarks tend to do well on finding security bugs, but it's not a 1:1 correlation, there are surprises.
- cmrdporcupine 3mo agoI like GLM 5.2... ish. It's ok. I'd be mostly fine switching to it. I just can't find a cost effective way to do that. z.AI's coding plan is both overpriced and unreliable. ollama's is also overpriced. Paying by the token for it on openrouter etc is more expensive than just having a Codex or Claude coding plan. If you have to pay by the token, it's clearly cheaper. It's not competitive with a coding plan though.
- TurdF3rguson 3mo agoIt also means giving up vision which I don't know how I would deal with. I think I would prefer a weaker model with vision than a stronger without.
- cmrdporcupine 3mo agoIf you using opencode or similar you can just temporarily switch models -- in the same session -- to something that has vision and have it look at your image. And then switch back.
- gazpachotron 3mo agoOr create an agent or subagent that just looks at images, and specify a vision model for that agent.
- TurdF3rguson 3mo agoI don't see how that helps, I would still need to somehow get the image into the coding model's context.
- nozzlegear 3mo agoWhy's that?
- gmerc 3mo agovision runs just fine locally for most usecases, so it's really just a skill to call that Ollama instance
- lenerdenator 3mo agoThe incentive to develop Claude further is to make money. The incentive to develop these Chinese models further is to trash the business case of most American AI labs.
- _s_a_m_ 3mo agoI tried GLM many times and it is bad, i have on clue what these people are talking about
- jeffnash 3mo agohave you tried 5.2? I agree that 5.1 and prior were below Kimi, Mimo, Qwen, Minimax, and probably Deepseek (depending on task), but 5.2 (especially unquantized) feels like something else. Now I feel like that I'm covered by GLM 5.2 and Minimax M3 (when I need vision or a second pass on something).
- throw10920 3mo agoBad for security research or for general coding? Having used GLM 5.2 for non-security software work, I can say it's better than Sonnet (but not Opus), and cheaper than both (because when you steal someone else's IP, you don't have to amortize the cost of their R&D).
- byzantinegene 3mo agostealing someone's ip... hmmmm
- thefourthchime 3mo agoSame. I asked it my Pac-Man question and it was the first to DNF. It just goes off getting confused about how to design the map for 15 minutes and then times out.
- Art9681 3mo agoThis is because of the safeguards and not the model capabilities. If these folks signed up for the proper cyber service offered by Anthropic where refusals are removed then the open weight model wouldn't look as capable.
- unnouinceput 3mo agoAnd just like Linux lost to Windows in consumer market due to devs/creator's stubbornness, same will happen with closed vs open LLM. In the end the one that is used the most will be the one that you train your kids on and therefore the one that wins the market. Eventually the closed one with too much guardrail will be left behind because people will stop using it. You need to read the market. Linus didn't read it in 90's, Gates did and that's why Windows is in almost every home.
- throwaway676712 3mo agoIs this 2006? Linux is present on literal billions of android phones, servers, supercomputers and other embedded devices. It's the most ubiquitous OS on the planet and it's not even close, even Microsoft contributes to it. The only niche where it doesn't utterly dwarf the competition is personal computers and it looks like we're all getting priced out of that anyway
- unnouinceput 3mo agoI was 100% sure that somebody will throw this, but I didn't actually expect to do it from a throwaway account. Maybe because you know calling Android Linux is like calling a human just an ape (but mirrored because you know, Android in this case is the ape). Oh, and Microsoft Loves Linux, right? Because that's why they invented WSL, to make people go and use Linux, right? riight!!
- throwaway676712 3mo agoMoving goalposts. Android is Linux (the GNU part of GNU/Linux is trivial to add), people use Linux knowingly or not, and Linux is on the most devices on Earth regardless of whether their users know it or not.
- TacticalCoder 3mo agoHow to reconcile that with the recent, highly upvoted, article titled: "The gap between open weights LLMs and closed source LLMs"? What explains it? Is TFA lying? Is the most upvoted comment here lying?
- Bigpet 3mo agoTop comment doesn't say it's better. Just says it's a "workhorse". The article itself doesn't say "it's better", basically just says "in this one specific benchmark it beat Claude with Claude code". Mind you with multimodality it Opus still beat GLM 5.2 very handily in that same benchmark. I can't find any contradiction and I don't see anyone lying directly. At most they lead you to imply false things, but they're not untrue at a literal reading.
- deleted 3mo ago[deleted]
- utunga 3mo agoJust popping in to say that no you can't use the word "tokenomics" to mean that. Argh.
- yieldcrv 3mo agowho is your favorite hosted GLM 5.2 provider? I'm looking for fastest tokens/sec and best cost additionally, reliable API, because z.ai can be finicky also, not for Enterprise use, but I like non-US providers, I don't care if the party happens to be the one reading my information and stealing my trade secrets, if they won't respond to a US subpoena
- slashdave 3mo agoAdvertisement
- dools 3mo agoI think Opus 4.8 is deliberately nobbled. Kimi k2.6 with Kimi code beats opus models at finding vulnerabilities, even though it produces some false positives, when I give the same issues to opus and ask it to verify most of the time it concurs it’s a real issue even though it failed to find the issue itself
- deleted 3mo ago[deleted]
- jackdawed 3mo agoI use GLM 5.2 via Neuralwatt and it's gotten so cheap I wouldn't mind cancelling my personal Claude subscription if work gave me one. I've spent 374M tokens this month and it only cost me $18 on energy-based pricing.
- cmrdporcupine 3mo agoHow's the reliability and speed?
- lowbloodsugar 3mo agoFelt like I was reading advertising for their harness.
- croemer 3mo agoThey should also at least run Opus through the same Pydantic harness they used for GLM. As is, it's apples vs pears. Where's the cost per vulnerability for all the other models than GLM? Also, without code this isn't very trustworthy. Could all be made up as well.
- CurbStomper 3mo ago[dead]
- XCSme 3mo agoDoes a bit worse than Opus 4.8 in my tests[0], but it's 5x cheaper and 3x slower. [0]: https://aibenchy.com/compare/anthropic-claude-opus-4-8-medium/z-ai-glm-5-2-medium/ https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...
- XCSme 3mo agoNote that being open-weights, "slower" is relative, as it depends on who's serving the model. This can drastically change over time too.
- nsoonhui 3mo agoNot sure what to make if your benchmark because GPT 5.5(low) ranks higher than GPT 5.5 (medium) -- #4 vs #9
- XCSme 3mo agoYou'd be surprised, some models on high do worse than on medium, because they start overthinking and doubting themselves, polluting the context with too much information, etc. It depends a lot on the task and harness too (using plans and to-do lists, vs one-shot answers), but for simply answering directly to an inquiry, often extra thinking doesn't necessarily improve the answer, especially if the answer is binary, or can be correct or wrong, as opposed to having more time to refine a creative output.
- XCSme 3mo agoAnother example was Gemini 3.1 flash lite, which on high was basically just burning tokens, costing like 30x more, while giving worse answers: https://aibenchy.com/compare/google-gemini-3-1-flash-lite-high/google-gemini-3-1-flash-lite-preview-medium/ https://aibenchy.com/compare/google-gemini-3-1-flash-lite-hi...
- rvz 3mo agoMany people here are now realizing that open weight models are now able to compete against frontier closed models. This is where we are heading and why many closed labs are terrified of this affecting their bottom line and the reason why they want them banned from being released.
- crazylogger 3mo agoActually they don't even need to compete against frontier closed models, they just need to work. 99.99% people's day jobs aren't competing for the Fields Medal or even finding security vulnerabilities. So it appears while TAM (total addressable market) of AI in general is huge, TAM for frontier LLMs is tiny. Efficiency gains at roughly the same performance might be all people care about from now on.
- deleted 3mo ago[deleted]
- gurjeet 3mo agoTwice in the text quotes Claude Code's F1 score as 32%, but the table shows the score is 37%. It's very likely that the actual score is 32% (because it is referenced 2 times, and a third time indirectly as the difference 'seven'). Oddly, this is a strong indication of the text being hand-written rather than LLM-assisted; it's very likely that a human made a mistake in creating the table. > ... beating Claude Code (32%) ... > ... GLM 5.2 ... beat Claude Code by seven points (39% vs. 32%). > Rank | Configuration | Harness | F1 > ... > 4 | Claude Code (Opus 4.6) | Claude Code SDK | 37%
- insiderphd 3mo agoHello author here, or one of them anyway. I can confirm that it was hand written, 32% was combined all the Claude models (4.6, 4.7, 4.8) mushed into one score, 37% was Opus 4.6 specifically (which did the best)
- Alien1Being 3mo agoThe current US administration has gone a long way towards handing over leadership in AI to China.
- a96 3mo agoAlong with everything else. Almost like having a fascist dictatorship isn't really a very competent way to run a country no matter what the size.
- rbbydotdev 3mo agoArgh, agent benchmarks are so bad and can be gamed easier than bmw emissions tests.
- synergy20 3mo agobut, it's $160/month(unless you buy a one-year plan that gets cheaper), not too far from $200/month from claude and codex? why should I switch?
- sidcool 3mo agoGenuinely curious. Say GLM 5.2 is better than Opus. But how does one go about using it by themselves?
- KronisLV 3mo agoThe simplest would be either OpenRouter: https://openrouter.ai/z-ai/glm-5.2 https://openrouter.ai/z-ai/glm-5.2 Or grabbing their GLM Coding Plan directly: https://z.ai/subscribe https://z.ai/subscribe I went with the second one to try it out, feels pretty okay (with OpenCode, though Claude Code would also work), however it feels like I reach the weekly limits somewhat fast with their 65 USD Pro subscription. They also have that whole peak times thing going on and apparently it will get worse after September: > Supported models and Visual Understanding MCP share the same usage quota. GLM-5.2 and GLM-5-Turbo consume quota at 3x during peak hours and 2x during off-peak hours. Limited-time benefit: off-peak usage is currently charged at only 1x quota through the end of September. Peak hours: 14:00–18:00 daily (UTC+8).
- Mashimo 3mo agoOpenRouter, Z.ai coding plan, OpenCode Go, OpenCode Zen .. and probably more.
- devld 3mo agoDigitalOcean hosts it.
- theptip 3mo agoBut… what effort level? “Opus 4.8” is a massive capability range. If you just ran it on medium that is a completely different result than vs. max.
- goyoon 3mo ago[dead]
- unnouinceput 3mo agoOK, half the article is on and on about harness and scaffolding and whatnot. I kept reading waiting for a benchmark where they give the same scaffolding to GLM like they did to Opus. Where is that one?
- andai 3mo agoMost interesting things to me from their benchmarks: GPT does way worse than Opus without their harness, but better with it. Opus 4.7 and 4.8 do way worse than 4.6. (Intentional nerfing?) Would have been interesting to see GLM in the custom harness. Would also be interesting to run GLM in Claude Code, which it has presumably been fine tuned on.
- jocelyner 3mo ago[dead]
- uluckydev 3mo agoI used Claude a lot, but with Claude Code it takes a lot of context window, and it's very pricey, to be honest. Then I shifted towards Minimax. I used the coding plan because it's cheaper, but it still gets the job done. When M3 came out, I started using it, and it was actually really good. After that, I shifted towards OpenCode for my AI agent, and that's been really good as well. The best thing I realized is that it uses less context, works better, and gives me access to a lot of different models from one place. I never actually used GLM, but I recently found QuanCode, which is amazing. I used it to build a full-stack application. Now I'm shifting my focus more toward SaaS distribution. I'm still figuring out how to automate different workflows, and using QuanCode has been really fast and effective for building those automations.
- Kiog-Aser 3mo ago[dead]
- fishonbike 3mo ago[flagged]
- modgate 3mo ago[flagged]
- questionreality 3mo agohope open source continues to improve
- mohitpaddhariya 3mo agoopen-weight models routinely match or even outperform previous-generation proprietary APIs
- zwJay 3mo ago[dead]
- 40four 3mo agoIt’s hard to argue against the open weight models if your only concern is coding. Which, for many of us hackers here in this forum, it is. But I would like to point out that the overwhelming majority of people using LLMs aren’t programmers, don’t care about coding, and couldn’t even be bothered to “vibe code”. So we should consider the bias of the output of these open weight models, and what that looks like, outside of the context of writing code.
- WinstonSmith84 3mo agoThere is no money made from these people though .. people who are using ChatGPT to plan for their next week-end or their next vacation aren't paying a $100 or $200 monthly subscription. As for non coder office workers (accountants, PMs, etc.), they use Microsoft or Google products which all integrate AI to some extent within their products - with RAG for Sharepoint to some basic AIs to generate text or automate work in spreadsheets .. the models used there are already capable enough for all what's needed (I think Microsoft is using GPT 5.1 or 5.2 in its latest iteration but for sure no GPT 5.4/5.5). The thing is, Software development is where money is made for these labs
- 40four 3mo agoYou’re making a good point. I don’t disagree with what you’re saying. But I think my point got lost. I don’t agree with “Software development is where money is made for these labs”. Coders will inevitably eat up the most tokens & buy the bigger $200 subscriptions because we want to keep working. But us coders are still the small minority of users. They aren’t counting on us to get to trillion dollar evaluations. They are counting on the regular folks to buy the $20/ month subscription. It’s really easy to run out your free tier usage these days, asking questions that have nothing to do with coding. So my point is what does that output look like for someone asking a question about politics or world news?
- bel8 3mo agoI think most of these $20/mo subscrptions will either be Apple's iCloud, Microsoft's Office 365 or Google's Drive+Office plans which already do or will offer bundled AI. I know Google gives me free Gemini AI from my Google Drive plan. Microsoft probably already does too, didn't test. Apple is probably crafting some arrangements if not offering already. My point is most people wont pay for AI. It will be bundled. And I think AI is going to be free for all, with ads.
- protonisafk 3mo agoIt seems benchmarks keep changing and preferring the latest AI agent literally every time.
- chonghaoju 3mo agoEvery agent run writes an audit record. Not for compliance theater — because when something breaks at 2am, you need to know exactly what happened and why.
- kelnos 3mo agoTitle is misleading (and is editorialized from the actual article title). GLM 5.2 did better than Claude in one specific cybersecurity-related benchmark (finding vulnerabilities of one certain type). I don't think you can draw any general conclusions about the relative utility of the two models.
- insiderphd 3mo ago1000% this, this was us internally testing if our harness worked, the motivation was never to test them in-depth 1v1. We were just really shocked at the results, there’s a lot more work to do here.
- croemer 3mo agoCan you run Claude Opus through the same Pydantic harness and add the cost to the benchmark result table? An isolated price is meaningless.
- bingemaker 3mo agoHow do you run GLM? Are there any hosted services?
- port3000 3mo agoOpencode Go subscription ($5 to try for one month) or Neuralwatt are what I use. Both through opensource Opencode harness (like Claude code)
- bingemaker 3mo agoThank you!
- childintime 3mo agoAbout running models locally and why data centers win (for now): they can stream the model weights to many neural engines at the same time, so each of these only needs enough RAM to hold the KV cache. So each engine is cheaper to operate, plus they are time-shared, resulting in massive wins for data centers. So one can see businesses owning their own such cluster, next to their database infra, in the near future.
- maxignol 3mo agoWould you recommand some ressources about how multiple neural engines are used in data centers ?
- ben8bit 3mo agoDefinitely a +1 from me. I've really enjoyed using it via OpenCode/Zen. Not loving the pricing with OC so will probably switch to OpenRouter once my credits are done.
- maxignol 3mo agoHave you tried opencode go ?
- _cs2017_ 3mo agoI don't feel the numbers without the harness are useful. People will use the model with the harness. I know that harness may not be optimized to this model, but it's still more useful to see the numbers from an imperfect harness than from a no harness setup.
- cake-rusk 3mo agoHow do you run this thing? What kind of hardware do you need?
- tmach32 3mo agoI think one thing people are missing about this article is that they are arguing that the harness can make a bigger difference than the model. They aren't merely hyping GLM 5.2.
- xlii 3mo agoI switch from Codex to GLM 5.2 when I'm out of tokens. The main difference for me is time to completion. GPT gets there <5 minutes, GLM 5.2 without context takes ~1H. Though the harness makes a significant difference. On Pi GLM5.2 dreams for minutes, with OpenCode it's more on the point and gets to editing quicker.
- jacomoRodriguez 3mo agoWhich harness do you recommend to run coding task with glm 5.2? Any good resources about this (also for setup and recommend config)?
- mpfect 3mo agoFeeling proud on these Open Models. Its just they need to focus on efficiency as well especially in terms of size.
- tomerbd 3mo agoGLM 5.2 - Super Clear GPT-5.5 - Super Smart Auto/Composer - Super Fast (cursor)
- armcat 3mo agoI find it astounding that ppl still comment “it’s still behind” or “it’s not the best model”. Everything is about the harness. Even the big AI labs are focusing on managing agents - sandboxes, memory, context, skills, loops. With the right harness GLM 5.2 can do no wrong.
- contentkraft 3mo ago[dead]
- spaceman_2020 3mo agoOpus 4.8 is genuinely one of the most frustrating models in casual use. It has a tendency to completely lose context in the middle of a conversation. It’s also too pedantic and nitpicky, and relies on language that’s way too specific to get any work done. I always end up being frustrated with it and revert to opus 4.6
- mattmcdonagh 3mo agoGLM-5.2 suggests long-horizon agentic work is becoming open, cheap, and deployable. What does that mean for the frontier? https://lifeinthesingularity.com/p/glm-52-proves-ai-comes-for-all-moats https://lifeinthesingularity.com/p/glm-52-proves-ai-comes-fo...
- mciair_ 3mo ago[flagged]
- dmix 3mo agoI hope someone is also building a Claude Design competitor. One that is similarly HTML based instead of the Figma/Magic Patterns approach. I have more vendor lock-in with Design than I do with Code, and will switch over as soon as Claude loses the smallest technical advantage
- kroaton 3mo agohttps://github.com/nexu-io/open-design https://github.com/nexu-io/open-design
- softwaredoug 3mo agoAre open labs just loss leaders backed by Chinese govt? Is this like electric cars where the goal is to flood the market with good enough quality for free so they end up dominating the market? Or is there a business model I’m missing?
- 34679 3mo agoUS EVs were also heavily subsidized, but they were all built using Chinese parts.
- someperson 3mo agoThe EV supply chain in the US back in say 2007 certainly had far fewer key parts sourced from China than recent years. As far as US EVs being subsidized early, if you take state and federal tax incentives, DoE grants and loan guarantees as subsidizes then that's true. It's debatable (I think incentives applied to all suppliers not just US ones) but a reasonable statement.
- nojvek 3mo agoTesla given $60M by Obama admin when they were deep in debt and may have gone out of business. so Tesla technically is subsidized by US govt. SpaceX too. Without NASA funding, they'd be long out of business. China and US ain't that different. China realizes that being a tech and industrial powerhouse working on future tech is great for their economy. They bet huge on it. That's how they win. Europe on the other hand is now a laggard.
- Rover222 3mo agoUS EVs were "lightly" subsidized compared to what the Chinese govt has done. In the ballpark of 250 billion dollars by the Chinese vs maybe 10% of that by the US.
- DiogenesKynikos 3mo agoNote that most of those subsidies are things like sales-tax exemptions for EVs and support for charging infrastructure in China. In other words, they're not subsidies for Chinese cars being exported abroad. They're not even directly paid to the manufacturers.
- CurbStomper 3mo ago[dead]
- Roark66 3mo agoHas anyone compared the costs between maxing out a Claude Max x5 subscription (one for €120 euro a month) and same amount of work on GLM5.2 via API at a cost of $4 per mln token out? I have a feeling Anthropic may still come out cheeper (mainly thanks to enterprises subsidising the Max subscriptions). But I'm very excited with the possibility of using fully EU based inference rivalling Opus in quality.
- m3kw9 3mo agoThere is 2 suspicious words "Beats" and "our benchmarks"
- ni5arga 3mo ago> We ran a set of popular open-source models against our IDOR benchmark. "our IDOR benchmark", there you go.
- stbenjam 3mo agoChinese models are almost certainly cheating on benchmarks, I would bet if you saw the training data that the benchmark canaries are in there. GLM may be a good model in general but it s benchmaxxed and definitely not as good as Opus 4.8.
- bel8 3mo agoWhy would you say that? I use DeepSeek V4 Flash (high) and MiMo 2.5 (non Pro, because vision) to work on medium sized projects (~1mil lines of code, C#, Go, TypeScript) with great success. And that is coming from someone who used Opus 4.7 and GPT 5.5 as workhorses before. And I'm pretty sure GLM 5.2 is better than the lighter models I use. My worflow is simple: plan -> clarify -> implement. 1) plan prompt template: I describe what I need and ask LLM to generate a markdown file containing an implementation plan plus at least 10 clarification questions for me to answer. 2) I answer the questions in the plan.md file. 3) implementation prompt template: I ask LLM to implement plan.md and tell me at the end if there were any deviations and new findings during the implementation (there ofter are).
- flowghost_24 3mo agoI am using this with a workflow of Claude Code, Codex, Kimi and GLM and the results are pretty astounding and almost 90% of the times Claude's findings and plans are overturned with Claude's agreement.
- stellamariesays 3mo ago[flagged]
- kraflio 3mo agoExactly the same i am now trying to use and will keep you updated
- dvduval 3mo agoIf it’s not quite as good as the hype yet, I expect it probably will be in the near future. To do a lot of the primary coating tasks needed for most situations, it’s probably gonna be good enough if it isn’t ready. The harness will be there as well.
- johnnyAghands 3mo agoThe title of the post on their blog is really misleading "We have Mythos at Home: GLM 5.2 beats Claude in our Cyber Benchmarks". Mythos (or Fable) isn't even benchmarked, and there's giant caveat literally at the bottom: "We have a caveat: This is one task, one dataset, one run." I think the post is still informative, but very a little disingenuous and clickbaity.
- nizbit 3mo ago[dead]
- aubanel 3mo agoThere's no question to me, after trying both, that Fable is much better than GLM-5.2 when left alone in front of hard coding tasks Now maybe what plateaus is the human collaboration efficiency, because at some point it will be bottlenecked by the human Thus companies who still try to have humans perform intertwined work with their AI won't see an improvement, while the ones who fin the right conditions to give their AI more free rein will see it. Kind of like it's no use having a workhorse pull a combine harvester : at some point, when machines reach sufficient efficiency, you just give wheels to the harvester and let it run.
- simplyluke 3mo agoI've been using it for a week via opencode in a large, mature codebase for some moderately ambitious feature development, and a bit of debugging. Explicit purpose is evaluating if it may be a good substitute to save money for many tasks. For several tasks I've had both it and opus 4.8 attempt the same task and compared them. In general, it's comparable across the board. Claude is less "verbose" -- GLM really likes to comment a ton. There were a few things where I think claude would have needed a little bit less back and forth. So opus still has an edge, but it's marginal, very much unlike previous open/competitor models where benchmarks looked good but actual day to day performance was pretty bad. I'm sure fable is "better" but it's so expensive + data retention policies are such that for the moment it was generally available I couldn't use it for work. This is still notably better performance than when claude code took the industry by storm. I'm understanding why Dario is trying to regulate open weight models away.
- Mona1 3mo ago[dead]
- brammertottens 3mo agoThis is an interesting finding, but very specialised. It would also be great to get some more information about the benchmark. Is it just a collection of files with vulnerabilities, or are they hidden in a real codebase, where LLM based approaches will not be able to scan every file like a static code scanner is able todo.
- mnauf 3mo agoexactly what I needed to hear!
- aussinholdn 3mo ago[dead]
- Herze 3mo ago[flagged]