11 ms·
GLM-5.1: Towards Long-Horizon Tasks
- dang 6mo ago[stub for offtopicness] [[you guys, please don't post like this to HN - it will just irritate the community and get you flamed]]
- louszbd 6mo ago[flagged]
- seven2928 6mo ago[flagged]
- zendi 6mo ago[flagged]
- smith7018 6mo agoHmm, three spam comments posted within 9 minutes of each other. The accounts were created 15 minutes ago, 51 days ago, and 3 months ago. Interesting. Hopefully these aren't bots created by Z.AI because GLM doesn't need fake engagement.
- dang 6mo agoThese comments are probably either by friends of the OP or perhaps associated with the project somehow, which is against HN's rules but not the kind of attack we're mostly concerned with these days. Old-fashioned voting rings and booster comments aren't existential threats and actually bring up somewhat nostalgic feelings at the moment! Thanks for watching out for the quality of HN...
- ray__ 6mo agoWould love to read a Tell HN post about the kinds of attacks you are concerned with!
- dang 6mo agoFor example, there are rings of accounts posting generated comments, presumably in order to build karma for spammy or (let's be kind) promotional reasons. There are also plenty of spam rings that create tons of accounts and whatnot. These are different from the submitter-passed-a-link-to-friends kind of upvoting and booster comments, which feel quaint by comparison. In this case people usually don't know they are breaking HN's rules, which is why they don't try to hide it.
- tadfisher 6mo agoI moderate a medium-sized development subreddit. The sheer volume of spam advertising some AI SaaS company has skyrocketed over the past few months, like 10000%. Comment spam is now a service you can purchase [0][1], and I would not be surprised if Z.ai engaged some marketing firm which ended up purchasing this service. There are YC members in the current batch who are spamming us right now [2]. They are all obvious engagement-bait questions which are conveniently answered with references to the SaaS. [0]: https://www.reddit.com/r/DoneDirtCheap/comments/1n5gubz/get_paid_to_post_comment_on_reddit_1_per_post_05/ https://www.reddit.com/r/DoneDirtCheap/comments/1n5gubz/get_... [1]: https://www.reddit.com/r/AIJobs/comments/1oxjfjs/hiring_paid_reddit_commenters_easy_daily_income/ https://www.reddit.com/r/AIJobs/comments/1oxjfjs/hiring_paid... [2]: https://www.reddit.com/r/androiddev/comments/1sdyijs/no_code_test_automations_for_android_feels_like_a/ https://www.reddit.com/r/androiddev/comments/1sdyijs/no_code...
- greenavocado 6mo agoZ.ai Discord is filled to the brim with people experiencing capacity issues. I had to cancel my subscription with Z.ai because the service was totally unusable. Their Discord is a graveyard of failures. I switched to Alibaba Cloud for GLM but now they hiked their coding plan to $50 a month which is 2.5x more expensive than ChatGPT Plus. Totally insane.
- sourcecodeplz 6mo agoEveryone has started either hiking their prices or limiting the tokens, gravy train is over. Glad we have open models that we can host; Sad RAM is so expensive..
- bigyabai 6mo agoIt's an okay model. My biggest issue using GLM 5.1 in OpenCode is that it loses coherency over longer contexts. When you crest 128k tokens, there's a high chance that the model will start spouting gibberish until you compact the history. For short-term bugfixing and tweaks though, it does about what I'd expect from Sonnet for a pretty low price.
- embedding-shape 6mo ago> It's an okay model. My biggest issue using GLM 5.1 in OpenCode is that it loses coherency over longer contexts Since the entire purpose, focus and motivation of this model seems to have been "coherency over longer contexts", doesn't that issue makes it not an OK model? It's bad at the thing it's supposed to be good at, no?
- wolttam 6mo agolong(er) contexts (than the previous model) It does devolve into gibberish at long context (~120k+ tokens by my estimation but I haven't properly measured), but this is still by far the best bang-for-buck value model I have used for coding. It's a fine model
- verdverm 6mo agoHave you tried gemma4? I'm curious how the bang for buck ratio works in comparison. My initial tests for coding tasks have been positive and I can run it at home. Bigger models I assume are still better on harder tasks.
- disiplus 6mo agoi have glm and kimi. kimi was in most of the cases better and my replacement for claude when i run out of tokens. Now im finding myself using glm more then kimi. Its funny that glm vs kimi, is like codex vs claude. Where glm and codex are better for backend and kimi and claude more for frontend. as kimi did a huge amount of claude distilation it seems to be somewhat based in data https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks https://www.anthropic.com/news/detecting-and-preventing-dist...
- Yukonv 6mo agoUnsloth quantizations are available on release as well. [0] The IQ4_XS is a massive 361 GB with the 754B parameters. This is definitely a model your average local LLM enthusiast is not going to be able to run even with high end hardware. [0] https://huggingface.co/unsloth/GLM-5.1-GGUF https://huggingface.co/unsloth/GLM-5.1-GGUF
- zozbot234 6mo agoSSD offload is always a possibility with good software support. Of course you might easily object that the model would not be "running" then, more like crawling. Still you'd be able to execute it locally and get it to respond after some time. Meanwhile we're even seeing emerging 'engram' and 'inner-layer embedding parameters' techniques where the possibility of SSD offload is planned for in advance when developing the architecture.
- adrian_b 6mo agoFor conversational purposes that may be too slow, but as a coding assistant this should work, especially if many tasks are batched, so that they may progress simultaneously through a single pass over the SSD data.
- QuantumNomad_ 6mo agoThree hour coffee break while the LLM prepares scaffolding for the project.
- RickHull 6mo agoI am on their "Coding Lite" plan, which I got a lot of use out of for a few months, but it has been seriously gimped now. Obvious quantization issues, going in circles, flipping from X to !X, injecting chinese characters. It is useless now for any serious coding work.
- kay_o 6mo agoI am on the mid tier Coding plan to trying it out for the sake of curiosity. During off peak hour a simple 3 line CSS change took over 50 minutes and it routinely times out mid-tool and leaves dangling XML and tool uses everywhere, overwriting files badly or patching duplicate lines into files
- unicornfinder 6mo agoI'm on their pro plan and I respectfully disagree - it's genuinely excellent with GLM 5.1 so long as you remember to /compact once it hits around 100k tokens. At that point it's pretty much broken and entirely unusable, but if you keep context under about 100k it's genuinely on par with Opus for me, and in some ways it's arguably better.
- airstrike 6mo ago100k tokens it's basically nothing these days. Claude Opus 4.6M with 1M context windows is just a different ball game
- 6mo ago
- andrewmcwatters 6mo ago[dead]
- alex7o 6mo agoTo be honest I am a bit sad as, glm5.1 is producing mich better typescript than opus or codex imo, but no matter what it does sometimes go into shizo mode at some point over longer contexts. Not always tho I have had multiple session go over 200k and be fine.
- MegagramEnjoyer 6mo agoWhy is that sad? A free and open source model outperforming their closed source counterparts is always a win for the users
- KaoruAoiShiho 6mo agoThe non-awesome context window is the sad part, but I think a better harness can deal with this.
- DeathArrow 6mo agoAfter the context gets to 100k tokens you should open a new session or run /compact.
- disiplus 6mo agoWhen it works and its not slow it can impress. Like yesterday it solved something that kimi k2.5 could not. and kimi was best open source model for me. But it still slow sometimes. I have z.ai and kimi subscription when i run out of tokens for claude (max) and codex(plus). i have a feeling its nearing opus 4.5 level if they could fix it getting crazy after like 100k tokens.
- DeathArrow 6mo agoWhy don't you start a new session or use the /compact command when context gets to 100k tokens? From my testing it was ok until 145k tokens, the largest context I had before switching to a new session. I think Z.ai officially said it should be good until 200k tokens. Using it in Open Code is compacting the context automatically when it gets too large.
- 6mo ago
- aplomb1026 6mo ago[dead]
- kirby88 6mo agoI wonder how that compare to harness methods like MAKER https://www.cognizant.com/us/en/ai-lab/blog/maker https://www.cognizant.com/us/en/ai-lab/blog/maker
- DeathArrow 6mo agoI am already subscribed to their GLM Coding Pro monthly plan and working with GLM 5.1 coupled with Open Code is such a pleasure! I will cancel my Cursor subscription.
- winterqt 6mo agoComments here seem to be talking like they've used this model for longer than a few hours -- is this true, or are y'all just sharing your initial thoughts?
- BeetleB 6mo agoIt's been out for a while.
- KaoruAoiShiho 6mo agoBlog post is new but the model is about 2 weeks in public.
- stavros 6mo agoMy local tennis court's reservation website was broken and I couldn't cancel a reservation, and I asked GLM-5.1 if it can figure out the API. Five minutes later, I check and it had found a /cancel.php URL that accepted an ID but the ID wasn't exposed anywhere, so it found and was exploiting a blind SQL injection vulnerability to find my reservation ID. Overeager, but I was really really impressed.
- arcanemachiner 6mo agoThat is both amazing and terrifying.
- bglazer 6mo agoThis is insane, I love it.
- disiplus 6mo agoYeah it seems they did not align it to much, at least for now. Yesterday it helped me bypass the bot detection on a local marketplace. that i wanted to scrap some listing for my personal alerting system. Al the others failed but glm5.1 found a set of parameters and tweaks how to make my browser in container not be detected.
- 6mo ago
- jaggs 6mo agoHow does it compare to Kimi 2.5 or Qwen 3.6 Plus?
- DeathArrow 6mo agoCompared to Kimi 2.5 or Qwen 3.6 Plus I don't know, but I ran GLM 5 (not 5.1) side by side with Qwen 3.5 Plus and it was visibly better.
- eis 6mo agoThe blog post has a benchmark comparison table with these two in it
- jaggs 6mo agoThanks, I missed that. It's very interesting. They're quite close, but I found Qwen 3.6 plus was just marginally better than Kimi 2.5. But looking at the stats I'll definitely give GLM 5.1 a try now. [edit: even though looking at it, it's not cheap and has a much smaller context size.And I can't tell about tool use.]
- XCSme 6mo agoGeneral intelligence (not coding) comparison: https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-medium/moonshotai-kimi-k2-5-medium/qwen-qwen3-6-plus-preview-medium/ https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-...
- BoorishBears 6mo agoIs there really no rule that discourages 99% of your interactions with HN from being peddling some useless slop benchmark?
- XCSme 6mo agoIf it's relevant to the discussion, I hope not. I've spent probably over100 hours working on this benchmarking/site platform, and all tests are manually written. For me (and many others that reached out to me) are not useless either. I use this myself regularly when choosing and comparing new models. I honestly beleive it is providing value to the conversation. Let me know if you know of a better platform you can use to compare models, I built this one because I didn't find any with good enough UX.
- gavinray 6mo agoI find the "8 hour Linux Desktop" bit disingenuous, in the fine print it's a browser page: > "build a Linux-style desktop environment as a web application" They claim "50 applications from scratch", but "Browser" and a bunch of the other apps are likely all <iframe> elements. We all know that building a spec-compliant browser alone is a herculean task.
- bredren 6mo agoIt is a big claim without the source and prompting.
- MrPowerGamerBR 6mo agoIn my opinion it would be way cooler if it actually created a real Linux desktop environment instead of only a replica. Would it succeed? Probably not, but it would be way more interesting, even if it didn't work. I find things like Claude's C compiler way more interesting where, even though CCC is objectively bad (code is messy, generates very bad unoptimized code, etc) it at least is something cool and shows that with some human guideance it could generate something even better.
- johnfn 6mo agoGLM-5.0 is the real deal as far as open source models go. In our internal benchmarks it consistently outperforms other open source models, and was on par with things like GPT-5.2. Note that we don't use it for coding - we use it for more fuzzy tasks.
- sourcecodeplz 6mo agoYep, haven't tried 5.1 but for my PHP coding, GLM-5 is 99% the same as Sonnet/Opus/GPT-5 levels. It is unbelievably strong for what it costs, not to mention you can run it locally.
- deepsquirrelnet 6mo agoI am working on a large scale dataset for producing agent traces for Python <> cython conversion with tooling, and it is second only to gemini pro 3.1 in acceptance rates (16% vs 26%). Mid-sized models like gpt-oss minimax and qwen3.5 122b are around 6%, and gemma4 31b around 7% (but much slower). I haven’t tried Opus or ChatGPT due to high costs on openrouter for this application.
- epolanski 6mo agoSame thing I noticed. My use cases are not code editing or authoring related, but when it comes to understanding a codebase and it's docs to help stakeholders write tasks or understand systems it has always outperformed american models at roughly half the price.
- foopod 6mo agoIt really bothers me that people refer to open weight models as being open source. They fundamentally aren't and are more akin to freeware than anything else.
- minimaxir 6mo agoThe focus on the speed of the agent generated code as a measure of model quality is unusual and interesting. I've been focusing on intentionally benchmaxxing agentic projects (e.g. "create benchmarks, get a baseline, then make the benchmarks 1.4x faster or better without cheating the benchmarks or causing any regression in output quality") and Opus 4.6 does it very well: in Rust, it can find enough low-level optimizations to make already-fast Rust code up to 6x faster while still passing all tests. It's a fun way to quantify the real-world performance between models that's more practical and actionable.
- tgtweak 6mo agoShare the harness for that browser linux OS task :)
- EddyAI 6mo ago[dead]
- maxdo 6mo agoOne of the bench maxed models . Every time I tried it , it’s not on par even with other open source models .
- wallmountedtv 6mo agoFeeling very much the same. Attempting to use it through Claude Code as a model it just completely lost all context on what it was doing after a few months and kept short circuiting even with the most helpful prompts I could give, outside of just writing out the answer myself. I really do not get the praise for this model. Being "better than Opus 4.6" is not really something a benchmark will tell you. It's much more a consensus of users liking the flavor of an answer, rather than fueling x% correct on a benchmark.
- kamranjon 6mo agoI'm crossing my fingers they release a flash version of this. GLM 4.7 Flash is the main model I use locally for agentic coding work, it's pretty incredible. Didn't find anything in the release about it - but hoping it's on the horizon.
- epolanski 6mo agoI was very satisfied with GLM5, I'm not gonna lie. Excited to test this.
- mark_l_watson 6mo agoI can’t wait to try it. I set up a new system this morning with OpenClaw and GLM-5, and I like GLM-5 as the backend for Claude Code. Excellent results.
- Alifatisk 6mo agoThere is also GLM-5-Turbo, have you tried it for your claw?
- simonw 6mo agoNot only did this one draw me an excellent pelican... it also animated it! https://simonwillison.net/2026/Apr/7/glm-51/ https://simonwillison.net/2026/Apr/7/glm-51/
- ipsum2 6mo agoIt made it realistic. A pelican is much more likely to be flying in the sky than riding a bicycle.
- _pdp_ 6mo agoSimon, you need to come up with improved benchmarks soon.
- lemonish97 6mo agoAgree. But you can keep the pelican theme in whatever new benchmark you choose to come up with. Iconic at this point.
- fancy_pantser 6mo agolet me see Tayne with a hat wobble
- stingraycharles 6mo agoSurely at this point it’s part of the training set and the benchmark has lost its value?
- Marciplan 6mo agothese comments are as useless as simon posting his pelicans
- blazespin 6mo agoAnthropic's reply? A model you can't use.
- minimaxir 6mo agoMythos is most definitely not in response to this announcement.
- dryarzeg 6mo agoA bit off-topic, but for some reason, even though I don't use LLMs for my job or for my hobbies, or in daily life frequently (and when I do, it's mostly some kind of "rubber duck brainstorm"), when I see open-weight releases like this one or the recent Gemma 4 (which is very good for local models); the first time was with DeepSeek-R1 (this one, despite being blamed for "censorship", was heavily censored only via DeepSeek API, the local model - full-weight 685B, not the distilled ones - was pretty much unhinged regarding censorship on any topic)... there's always one song coming to mind and I simply can't get rid of it no matter how hard I try. "I am the storm that is approaching, provoking..." : )
- dvt 6mo agoEvery single day, three things are becoming more and more clear: (1) OpenAI & Anthropic are absolutely cooked; it's obvious they have no moat (2) Local/private inference is the future of AI (3) There's *still* no killer product yet (so get to work!)
- kcb 6mo agoWhat benefit is there to dropping $50k on GPUs to run this personally besides being a cool enthusiast project?
- blizdiddy 6mo agoIs it so hard to project out a couple product cycles? Computers get better. We’ve gone from $50k workstation to commodity hardware before several times
- kcb 6mo agoSubscription services get all the same benefits from computer hardware getting better. But actually due to scale, batching, resource utilization, they'll always be able to take more advantage of that.
- deminature 6mo agoIntel has just released a high VRAM card which allows you to have 128GB of VRAM for $4k. The prices are dropping rapidly. The local models aren't adapted to work on this setup yet, so performance is disappointing. But highly capable local models are becoming increasingly realistic. https://www.youtube.com/watch?v=RcIWhm16ouQ https://www.youtube.com/watch?v=RcIWhm16ouQ
- kcb 6mo agoThat's 4 32GB GPUs with 600GB/s bw each. This model is not running on that scale GPUs. I think something like 96GB RTX PRO 6000 Blackwells would be the minimum to run a model of this size with performance in the range of subscription models.
- philipwhiuk 6mo agoThis is the flip side of the Project Glasswing stuff... Everyone else isn't that far behind and they aren't all gonna just wall off their new model. A reason that Anthropic will eventually give is 'the competition can do what Glasswing can do so what's the point limiting it'.
- XCSme 6mo agoGLM 5.1 does worse than GLM 5 in my tests[0] (both medium reasoning OR no reasoning). I think the model is now tuned more towards agentic use/coding than general intelligence. [0]: https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-medium/z-ai-glm-5-none/z-ai-glm-5-1-none/ https://aibenchy.com/compare/z-ai-glm-5-medium/z-ai-glm-5-1-...
- XCSme 6mo agoThe (none) version especially shows considerable degradation.
- meidad_g 6mo ago[dead]
- 8dazo 6mo agoJust saw the Claude Mythos post. Not sure when it’s going public, but this feels like a real jump, not just incremental progress. Also waiting for the next GLM release coz specs are looking kind of insane.
- zozbot234 6mo agoGemini and GPT have Deep Research models already, Mythos looks like much the same thing.
- gertlabs 6mo agoWe're still adding samples, but some early takeaways from benchmarking on https://gertlabs.com https://gertlabs.com: Contrary to the model card, its one-shot performance is more impressive than its agentic abilities. On both metrics, GLM 5.1 is competitive with frontier models. But keeping in mind this is an open source model operating near the frontier, it's nothing short of incredible. I suspect 2 issues with the model are keeping it from fully realizing its potential in agentic harnesses: - Context rot (already a common complaint). We are still working on a metric to robustly test and visualize this on the site. - The model was most likely overtrained on standardized toolsets and benchmarks, and isn't as adaptive in using arbitrary tooling in our custom harness simulations. We've decided to commit to measuring intelligence as the ability to use custom, changing tools, instead of being trained to use specific tools (while still always providing a way to run local bash and other common tools). There are arguments to be made for either, but the former is more indicative of general intelligence. Regardless, it's a subtle difference and GLM 5.1 still performs well with tooling in our environments. Crazy week for open source AI. Gemma 4 has shown that large model density is nowhere near optimized. Moats are shrinking. If there are more representations of model performance you'd like to see, I'm actively reading your feedback and ideas.
- IceHegel 6mo ago[dead]
- nareyko 6mo ago[dead]
- DeathArrow 6mo agoIt would be nice if you can test the model with different harnesses, Z.ai's own Z Code, Claude Code, Open Code, Pi, Cursor etc. My impression is that the choice of harness matters a lot.
- gertlabs 6mo agoInteresting idea. The metric I'd intuitively want to see is low variance between harnesses for a smarter model. But if a large sample of models statistically outperformed with a certain harness, that's indeed a valuable signal for a developer.
- EITB_2026 6mo agoGood One Though
- Ms-J 6mo agoZ.ai and their GLM models are pretty low quality. I've been testing it for awhile now since it seemed to have potential as a local model. With this new update it still cannot parse simple, test PDFs correctly. It inconsistently tells me that the value in the name field in the document is incorrect, and has the name reversed to put the last name first. Or that a date is wrong as it's in the past/future, when it is not. Tons of fundamental errors like that. Even when looking at the thinking process there are issues: I used a test website for it to analyze and it says that the sites copyright year states 2026 which is in the future and to investigate as it could be an attack, but right after prints today's correct date. I'm in the process of trying to get it uncensored. Hopefully that will create some use out of z.ai Edit: by the way, which is the best uncensored model at the moment?
- uvu 6mo agoCompletely agree with this statement "Z.ai and their GLM models are pretty low quality." I have been trying out and it's kind of useless compare to SOTA models.
- Ms-J 6mo ago[flagged]
- adrian_b 6mo agoI do not doubt your experience, but such statements should always be qualified by specifying the kind of tasks for which you have tried the models. For all existing models, including for all SOTA models, you can find contradictory statements, that they suck and that they are great. It is very likely that all these statements are true simultaneously, because each model may succeed for some tasks and fail for others, so without specifying the tested tasks any claim that a model was good or bad is worthless.
- rednb 6mo agoI'e been using their models pretty much daily for the past 2 months to work on the codebase of a very complex B2B2C platform written in an unusual functional language (F#) with an angular frontend. I also use Claude premium daily for another client, and i use Codex. and i can tell you that GLM5 is at this point much more capable than Claude and Codex for complex backend end work, complex feature planning, and long horizon tasks. One thing i've noticed is that it is particularly good at following instructions and guidelines, even deep into the execution of a plan. To me the only problem is that z.ai have had trouble with inference : the performance of their API has been pretty poor at times. It looks like this is an hardware issue related to the Huawei chips they use rather than an issue with the model itself. The situation has been substantially improving over the past few weeks. GLM5.1, GLM5-Turbo and GLM5v are at this point better than Opus, Codex, Gemini and other claude source models. We have reached a major turning point. To me, the only closed source model still in the game is codex as it is much faster at executing simple tasks and implementing already created plans. Try GLM5v for your PDF work, it's their last generation vision model that has been released a couple of days ago.
- aryehof 6mo ago[dead]
- clark1013 6mo agoI’ve been using GLM 5.1 instead of GPT 5.4 for a few days now, and it’s working smoothly.
- deleted 6mo ago[deleted]
- bdeol22 6mo agoLong-horizon demos are fun; the product test is still interrupted real life—can it pick up three days later without you re-teaching context?
- claud_ia 6mo ago[dead]
- Manchitsanan 6mo ago[dead]
- redoh 6mo ago[flagged]
- deleted 6mo ago[deleted]