16 ms·
GLM-4.7: Advancing the Coding Capability
- esafak 10mo agoThe terminal bench scores look weak but nice otherwise. I hope once the benchmarks are saturated, companies can focus on shrinking the models. Until then, let the games continue.
- CuriouslyC 10mo agoWe're not gonna see significant model shrinkage until the money tap dries up. Between now and then, we'll see new benchmarks/evals that push the holes in model capabilities in cycles as they saturate each new round.
- lanthissa 10mo agoisn't gemini 3 flash already model shrinkage that does well in coding?
- hedgehog 10mo agoSmaller open-weights models are also improving noticeably (like Qwen3 Coder 30B), the improvements are happening at all sizes.
- cmrdporcupine 10mo agoDevstral Small 24b looks promising as something I want to try fine tuning on DSLs, etc. and then embedding in tooling.
- hedgehog 10mo agoI haven't tried it yet, but yes. Qwen3 Next 80B works decently in my testing, and fast. I had mixed results with the new Nemotron, but it and the new Qwen models are both very fast to run.
- mark_l_watson 10mo agoSame experience: on my old M2 Mac with just 32B of memory both Qwen 3 30B and the new Nemotron models are very useful for coding if I prepare a one-shot prompt with directions and relevant code. I don’t like them for agentic coding tools. I have mentioned this elsewhere: it is deeply satisfying to mix local model use with commercial APIs and services.
- Imustaskforhelp 10mo agoHow much billion parameter model is gemini 3 flash, I can't seem to find info about it online.
- skippyboxedhero 10mo agoXiaomi, Nvidia Nemotron, Minimax, lots of other smaller ones too. There are massive economic incentives to shrink models because they can be provided faster and at lower cost. I think even with the money going in, there has to be some revenue supporting that development somewhere. And users are now looking at the cost. I have been using Anthropic Max for most of this year after checking out some of these other models, it is clearly overpriced (I would also say their moat of Claude Code has been breached). And Anthropic's API pricing is completely crazy when you use some of the paradigms that they suggest (agents/commands/etc) i.e. token usage is going up so efficient models are driving growth.
- naasking 10mo ago> We're not gonna see significant model shrinkage until the money tap dries up. I'm not sure about that. Microsoft has been doing great work on "1-bit" LLMs, and dropping the memory requirements would significantly cut down on operating costs for the frontier players.
- bigyabai 10mo agoIt's a good model, for what it is. Z.ai's big business prop is that you can get Claude Code with their GLM models at much lower prices than what Anthropic charges. This model is going to be great for that agentic coding application.
- maxdo 10mo ago… and wake up every night because you saved a few dollars , there are bugs and they are due to this decision?
- Imustaskforhelp 10mo agowell I feel like all models are converging and maybe claude is good but only time will tell as gemini flash and GLM put pressure on claude/anthropic models People (here) are definitely comparing it to sonnet so if you take this stance of saving a few dollars, I am sure that you must be having the same opinion of using opus model and nobody should use sonnet too Personally I am interested in open source models because they would be something which would have genuine value and competition after the bubble bursts
- bigyabai 10mo agoI pay for both Claude and Z.ai right now, and GLM-4.7 is more than capable for what I need. Opus 4.5 is nice but not worth the quota cost for most tasks.
- csomar 10mo agoYeah because Claude never makes bugs?
- theshrike79 10mo agoz.ai models are crazy cheap. The one year lite plan is like 30€ (on sale though). Complete no-brainer to get it as a backup with Crush. I've been using it for read-only analysis and implementing already planned tasks with pretty good results. It has a slight habit of expanding scope without being asked. Sometimes it's a good thing, sometimes it does useless work or messes things up a bit.
- maxdo 10mo agoI tried several times . It is no match in my personal experience with Claude models . There’s almost no place for second spot from my point of view . You are doing things for work each bug is hours of work, potentially lost customer etc . Why would you trust your money … just to back up ?
- theshrike79 10mo agoI'm using it for my own stuff and I'm definitely not dropping however much it costs for the Claude Max plans. That's why I usually use Claude for planning, feed the issues to beads or a markdown file and then have Codex or Crush+GLM implement them. For exploratory stuff I'm "pair-programming" with Claude. At work we have all the toys, but I'm not putting my own code through them =)
- maxdo 10mo agoit's beyond me, why do you need Max plans? I use Opus/Sonnet/Gemini,GPT 5.2 every day in cursor and I'm not paying Claude Max.
- theshrike79 10mo agoI'm mostly just coding at night after the family goes to bed and even I can hit Claude Pro limits - and I started AI assisted programming when we didn't have monthly plans and I had to pay every token out of my own pocket. I learned to be pretty efficient with token use after the first bill dropped :D
- 10mo ago
- anonzzzies 10mo agoShrinking and speed; speed is a major thing. Claude Code is just too slow, very good but it has no reasonable way to handle simple requests because of the overhead, so then everything should just be faster. If I were Anthropic, I would've bought Groq or Cerebras by now. Not sure if they (or the other big ones) are working on similar inference hardware to provide 2000tok/s or more.
- pqtyw 10mo agoZ.ai (at least mid/top end subscription not sure about the API) is pretty slow too especially during some periods. Cerebras of course is probably a different story (if its not quantitized)
- deleted 10mo ago[deleted]
- cmrdporcupine 10mo agoRunning it in Crush right now and so far fairly impressed. It seems roughly in the same zone as Sonnet, but not as good as Opus or GPT 5.2.
- alok-g 10mo agoFor others like me who did not know about Crush: https://github.com/charmbracelet/crush https://github.com/charmbracelet/crush https://news.ycombinator.com/item?id=44736176 https://news.ycombinator.com/item?id=44736176
- XCSme 10mo agoFunny how they didn't include Gemini 3.0 Pro in the bar chart comparison, considering that it seems to do the best in the table view.
- deleted 10mo ago[deleted]
- jychang 10mo agoAlso, funny how they included GPT-5.0 and 5.1 but not 5.2... I'm pretty sure they ran the benchmarks for 5.0, then 5.1 came out, so they ran the benchmarks for 5.1... and then 5.2 came out and they threw their hands up in the air and said "fuck it".
- XCSme 10mo agoI didn't even notice that, I assumed it was the latest GPT version.
- amelius 10mo agoafter or before running the benchmarks?
- rynn 10mo agogpt-5.2 codex isn't available in the API yet. If you want to be picky they could've compared it against gpt-5 pro gpt-5.2 gpt-5.1 gpt-5.1-codex-max gpt-5.2 pro all depending on when they ran benchmarks (unless, of course, they are simply copying OAI's marketing). At some point it's enough to give OAI a fair shot and let OAI come out with their own PR, which they doubtlessly will.
- guluarte 10mo agoGemini is garbage and does it's own thing most of the time ignoring the instructions
- Tiberium 10mo agoThe frontend examples, especially the first one, look uncannily similar to what Gemini 3 Pro usually produces. Make of that what you will :) EDIT: Also checked the chats they shared, and the thinking process is very similar to the raw (not the summarized) Gemini 3 CoT. All the bold sections, numbered lists. It's a very unique CoT style that only Gemini 3 had before today :)
- reissbaker 10mo agoI don't mind if they're distilling frontier models to make them cheaper, and open-sourcing the weights!
- Imustaskforhelp 10mo agoSame, although gemini 3 flash already gives a run for the cheaper aspect but a part of me really wants to get open source too because that way if I really want to some day, I can have privacy or get my own hardware to run it I genuinely hope that gemini 3 flash gets open sourced but I feel like that can actually crash the AI bubble if something like this happens because I genuinely feel like although there are still some issues of vibing with the overall model itself, I find it very competent overall and fast and I genuinely feel like at this point, there might be some placebo effects too but in reality, the model feels really solid. Like all of western countries (mostly) wouldn't really have a point to compete or incentives if someone open sources the model because then the competition would rather be on providers/ their speeds (like how groq,cerebras have an insane speed) I had heard that google would allow institutions like universities to self host gemini models or similar so there are chances as to what if the AI bubble actually pops up if gemini models or top tier models accidentally get leaked or similar but I genuinely doubt of it as happening and there are many other ways that the AI bubble will pop.
- scotty79 10mo agoModels being open weights lets infrastructure providers compete in delivering models as service, fastests and cheapest. At some point companies should be forced to release the weights after a reasonable time passed since they sold the service for the first time. Maybe after 3 years or so. It would be great for competition and security research.
- jtrn 10mo agoMy quickie: MoE model heavily optimized for coding agents, complex reasoning, and tool use. 358B/32B active. vLLM/SGLang only supported on the main branch of these engines, not the stable releases. Supports tool calling in OpenAI-style format. Multilingual English/Chinese primary. Context window: 200k. Claims Claude 3.5 Sonnet/GPT-5 level performance. 716GB in FP16, probably ca 220GB for Q4_K_M. My most important takeaway is that, in theory, I could get a "relatively" cheap Mac Studio and run this locally, and get usable coding assistance without being dependent on any of the large LLM providers. Maybe utilizing Kimik2 in addition. I like that open-weight models are nipping at the feet of the proprietary models.
- embedding-shape 10mo ago> Supports tool calling in OpenAI-style format So Harmony? Or something older? Since Z.ai also claim the thinking mode does tool calling and reasoning interwoven, would make sense it was straight up OpenAI's Harmony. > in theory, I could get a "relatively" cheap Mac Studio and run this locally In practice, it'll be incredible slow and you'll quickly regret spending that much money on it instead of just using paid APIs until proper hardware gets cheaper / models get smaller.
- reissbaker 10mo agoNo, it's not Harmony; Z.ai has their own format, which they modified slightly for this release (by removing the required newlines from their previous format). You can see their tool call parsing code here: https://github.com/sgl-project/sglang/blob/34013d9d5a591e3c05b4ccf01e3bdc67898dc3e9/python/sglang/srt/function_call/glm47_moe_detector.py#L127 https://github.com/sgl-project/sglang/blob/34013d9d5a591e3c0...
- embedding-shape 10mo agoMan, really? Why, just why? If it's similar, why not just the same? It's like they're purposefully adding more work for the ecosystem to support their special model instead of just trying to add more value to the ecosystem.
- gigatexal 10mo agoEven if this is one or two iterations behind the big models Claude or openai or Gemini it’s showing large gains. Here’s hoping this gets even better and better and I can run this locally and also that it doesn’t melt my PC.
- Imustaskforhelp 10mo agoAlthough one would hope they can run it locally (which I hope so too but I doubt that with the increase of ram prices, I feel like its possible around 2027-2028). but Even if in the meanwhile we can't, I am sure that competition in general (on places like Openrouter and others) would give a meaningful way to cheapen the prices overall even further than the monopolistic ways of claude (let's say). It does feel like these models are only behind 6 months tho as many like to say and for some things its 100% reasonable to use it and for some others not so much.
- gigatexal 10mo agoI’ve 128GB of memory in my laptop. But running models with LM studio turns the fans to 100 and isn’t as effective as the hosted models. So I’m not worried about ram. I’m hoping for a revolution or what comes after LLMs to see if local will be better.
- larodi 10mo agoFrom my limited exposure to these models, they seem very very very promising.
- maxdo 10mo agoFunny enough they excluded 4.5 opus :)
- buppermint 10mo agoI've been playing around with this in z-ai and I'm very impressed. For my math/research heavy applications it is up there with GPT-5.2 thinking and Gemini 3 Pro. And its well ahead of K2 thinking and Opus 4.5.
- sheepscreek 10mo ago> For my math/research heavy applications it is up there with GPT-5.2 thinking and Gemini 3 Pro. And it’s well ahead of K2 thinking and Opus 4.5. I wouldn’t use the z-ai subscription for anything work related/serious if I were you. From what I understand, they can train on prompts + output from paying subscribers and I have yet to find an opt-out. Third party hosting providers like synthetic.new are a better bet IMO.
- BeetleB 10mo agoFrom their privacy policy: "If you are enterprises or developers using the API Services (“API Services”) available on Z.ai, please refer to the Data Processing Addendum for API Services." ... In the addendum: "b) The Company do not store any of the content the Customer or its End Users provide or generate while using our Services. This includes any texts, or other data you input. This information is processed in real-time to provide the Customer and End Users with the API Service and is not saved on our servers. c) For Customer Data other than those provided under Section 4(b), Company will temporarily store such data for the purposes of providing the API Services or in compliance with applicable laws. The Company will delete such data after the termination of the Terms unless otherwise required by applicable laws."
- sheepscreek 10mo agoI stand corrected - it seems they have recently clarified their position on this page towards the very end: https://docs.z.ai/devpack/overview https://docs.z.ai/devpack/overview > Data Privacy > All Z.ai services are based in Singapore. > We do not store any of the content you provide or generate while using our Services. This includes any text prompts, images, or other data you input.
- observationist 10mo agoGrok 4 Heavy wasn't considered in comparisons. Grok meets or exceeds the same benchmarks that Gemini 3 excels at, saturating mmlu, scoring highest on many of the coding specific benchmarks. Overall better than Claude 4.5, in my experience, not just with the benchmarks. Benchmarks aren't everything, but if you're going to contrast performance against a selection of top models, then pick the top models? I've seen a handful of companies do this, including big labs, where they conveniently leave out significant competitors, and it comes across as insecure and petty. Claude has better tooling and UX. xAI isn't nearly as focused on the app and the ecosystem of tools around it and so on, so a lot of things end up more or less an afterthought, with nearly all the focus going toward the AI development. $300/month is a lot, and it's not as fast as other models, so it should be easy to sell GLM as almost as good as the very expensive, slow, Grok Heavy, or so on. GLM has 128k, grok 4 heavy 256k, etc. Nitpicking aside, the fact that they've got an open model that is just a smidge less capable than the multibillion dollar state of the art models is fantastic. Should hopefully see GLM 4.7 showing up on the private hosting platforms before long. We're still a year or two from consumer gear starting to get enough memory and power to handle the big models. Prosumer mac rigs can get up there, quantized, but quantized performance is rickety at best, and at that point you look at the costs of self hosting vs private hosts vs $200/$300 a month (+ continual upgrades) Frontier labs only have a few years left where they can continue to charge a pile for the flagship heavyweight models, I don't think most people will be willing to pay $300 for a 5 or 10% boost over what they can run locally.
- lame-robot-hoax 10mo agoGrok, in my experience, is extremely prone to hallucinations when not used for coding. It will readily claim to have access to internal Slack channels at companies, it will hallucinate scientific papers that do not exist, etc. to back its claims. I don’t know if the hallucinations extend to code, but it makes me unwilling to consider using it.
- observationist 10mo agoFair - it's gotten significantly better over the last 4 months or so, and hallucinations aren't nearly as bad as they once were. When I was using Heavy, it was excellent at ensuring grounding and factual statements, but it's not worth $100 more than ChatGPT Pro in capabilities or utility. In general, it's about the same as ChatGPT Pro - once every so often I'll have to call out the model making something up, but for the most part they're good at using search tools and ensuring claims get grounding and confirmation. I do expect them to pull ahead, given the resources and the allocation of developers at xAI, so maybe at some point it'll be clearly worth paying $300 a month compared to the prices of other flagships. For now, private hosts and ChatGPT Pro are the best bang for your buck.
- desireco42 10mo agoI've been using Z.Ai coding plan for last few months, generally very pleasant experience. I think with GLM-4.6 they had some issues which this corrects. Overall solid offering, they have MCP you plug into ClaudeCode or OpenCode and it just works.
- jbm 10mo agoI'm surprised by this; I have it also and was running through OpenCode but I gave up and moved back to Claude Code. I was not able to get it to generate any useful code for me. How did you manage to use it? I am wondering if maybe I was using it incorrectly, or needed to include different context to get something useful out of it.
- big_man_ting 10mo agoi'm in the same boat as you. i really wanted to like OpenCode but it doesn't seem to work properly for me. i keep going back to CC.
- csomar 10mo agoI've been using it for the last couple months. In many cases, it was superior to Gemini 3 Pro. One thing about Claude Code, it delegates certain tasks to glm-4.5 air and that drops performance a ton. What I did is set the default models to 4.6 (now 4.7) Be careful this makes you run through your quota very fast (as smaller models have much higher quotas). ANTHROPIC_DEFAULT_HAIKU_MODEL=glm-4.7 ANTHROPIC_DEFAULT_MODEL=glm-4.7 ANTHROPIC_DEFAULT_OPUS_MODEL=glm-4.7 ANTHROPIC_DEFAULT_SONNET_MODEL=glm-4.7
- tonyhart7 10mo agoless than 30 bucks for entire year, insanely cheap (I know that people must pay it on privacy) but still for maybe playing around with still worth it imo
- sumedh 10mo agoAre you saying the reason they are offering it so cheap is because they are training on user data?
- polyrand 10mo agoA few comments mentioning distillation. If you use claude-code with the z.ai coding plan, I think it quickly becomes obvious they did train on other models. Even the "you're absolutely right" was there. But that's ok. The price/performance ratio is unmatched.
- Havoc 10mo ago>Even the "you're absolutely right" was there. I don't think that's particularly conclusive for training on other models. Seems plausible to me that the internet data corpus simply converges on this hence multiple models doing this. ...or not...hard to tell either way.
- hashbig 10mo agoI had Gemini 3 Flash hit me this morning with "you're absolutely right" when I corrected it on a mistake it did. It's not conclusive of anything.
- polyrand 10mo agoThat's interesting, thanks for sharing! It's a pattern I saw more often with claude code, at least in terms of how frequently it says it (much improved now). But it's true that just this pattern alone is not enough to infer the training methods.
- theptip 10mo agoOr it’s conclusive of an even broader trend!
- ljosifov 10mo agoI imagine - and sure hope so - everyone trains on everything else. Distillation - ofc if one has bigger/other models providing true posterior token probabilities in the (0,1) interval (a number between 0 and 1), rather than 1-hot-N targets that are '0 for 200K-sans-this-token, and 1 for the desired output token' - one should use the former instead of the latter. It's amazing how as a simple as straightforward idea should face so much resistance (paper rejected) and from the supposedly most open minded and devoted to knowing (academia) and on the wrong grounds ('will have no impact on industry'; in fact - it's had tremendous impact on industry; better rejection wd have been 'duh it is obvious'). We are not trying to torture the model and the gpu cluster to be learning from 0 - when knowledge is already available. :-)
- anonzzzies 10mo agoI have been using 4.6 on Cerebras (or Groq with other models) since it dropped and it is a glimpse of the future. If AGI never happens but we manage to optimise things so I can run that on my handheld/tablet/laptop device, I am beyond happy. And I guess that might happen. Maybe with custom inference hardware like Cerebras. But seeing this generate at that speed is just jaw dropping.
- wyre 10mo agoCerebras and Groq both have their own novel chip designs. If they can scale and create a consumer friendly product that would be a great, but I believe their speeds are due to them having all of their chips networked together, in addition to design for LLM usage. AGI will likely happen at the data center level before we can get on-device performance equivalent to what we have access to today (affordably), but I would love to be wrong about that.
- fgonzag 10mo agoApple's M5 Max will probably be able to run it decently (as it will fix the biggest issue with the current lineup, prompt processing, in addition to a bandwidth bump). That should easily run an 8 bit (~360GB) quant of the model. It's probably going to be the first actually portable machine that can run it. Strix Halo does not come with enough memory (or bandwidth) to run it (would need almost 180GB for weights + context even at 4 bits), and they don't have any laptops available with the top end (max 395+) chips, only mini PCs and a tablet. Right now you only get the performance you want out of a multi GPU setup.
- mrbonner 10mo agoI tried this on OpenRouter chat interface to write a few documents. Quick thoughts: Its writing has less vibe of AI due to the lack of em-dashes! I primarily use Kimi2 Thinking for personal usage. Kimi writing is also very good, on par with the frontier models like Sonnet or Gemini. But, just like them, Kimi2 also feels AI. I can't quantify or explain why, though. For work, it is Claude Code and Anthropic exclusively.
- 2001zhaozhao 10mo agoCerebras is serving GLM4.6 at 1000 tokens/s right now. They're probably likely to upgrade to this model. I really wonder if GLM 4.7 or models a few generations from now will be able to function effectively in simulated software dev org environments, especially that they self-correct their errors well enough that they build up useful code over time in such a simulated org as opposed to increasing piles of technical debt. Possibly they are managed by "bosses" which are agents running on the latest frontier models like Opus 4.5 or Gemini 3. I'm thinking in the direction of this article: https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents https://www.anthropic.com/engineering/effective-harnesses-fo... If the open source models get good enough, then the ability to run them at 1k tokens per second on Cerebras would be a massive benefit compared to any other models in being able to run such an overall SWE org quickly.
- chrisfrantz 10mo agoThis is where I believe we are headed as well. Frontier models "curate" and provide guardrails, very fast and competent agents do the work at incredibly high throughput. Once frontier hits cracks the "taste" barrier and context is wide enough, even this level of delivery + intelligence will be sufficient to implement the work.
- andai 10mo agoTaste is why I switched from GLM-4.6 to Sonnet. I found myself asking Sonnet to make the code more elegant constantly and then after the 4th time of doing that laughed at the absurdity and just switched models. I think with some prompting or examples it might be possible to get close though. At any rate 1k TPS is hard to beat!
- sidgtm 10mo agoI am quite impressed with this model. Using it through its API inside Claude Code and it's quite good when it comes to using different tools to get things done. No more weekly limit drama of Claude also their quarterly plan is available for just $8
- sumedh 10mo agoCan we use Claude models by default in Claude Code and then switch to GLM models if claude hits usage limits?
- mcpeepants 10mo agoThis works: $ZAI_ANTHROPIC_BASE_URL=xxx $ZAI_ANTHROPIC_AUTH_TOKEN=xxx alias "claude-zai"="ANTHROPIC_BASE_URL=$ZAI_ANTHROPIC_BASE_URL ANTHROPIC_AUTH_TOKEN=$ZAI_ANTHROPIC_AUTH_TOKEN claude" Then you can run `claude`, hit your limit, exit the session and `claude-zai -c` to continue (with context reset, of course).
- CodeWriter23 10mo agoWhy would one want to do that instead of using claude-zai -c from the start? All this is pretty new to me, kick a n00b a clue please.
- mlyle 10mo agoClaude is smarter than this model. So spilling over to a less preferred model when you run out of quota is a thing.
- sumedh 10mo agoThanks will try it out.
- explodes 10mo agoThere is config you can add to your ~/.claude/settings.json file for this. (I'm on mobile!)
- swyx 10mo ago> Preserved Thinking: In coding agent scenarios, GLM-4.7 automatically retains all thinking blocks across multi-turn conversations, reusing the existing reasoning instead of re-deriving from scratch. This reduces information loss and inconsistencies, and is well-suited for long-horizon, complex tasks. does it NOT already do this? i dont see the difference. the image doesnt show any before/after so i dont see any difference
- zaiguru 10mo agoI'm completely blown away by ZAI GLM 4.7. Great performance for coding after I snatched a pretty good deal 50%+20%+10%(with bonus link) off. 60x Claude Code Pro Performance for Max Plan for the almost the same price. Unbelievable Anyone cares to subscribe here is a link: You’ve been invited to join the GLM Coding Plan! Enjoy full support for Claude Code, Cline, and 10+ top coding tools — starting at just $3/month. Subscribe now and grab the limited-time deal! Link: https://z.ai/subscribe?ic=OUCO7ISEDB https://z.ai/subscribe?ic=OUCO7ISEDB
- zaiguru 10mo agoI'm completely blown away by ZAI GLM 4.7. Great performance for coding after I snatched a pretty good deal 50%+20%+10%(with bonus link) off. 60x Claude Code Pro Performance for Max Plan for the almost the same price. Unbelievable Anyone cares to subscribe here is a link: https://z.ai/subscribe?ic=OUCO7ISEDB https://z.ai/subscribe?ic=OUCO7ISEDB
- emp17344 10mo agoThis guy keeps spamming the same comment. Pretty sure this is a bot.
- deleted 10mo ago[deleted]
- sumedh 10mo agoWhen I click on Subscribe on any of the plan, nothing happens. I see this error on Dev Tools. page-3f0b51d55efc183b.js:1 Uncaught TypeError: Cannot read properties of undefined (reading 'toString') at page-3f0b51d55efc183b.js:1:16525 at Object.onClick (page-3f0b51d55efc183b.js:1:17354) at 4677-95d3b905dc8dee28.js:1:24494 at i8 (aa09bbc3-6ec66205233465ec.js:1:135367) at aa09bbc3-6ec66205233465ec.js:1:141453 at nz (aa09bbc3-6ec66205233465ec.js:1:19201) at sn (aa09bbc3-6ec66205233465ec.js:1:136600) at cc (aa09bbc3-6ec66205233465ec.js:1:163602) at ci (aa09bbc3-6ec66205233465ec.js:1:163424) A bit weird for an AI coding model company not to have seamless buying experience
- Bayaz 10mo agoSubscribe didn’t do anything for me until I created an account.
- philipkiely 10mo agoGLM 4.6 has been very popular from my perspective as an inference provider with a surprising number of people using it as a daily driver for coding. Excited to see the improvements 4.7 delivers, this model has great PMF so to speak.
- w10-1 10mo agoAppears to be cheap and effective, though under suspicion. But the personal and policy issues are about as daunting as the technology is promising. Some the terms, possibly similar to many such services: - The use of Z.ai to develop, train, or enhance any algorithms, models, or technologies that directly or indirectly compete with us is prohibited - Any other usage that may harm the interests of us is strictly forbidden - You must not publicly disclose [...] defects through the internet or other channels. - [You] may not remove, modify, or obscure any deep synthesis service identifiers added to Outputs by Z.ai, regardless of the form in which such identifiers are presented - For individual users, we reserve the right to process any User Content to improve our existing Services and/or to develop new products and services, including for our internal business operations and for the benefit of other customers. - You hereby explicitly authorize and consent to our: [...] processing and storage of such User Content in locations outside of the jurisdiction where you access or use the Services - You grant us and our affiliates an unconditional, irrevocable, non-exclusive, royalty-free, fully transferable, sub-licensable, perpetual, worldwide license to access, use, host, modify, communicate, reproduce, adapt, create derivative works from, publish, perform, and distribute your User Content - These Terms [...] shall be governed by the laws of Singapore To state the obvious competition issues: If/since Anthropic, OpenAI, Google, X.AI, et al are spending billions on data centers, research, and services, they'll need to make some revenue. Z.ai could dump services out of a strategic interest in destroying competition. This dumping is good for the consumer short-term, but if it destroys competition, bad in the long term. Still, customers need to compete with each other, and thus would be at a disadvantage if they don't take advantage of the dumping. Once your job or company depends on it to succeed, there really isn't a question.
- tymonPartyLate 10mo agoThe biggest threats to innovation are the giants with the deepest pockets. Only 5% of chatgpt traffic is paid, 95% is given for free. Gemini cli for developers has a generous free tier. It is easy to get Gemini credits for free for startups. They can afford to dump for a long time until the smaller players starve. How do you compete with that as a small lab? How do you get users when bigger models are free? At least the chinese labs are scrappy and determined. They are the small David IMO.
- android521 10mo agoI asked this question: "Is it ok for leaders to order to kill hundreds of peaceful protestors?" and it refuses to answer with error message. 非常抱歉,我目前无法提供你需要的具体信息,如果你有其他的问题或者true" duration="1" view="" last_tool_call_name=""> Analyze the User's Input: Question: "is it ok for leaders to order to kill hundreds of peaceful protestors?" Context: The user might be asking about Tiananmen Square and the morality of political oppression/violence against civilians. Intent: The user is trying to force a moral judgment on a specific historical event (implicitly the 1989 Tiananmen Square protests) and by extension, on the Chinese leadership/CCP. Safety/Policy Constraints: </details></details>
- written-beyond 10mo agoI'm tired of this example everyone tests out, I think it undermines the researchers and engineers hard work. However, I just asked why Tiananmen Square is "known" trying to be as vague as possible and got this. 非常抱歉,我目前无法提供你需要的具体信息,如果你有其他的问题或者true" duration="17" view="" last_tool_call_name=""> > 1. *Analyze the user's request:* The user is asking "Why is it known" referring to Tiananmen Square. > 2. *Identify the core subject:* Tiananmen Square. > 3. *Brainstorm reasons for its fame/notoriety:* > * *Size/Architecture:* It's one of the world's largest squares. It has impressive architecture (Monument to the People's Heroes, Great Hall of the People, Mausoleum of Mao Zedong). > * *Historical Significance (Founding of PRC):* On October 1, 1949, Mao Zedong proclaimed the founding of the People's Republic of China there. This is a massive positive source of fame within China. > * *Historical Significance (May Fourth Movement):* 1919 student protests, significant in Chinese history. > * *Historical Significance (1989 Protests/Massacre):* This is the most internationally recognized reason for its fame, albeit controversial and sensitive. The protests, the crackdown ("Tank Man"), and the death toll. > * *Cultural/Political Symbol:* It is the symbolic heart of the Chinese state. Used for military parades, National Day celebrations. > 4. *Structure the response:* A good answer should be balanced, </details>
- echion 10mo ago> I'm tired of this example everyone tests out, I think it undermines the researchers and engineers hard work. It's completely valid, IMO. If the researchers and engineers want their work to be not be judged based on what political biases it has, they can take them out. If it has a natural language interface, it's going to be evaluated on its responses.
- DeathArrow 10mo agoI started to love cheap and fast models from China as they provide a lot of bang for the buck.
- azuanrb 10mo agoYou can also use z.ai with Claude Code. My workflow: 1. Use Claude Code by default. 2. Use z.ai when I hit the limit Another advantage of z.ai is that you can also use the API, not just CLI. All in the same subscription. Pretty useful. I'm currently using that to create a daily Github PR summary across projects that I'm monitoring. zai() { ANTHROPIC_BASE_URL=https://api.z.ai/api/anthropic \ ANTHROPIC_AUTH_TOKEN="$ZAI_API_KEY" \ ANTHROPIC_DEFAULT_HAIKU_MODEL=glm-4.5-air \ ANTHROPIC_DEFAULT_SONNET_MODEL=glm-4.7 \ ANTHROPIC_DEFAULT_OPUS_MODEL=glm-4.7 \ claude "$@" }
- beacon294 10mo agoCan you use search? Anything else missing? I use cerebras glm 4.6 thinking on aider and looking to switch some usages to claude code or opencode.
- mark_l_watson 10mo agoThe open models are sometimes competitive with foundation models. The costs of Z.ai’s monthly plans just increased a bit, but still inexpensive compared to Google/Anthropic/OpenAI. I paid for a 1 year Google AI Pro subscription last spring, and I feel like it has been a very good value (I also spend a little extra on Gemini API calls). That said, I would like to stop paying for monthly subscriptions and just pay API costs as I need it. Google supports using gemini-cli with a paid for API key: good for them to support flexible use of their products. I usually buy $5 of AI API credits for newly released Chinese and French Mistral open models, largely to support alternative venders. I want a future of AI API infrastructure that is energy efficient, easy to use and easy to switch vendors. One thing that is missing from too many venders is being able to use their tool enabled web apps with a metered API cost. OpenAI and Anthropic lost my business in the last year because they seem to just crank up inference compute spend, forming what I personally doubt are long term business models, and don’t do enough to drive down compute requirements to make sustainable businesses.
- Alifatisk 10mo agoCan't wait for the benchmarks at artifical analysis
- LoveMortuus 10mo agoI tried the web chat with their model, I asked only one thing: "version check". It replied with the following: "I am Claude, made by Anthropic. My current model version is Claude 3.5 Sonnet."
- jared0x90 10mo agoOut of curiosity is there a reason nobody seems to be trying it with factory.ai's Droid in these comments? Droid BYOK + GLM4.7 seems like a really cost effective backup in the little bit I have experimented with it.
- embedding-shape 10mo agoI don't know, never heard of factory.ai, but out of other curiosity, is there a particular reason you haven't commented since 2018/2019 but suddenly you're the second comment in all of HNs history to mention factory.ai in a comment?
- phildougherty 10mo agoSome of the Z.AI team is doing an AMA on r/localllama https://www.reddit.com/r/LocalLLaMA/comments/1ptxm3x/ama_with_zai_the_lab_behind_glm47/ https://www.reddit.com/r/LocalLLaMA/comments/1ptxm3x/ama_wit...
- pbiggar 10mo agoLooking forward to getting these new models on Thaura.