5 ms·
Z.ai confirms Ox Alpha is a new GLM-series model and will release its weights
- hncsiocp9x 1mo ago[dead]
- garo-pro 1mo agoUnfortunately I can't find sources other than this for now but this seems to be legit.
- mohsen1 1mo ago> The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News. Seems legit. It's really hard to know how good it is. So much hype around it.
- eli 1mo agoI mean, you can try it for free.
- KaseyKim 1mo agothey have confirmed it officially
- dgellow 1mo agoDo we know the size of the model?
- daveyoung 1mo ago[dead]
- j_maffe 1mo agoAnyone has a link to a report of its capabilities? I can't find a reliable source.
- vblanco 1mo agocompletely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range.
- daveyoung 1mo agolikely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled
- kristofferR 1mo ago63% at DeepSWE. https://x.com/davis7/status/2091285712566140986 https://x.com/davis7/status/2091285712566140986 Wenghi is behind DeepSWE, one of the best benchmarks.
- deleted 1mo ago[deleted]
- esskay 1mo agoI'd be interested to know what was going on with it during the public test as there were numerous reports of it improving considerably at tasks it was asked to do early on in the test compared to later in it.
- rfoo 1mo agolol don't shout out the obvious
- daveyoung 1mo agoTwo potentials from my pov: 1. Just variance in pass@K. If you prompt any model multiple times you'll see a large variance. N=1, but I find chinese open source models have a higher variance than higher-RL'd models like fable/opus. 2. They legitimately shipped a new RL checkpoint over the 7 days, which I find hard to believe. I am leaning towards 1.
- re-thc 1mo ago2. There was a new checkpoint. Official.
- daveyoung 1mo agodo you have reference to where it was said?
- zarzavat 1mo ago3. Deployment problems unrelated to the weights causing degraded performance
- swiftcoder 1mo agoFor sure the version accessible from OpenCode had a massive timeout problem the first day or so, which seemed to heavily degrade its task completion rate
- dannyw 1mo ago
- WithinReason 1mo agoMixed signals, here it's performing below even GPT-5.4 Nano: https://livebench.ai/ https://livebench.ai/ while here it outperforms Fable by a significant margin: https://oxalpha.com/ https://oxalpha.com/ but if the latter is true, will people still say it was "distilled" from Fable?
- deleted 1mo ago[deleted]
- sunbum 1mo agothe 2nd website is not official, just something someone slopped together for some reason.
- Alifatisk 1mo agoI have plenty of these websites, I can’t understand why someone is doing this.
- colesantiago 1mo agoIt is called phishing and grifting. Many people and even software engineers fall for this all the time. Most of these people are from crypto pivoting to AI doing this. AI has made this easier and cheaper and it is going to get a LOT worse. Imagine lots of websites with typosquatting and looking exactly the same as another website, vibe coded and cloned within seconds. The public have no chance.
- Alifatisk 1mo agoWhat is there to phish? These are simple vibe coded websites providing information for a certain topic, nothing else. In this case, that 2nd url is a website with information regarding the new model as well as a broken chat interface to try out.
- colesantiago 1mo ago
- tosh 1mo agomy guess is this is a small model punching way above its weight on toy benches it made quite a few mistakes but was able to fix all of them on its own (meaning more tokens, more turns, more tool calls — but same outcome as gpt 5.6 sol)
- daveyoung 1mo ago[dead]
- giamma 1mo agohttps://unwall.app/www.bloomberg.com/news/articles/2026-08-26/china-s-z-ai-made-ox-alpha-stealth-model-that-rivals-deepseek https://unwall.app/www.bloomberg.com/news/articles/2026-08-2...
- KellyCriterion 1mo agothanks for pointing me out on Unwall.App! Didnt know they exist - looks very good, maybe even better than Archive.ph
- seydor 1mo agoFunny how all china companies are expected to release weights by default
- respectattentio 1mo agothey are playing a completely different game than the US
- Aurornis 1mo agoAll smaller models and models behind frontier are expected to be released by default. Otherwise there’s no reason to produce them. Chinese labs are not releasing all of their model weights. Qwen is known as an open weight model by most, but their top model is not open weight. Releasing weights is a marketing strategy for newer labs to get their brand out there.
- cute_boi 1mo agoWell, i don't see demands from people to release chatgpt 4o.
- owebmaster 1mo agoIt's funny that you chose the exactly model that has a huge fanbase asking for it to be added back. ChatGPT 4o got a lot of people addicted. https://mashable.com/article/chatgpt-gpt-4o-ai-retirement-protest-rage-openai-reddit https://mashable.com/article/chatgpt-gpt-4o-ai-retirement-pr...
- respectattentio 1mo agoit's for sure better than deepseek flash 07/31
- mark_l_watson 1mo agoThat is saying a lot if Ox Alpha is also small and relatively cheap computationally. I hope so; I love deepseek-v4-flash-0731 and use it frequently. Fast inference is good and fits with my dev style: I like to be in the loop, not let an agent code on its own for long periods of time.
- SyneRyder 1mo agoFrom their blog post, it's 320B total parameters and 18B active parameters, so a similar size, but slightly bigger. Regular pricing is $0.15 input, $0.50 output... but currently 50% off, making it $0.075 input and $0.25 output. That beats most of the V4 Flash providers, but not all, and obviously tokens per task may not be equivalent. I've also just noticed the blog post reveals the Artificial Analysis score - it's a 57, so it's Opus 4.8 / 5.6 Terra level. https://z.ai/blog/glm-5.3-flash https://z.ai/blog/glm-5.3-flash
- sbinnee 1mo agoIt is not going to be cheaper though. I may choose the cheaper one in the end because performance will be marginal, both being flash.
- harlan_pdx 1mo agoReleasing weights is the right move. Keeps them competitive with DeepSeek on the open side.
- stanac 1mo agoI had good experience with GLM 5.3, but... Z.AI is the only provider for GLM 5.3 on OpenRouter. I don't see 5.3 on Hugging Face. Not sure if this new model is "full GLM" or something smaller, or if they will like Moonshot AI publish weights but put restrictive license [1], which will again leave Z.AI as single GLM model provider on OpenRouter. [1] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE
- cute_boi 1mo agoI am happy if they publish under restrictive license. Developing model costs a tons of money and company need to make money somehow by still open sourcing project.
- xienze 1mo agoIt's not released yet, just announced.
- birdboy1 1mo agoGLM 5.3 weights are not yet released
- jijji 1mo agoThe release date is supposed to be August 28th 2026
- glimshe 1mo agoThere's a lot of brand confusion among the Chinese models right now. Kimi, Qwen, GLM, Z.ai, Ox. We might know the difference (or I should say, someone does because I'm losing track already) but these models have no chance at end user penetration and loyalty until there's a single focused survivor. It took me a year talking about it until my wife knew that ChatGPT and Gemini are two different things. PS: some replies, especially if you do a deep dive on comment history, clearly expose the joint effort to drum up support for Chinese models. This has been clear on HN lately as anything even slightly critical of Chinese tech gets downvoted unnaturally quickly. One can just wonder what's behind the effort...
- giwook 1mo agoI disagree. I think most users who are savvy enough to be using openweight models and/or running models locally are not dealing with the same level of confusion you are. Ox is just GLM. And z.ai is the maker of GLM. The main players in the openweight model market have been known for a while. And they already have significant user penetration.
- tokai 1mo agoJust because you're confused doesn't mean that there is general confusion here. Its really not that complicated.
- marclove 1mo agoConsumers aren’t the customer.
- mark_l_watson 1mo agoI have seen studies from MIT and Stanford that the majority or US startups are using much less expensive open weight models so consumers of their products are open model users whether they know it or not. These are often Chinese models. Not to go off topic but I am pleased to see open model support from US companies like Poolside.ai, NVIDIA, IBM, Google, etc.
- seaal 1mo agoThere's a lot of brand confusion among the American models right now. ChatGPT, Claude, Gemma, OpenAI, Meta, Google, Muse Spark, Anthropic, Microsoft, Gemini. We might know the difference (or I should say, someone does because I'm losing track already) but these models have no chance at end user penetration and loyalty until there's a single focused survivor. It took me a year talking about it until my wife knew that Kimi K3 and GLM 5.3 are two different things.
- kosolam 1mo agoOnly reason people are interested is it’s free at the moment. I wasn’t impressed by its performance. Once the model gets a price tag it’s usage will be negligible.
- kosolam 1mo agoIt doesn’t rival deepseek v4 flash, and of course not deepseek v4 pro. This is my own impression.
- amritbir1 1mo ago[dead]
- esafak 1mo agoYou used it for visual tasks, right?
- segmondy 1mo agoThe price is free for some of us, we can run it at home.
- conradludgate 1mo agoHow much did you spend on hardware and electricity to run your free models
- garbageman 1mo agoLook, it's kind of like that 1982 HydraTech 16ft bass boat with a pearl glitter paint job thats on Craigslist. When the wife asks, you low balled it and they accepted - an offer too good to refuse.
- segmondy 1mo agoI have many rigs, but let's take 1 for example. 160gb. $1000 that's what it cost. 10 16gb MI50 gpus from ebay at $90. $900. Plugged them into an $100 octominer case from Facebook marketplace. I'm sure that doesn't satisfy you, keep coming up with excuses instead of finding ways to make this happen for you. You either find a way to get in and play or you sit on the sideline and moan about those in the field.
- fen_wick 1mo agoGood to see more competition in the open weights space. The more players the better.
- stingraycharles 1mo agoBut they’re not at all a new player.
- freakynit 1mo agoIt one-shotted generation of Java bindings for this project: https://github.com/jeffhajewski/latticedb https://github.com/jeffhajewski/latticedb Related PR: https://github.com/jeffhajewski/latticedb/pull/5 https://github.com/jeffhajewski/latticedb/pull/5 The session used ~100K input tokens, ~60K output tokens, and ~80K thinking tokens. I reviewed it using gpt-sol-medium, and it seems to be satisfied with it's work.
- hypfer 1mo ago> The company on Wednesday confirmed speculation that the Ox Alpha model is a new iteration of its GLM series and said it will release the weights for it tonight, in response to queries by Bloomberg News. Where? And "Tonight" in which timezone?
- tokai 1mo agoSingapore I would assume. Z.ai usually peg everything to Singapore time.
- a012 1mo agoChina is GMT+8
- NitpickLawyer 1mo ago> in which timezone? Apparently someone working at a 3rd party inference provider also got confused and posted confirmation about it being a glm-flash model, despite having an embargo on that info. Someone jumped in the comments and told them they missed the timezone :) In any case it should be releasing in a few hours. Timezones are hard.
- ricardobeat 1mo agoI had Ox Alpha working on coding tasks for a couple days non-stop, via OpenRouter and OpenCode Zen. Crush harness. It was able to complete tasks at a level that I'd put between Sonnet and Opus. It makes few mistakes, but is not that smart. The main issue for me, is that it degraded into a doom loop several times. One of them was running the same bash command about a thousand times. The last model I've used that had this problem was Mimo 2.5, which is quite dated at this point. As a result of this, you cannot leave it unattended / not usable for agents.
- cyanydeez 1mo agoI usually see doom loops when working with quants. Likely theyre trying to maximize the viability of a efficient model quant that can bw upgraded. Like cutting coke to get crack, quantiry over quality.
- echelon_musk 1mo agoNit pick; cutting (adulterating) cocaine doesn't produce crack cocaine.
- redox99 1mo agoOx alpha at moments felt like it was quantized to hell. I think the last few days it might have improved.
- knuckleheads 1mo agoI couldn't get past all the network errors on OpenCode. Seemed smart enough, and was useful when I was low on usage on Claude, but beyond that, really hard for me to say whether it was Good or Bad.
- SyneRyder 1mo agoRather than a pelican, for fun I showed it a couple of screenshots from Niu Lai and asked it to create an SVG inspired by the images. I explained a little about how the movie had been made by a mother & son team, initially derided but then went on to surprise cult box office success. It came up with this: https://x.com/syneryder/status/2091978367579156569/photo/1 https://x.com/syneryder/status/2091978367579156569/photo/1 Created in a single turn - but technically not a "one-shot", because I gave it a tool to convert SVG to PNG so it could visualize what it had made. I asked it to keep iterating with tools during the same turn until it was happy. I've also been using Ox Alpha for tasks that better resemble real work, and I'm really enjoying working with it. I've downgraded my Anthropic account so I can put some budget towards Ox Alpha instead, with the rumors that this one is going to be cheap. Opus & Fable are still better at getting large tasks / features done autonomously, but Ox Alpha can work autonomously too, and it's fun. I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models. (As much as I don't want to say that, as someone with Claude /stickers on their laptop.)
- deleted 1mo ago[deleted]
- netniuq 1mo ago> …and it's fun. I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models hard agree. it does not really feel "smart", but the personality is super refreshing
- Aurornis 1mo ago> I'm enjoying working with Ox in a way that I'm just not enjoying talking to the 5.0 Anthropic models. That’s very valid, but right now every other model I use is easier to talk to than Opus 5.0 Opus 5.0 has an impenetrable way of communicating. I can parse it, but it takes so much more work than it should.
- SyneRyder 1mo agoYeah, that's a fair point. "More intelligible than Adriano Celentano in Prisencolinensinainciusol" is not a high bar. As another comparison, I went back to MiniMax M3 for a while last night. It was significantly faster than Ox, but I felt M3's replies were harder to parse, not quite getting to the point. But I guess I could curb that with some prompts. It depends if the Ox Alpha pricing is as cheap as was being rumored. If it's competitive with DeepSeek Flash and significantly undercutting Luna, that feels like it will be significant.
- xbmcuser 1mo agowill we reach the singularity once the llm can be used to program the llm?
- syntaxing 1mo agoI’m more curious on the size. If it’s smaller than or equal size to GLM 5.3, this would be a crazy good model. If it’s closer to deepseek pro, it would be a good model. If it’s near Kimi K3, I think it’s competitive but nothing particularly differentiating.
- SomeonesAccount 1mo agoDefinitely agree. If it is small (eg. Qwen 3.8 28b or gpt-oss-120) then this might be amazing. If it is anywhere near Kimi K3 it would need to have some other differentiating factor than intelligence.
- m00dy 1mo agoYeah, it was identified as a GLM-series model quite a while ago. You can also check out this AI model fingerprinting resource [0]. [0]: https://openrating.io/blog/current-state-of-ai-model-fingerprinting https://openrating.io/blog/current-state-of-ai-model-fingerp...
- ThouYS 1mo agoCalling it now: The big deal about this model is the sheer volume they were offering through openrouter and OpenCode. How? Chinese AI accelerators / nvidia-free stack
- redox99 1mo agoOx alpha is better at UI than GPT 5.6 Sol. Not a high bar considering Sol sucks at UI, but as someone who just has a codex sub, I've used almost 1B tokens of ox alpha these last few days to complement Sol smartness. Inference was atrocious in terms of speed and constant timeouts. If it's served fast it will be a delight to use.
- lampcord 1mo ago[dead]
- _pdp_ 1mo agoOx Alpha has been running on auto-pilot for the past 5 days on various experiments. Very impressive model. Here are some examples, open-source documented and the data available in HF datasets: https://openzot.github.io/whetstone/ https://openzot.github.io/whetstone/ - https://github.com/openzot/whetstone https://github.com/openzot/whetstone https://openzot.github.io/arcade/ https://openzot.github.io/arcade/ - https://github.com/openzot/arcade https://github.com/openzot/arcade https://openzot.github.io/machinery/ https://openzot.github.io/machinery/ - https://github.com/openzot/machinery https://github.com/openzot/machinery
- 13639366668 1mo ago[flagged]
- yipinwong 1mo agoOx Alpha was working good for me but I do not it if a trend starts where openrouter hides where the traffic is going to.
- RataNova 1mo agoIt writes pretty clean code and holds context alright, but it starts stumbling and losing the plot on complex bash scripts with pipelines. Waiting for the weights to drop so we can dig under the hood and see what is going on there
- danieltk76 1mo agoto be honest I found it underwhelming.
- amathur2k 1mo agoWhich harness are u folks using, I have tried opencode and claude code. Both absolutely keep hanging due to the model running into loops and becoming unavailable. Unable to do even simple things
- codybontecou 1mo agoI happily ran it in Pi without issue.
- itsryanlenk 1mo agoLooking forward to seeing the stats. I gave it an abandoned repo for an Aseprite MCP someone made and told it to iterate with a laundry list of things I wanted from it to include thousands of plugins. Came back 20 hours later and it shit out a pretty surprising little tool, will post the public repo when I get time.
- garo-pro 1mo agohttps://z.ai/blog/glm-5.3-flash https://z.ai/blog/glm-5.3-flash
- ashing 1mo agoI'm going to use this model hard.
- akshay_akula 1mo agoReads like the usual pattern, somewhere between Sonnet and Opus but not something you leave unattended. Curious if the weights release changes that.
- keheai_harvey 1mo ago[flagged]
- fenderblender 1mo ago[dead]