17 ms·
GLM-5.3-Flash
https://news.ycombinator.com/item?id=49450353 https://news.ycombinator.com/item?id=49450353
- tianji-astrolog 2mo ago[flagged]
- rahimnathwani 2mo agoRelated: https://news.ycombinator.com/item?id=49446422 https://news.ycombinator.com/item?id=49446422 (281 points, 118 comments)
- iamsyr 2mo agoStandard API Pricing for GLM-5.3-Flash (per 1M tokens) - Input: $0.15 - Output: $0.50 - Cached input: $0.03
- Xunjin 2mo agoIs that cheaper than DS4 flash?
- javier123454321 2mo agoAll I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
- denysvitali 2mo agoTbh it was also slow because it was being hammered by everyone making use of the free tokens
- javier123454321 2mo agoPossibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.
- swiftcoder 2mo agoIt's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now
- arizen 2mo agoFew weeks ago, I wouldn't expect this statement to be true. Accelerate!
- nateb2022 2mo agoSlightly more expensive than the (post-price hike) DS4 flash pricing, but in the ballpark. https://openrouter.ai/compare/deepseek/deepseek-v4-flash-0731/z-ai/glm-5.3-flash https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...
- walrus01 2mo agoComparison should be to 0731
- nateb2022 2mo agoupdated thanks
- drob518 2mo agoHm. GLM is more expensive in all dimensions than DS but it has a lower weighted average input? How is that?? Something seems off. EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.
- desterothx 2mo agoI mean, it implies that it has an even better cache hit rate that DS flash, which is impressive, as the chr on DS flash was already really good in my experience
- nateb2022 2mo agoI think the weighted average takes into account all providers (some DS4 flash providers are 'premium' providers and offering higher speeds for higher pricing) and these are tilting the scale
- deleted 2mo ago[deleted]
- epolanski 2mo agoI'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field. It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.
- esperent 2mo agoThis has been clearly stated as what would happen going back several decades at least.
- ricardobeat 2mo agoStarting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.
- himata4113 2mo agoWell the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.
- nananana9 2mo agoThat's how you catch up when you're behind. Now the US is behind in EVs can you guess what they're doing? [1] [1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly-chinese-evs-to-detroit-to-study-them https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...
- himata4113 2mo ago"argument is very weak" regardless as I said.
- cyanydeez 2mo agoyeah, America is totally out there respecting international law. "problem" indeed.
- sunbum 2mo ago> with all of this traffic served on Chinese AI chips RIP Nivida shareholders
- ChoosesBarbecue 2mo agoGod I wish I could’ve shorted NVIDIA right now
- browningstreet 2mo agoIt's earnings day for them...
- kingstnap 2mo agoWhats stopping you? You could buy puts right now. Get a 210 strike put contract and if your thesis is that nvidias current 10 day slide continues you could make some money.
- outworlder 2mo agoUnless NVidia craters you are likely to lose money given the IV crush that will happen today.
- mmastrac 2mo agoWeights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash https://huggingface.co/zai-org/GLM-5.3-Flash I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through. I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level. I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.
- kilroy123 2mo ago> get myself four sparks at a decent price Wow, if you don't mind me asking. How and where?
- mmastrac 2mo agoI bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely. They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.
- swiftcoder 2mo ago> it's the only model in the whole lineup that isn't priced insanely $4,000 isn't priced insanely? ye gads
- a3w 2mo agoI thought 4000 in sum. No wait, 4000 per, plus tax. Or EUR pricing to similar accord. Ouch.
- packetlost 2mo agoFor those who didn't read, this is the identity of the mysterious "Ox Alpha" model
- Bluestein 2mo agoThey even give this over the API now: │ https://openrouter.ai/api/v1/chat/completions https://openrouter.ai/api/v1/chat/completions model: stealth/ox-alpha auth: OPENROUTER_API_KEY status: 404 Not Found response: {"error":{"message":"Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash. │ Use it now: https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_id https://openrouter.ai/z-ai/glm-5.3-flash","code":404},"user_...":"}
- AbsurdCensor 2mo agoYeah, made me suspicious of how well the Ox Alpha was performing that it wasn't some 'new group' making the model.
- revolvingthrow 2mo ago> 320B total parameters and just 18B active parameters This is pretty hefty for a "flash" model, even a 256 GB setup is insufficient at q4 - and q4 is already the worst-but-still-acceptable quant in my experience. The benchmarks look great, especially since GLM tends to be more honest than the average Chinese lab, but you’ll need to splurge to run it at home. @edit: so many releases that I forgot to math. This fits just fine in q4, realistically the minimal hardware would be 192gb - so blazing fast on double rtx 6000 pro and usable on 256gb unified memory. You could even go with 5bit quant on 256gb. … you’ll still need to splurge, though.
- colingauvin 2mo agoThat's 160GB-ish for Q4...how is 256 insufficient?
- dannyw 2mo agoLooks like the M5 Ultra Studio wait times are going to increase again. Already at 10-12 weeks, I wonder how long it'll go?
- speedgoose 2mo agoI guess like the M3 Ultra, at some point normal customers won’t be able to buy it.
- Destiner 2mo agofrom the article, pareto frontier for open source models is completely dominated by GLM now.
- montroser 2mo agoWell, it will be interesting to see where Qwen3.8-Flash-Next ends up landing, also released today. These are exciting times!
- Lalabadie 2mo agoI find GLM's idea of fast/flash is not really competitive with the speed DS4 Flash has, and it's hard to see them as being in the same segment for that reason.
- knollimar 2mo agoEven vision? Thought k3 might have an edge there
- TaLiTr 2mo ago> it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. From a biased source, but would be big if true. I've had great results with GLM 5.2. From their subscription page, the smallest plan gives you about 97M tokens weekly for 5.3 but 292M for 5.3 Flash. Not exactly 10x the limit.
- re-thc 2mo ago> From a biased source, but would be big if true. I've had great results with GLM 5.2. It's at least close (even if not better) from the Ox Alpha runs. For the price it's definitely great.
- wolttam 2mo agoThe recent and slightly smaller DSv4 Flash is also GLM 5.2 equivalent (or close enough)
- tokai 2mo agoDSv4 hallucinates much more than GLM-5.2 though.
- mariopt 2mo agoIt's only 320B, local frontier AI is getting closer, sooner than expected.
- oceansky 2mo agoCan't come soon enough!
- saberience 2mo agoIt's not possible to keep shrinking down parameters and keep "frontier" performance, it's like saying it's possible to take a 3 hour movie and compress it down to 3 megabytes, there are information theoretic limits on the amount of bits of information that can be compressed. What I'm saying is, if you're expecting a model that can be run on a 16GB or 32GB machine with the intelligence/knowledge of Mythos or Sol, it will never happen. It cannot happen, just like you cannot watch the Odyssey saved as a 16MB file. Smaller models can get faster and smarter, but by definition they can never compress all of the knowledge of a frontier model and they will approach a limit by which they cannot get better.
- deleted 2mo ago[deleted]
- hypfer 2mo agoYou can fit Shrek 1 into a 13mb gif tho https://www.deviantart.com/sssfjknfvdknj/art/the-ENTIRE-shrek-movie-as-a-gif-1283041952 https://www.deviantart.com/sssfjknfvdknj/art/the-ENTIRE-shre...
- twobitshifter 2mo agoThe current models are not close to approaching the limit of compression for intelligence. They aren’t even focused on it like Chinese labs are. The training of Qwen’s 27B parameter model showed that by structuring model training from fundamentals to more difficult topics they were able to drastically reduce the number of parameters needed. The ‘frontier’ models rely on scale to achieve their results but that’s not the only approach. Eventually we will hit up against the fundamental limits but we are not close with Sol and Mythos.
- garo-pro 2mo ago> Combined with our latest 30T-token multimodal pre-training corpus [...] Is the optimal formula still 20x the amount of model params in tokens for training? Could this mean we're getting a GLM with 1.5t params?
- swingboy 2mo agoHow much is the “discounted” pricing they mention?
- xena 2mo ago50%
- Imustaskforhelp 2mo ago> To overcome the relatively limited compute and memory capacity of individual chips, we built a dedicated inference engine for this architecture on top of SGLang. Notably, this effort was accelerated by our GLM-5.3-powered infrastructure agent, which assisted engineers in developing and optimizing kernels, diagnosing performance bottlenecks, and improving the serving stack — creating a feedback loop in which the model helped optimize the system serving the model itself. > (...) Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale. It might be one of the most actually practical tasks that AI might've done because the compounding effects of it and also its implications are/feels so immense. It feels as if Nvidia might be in a slight turbulence from it.
- freakynit 2mo agoNow translate this to physical world, robots building and optimizing other robots... getting iRobot (2004) vibes
- toppy 2mo agoBy clicking this link you download some PDF in the background
- krystofee 2mo agoIts displayed in the html...
- deleted 2mo ago[deleted]
- smilingPanda 2mo ago[dead]
- kayleykiwi 2mo agoThis looks like it goes hard, can't wait to try it
- yipinwong 2mo agoWhen reading this type of announcements, always have keen eyes on graphs. e.g. "Agent Coding Performance by Effort Level" cuts Y-axis from 0~20. - This makes it as if GLM-5.3-Flash made a bigger jump than it claimed as the Y-axis does not increase much (stupid trick used in biz reports) I did mention that ox was working ok for me, and having an open-weight comparable to close to SOTA makes it very compelling for me to try it out locally (well, only if I got more VRAM)
- nchmy 2mo agothey also conspicuously omitted GPT 5.6 Luna from comparison. It scores lower, but is also cheaper. MiMo 2.5 is not a valid comp at this point edit: nevermind. it is there in the artifical analysis scatter plot, but is greyed-out. MUCH more interesting is that in that chart, their cost is WAY off. The actual chart shows GLM 5.3 Flash at $0.09, but their chart shows $0.045...
- mrtesthah 2mo agoThe web page says 5.3 flash is discounted right now.
- drob518 2mo agoSeems disingenuous to draw frontier graphs with starter pricing.
- seaal 2mo agoWell, Luna debuted with 5x higher pricing than is currently available. With the pace of recent development these models might not be relevant by Thanksgiving.
- drob518 2mo agoOf course. Pricing is always changing, but typically it goes down over time, not up. So, if you're showing artificially low pricing from the start based on a teaser rate, IMO, you shouldn't be using that to show where you appear on a frontier graph. Place yourself on the graph based on your expected long-term pricing. Then, over time, adjust your position based on your standard rate, whatever that might be. Games are always being played for things like this, but this seems excessive.
- cootsnuck 2mo agoIf we fast forward say 5 years, I don't see how we don't end up in world where people (and enterprises) are more savvy with how they use LLMs. Meaning, more models, smaller models, weirder models, more specialized models, etc. And all of it running on a variety of hardware (edge devices, personal computers, on-demand cloud compute). I don't see how NVIDIA can keep their spot as belle of the ball. If LLMs and friends are truly to become as useful and ubiquitous as everyone thinks they will, then commoditization is the only option.
- drob518 2mo agoWe need to figure out what the real pricing is for a going concern. Right now, everyone is subsidizing and discounting to grow (or maintain) market share. The big question is whether the steady state, market derived inference pricing is above or below what we’re seeing today. I honestly don’t know. Anthropic had said that inference is profitable, but they’re clearly not yet profitable overall with training and buildouts still happening.
- apitman 2mo agoHas nobody from any of the companies hosting open weights models released detailed information on how much it really costs?
- drob518 2mo agoI’m sure someone does, but I’ve never seen anything other than vague statements like Anthropic’s “inference is profitable” comment. I suspect everyone is playing everything close to the vest because they aren’t yet public and they want to control the information flow to the street.
- bigyabai 2mo ago> I don't see how NVIDIA can keep their spot as belle of the ball. FWIW, people were saying "ASICs will kill CUDA demand!" since the crypto mining boom. Then a few months later, CUDA found another niche application in LLM applications. With the mounting demand for robotics, surveillance and autonomous weapons, I don't see how Nvidia couldn't keep their spot. They have their pick of the litter with hundreds of market segments, and unlike the rest of FAANG they're not afraid to branch out.
- ammmw 2mo ago[dead]
- tokai 2mo agoWhy is their own coding plan always the last place z.ai release their models? Its even online, you just have to guess the model settings.
- woadwarrior01 2mo agoCaptive audience.
- mrngld 2mo agoChinese labs are so used to manipulating benchmarks to try to flatter inferior models that when they finally have one that's really pretty good I think the official announcement here undersells it. https://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/ That's pretty solid. Smarter and cheaper than Luna xhigh, not as smart but less expensive than Luna max. Smashes deepseek v4 flash, and even worse it matches v4 pro at a tiny fraction the cost. Roughly equivalent to sol medium, at a fraction the cost. They should've just lead with real, up to date data, because it's good, not the silly old tactics like comparing to Opus 4.8 when 5.0 is out in many of their charts. Congrats to them!
- stavros 2mo agoOpus 5 is better than Fable in this benchmark?
- zarzavat 2mo agoEven Artificial Analysis has Opus 5 better than Fable in their aggregated "Intelligence Index" which combines 9 benchmarks. Opus 5 is heavily benchmaxxed.
- nijave 2mo agoDepends on the benchmark but yes. I think Opus is more heavily optimized for coding. On the usability side, its output is almost intolerable to read. It seems to code fairly well. Fable is more enjoyable to use for planning/interacting with
- seaal 2mo agoOnly 73K output tokens too. Anthropic should really be embarrassed with their Sonnet 5 price/performance.
- redox99 2mo agoIt's also better than Sol (at whatever effort) at designing pretty UIs. I have a Codex sub and I've been using this model for UI stuff.
- deleted 2mo ago[deleted]
- AnodicElegy 2mo agoArtificial Analysis benchmark is out: https://news.ycombinator.com/item?id=49450353 https://news.ycombinator.com/item?id=49450353
- claudeIsDown 2mo agoOn OpenRouter the pricing is: Input $0,075/M - Output $0,25/M - Cache Read $0,015 /M How is the business model of Anthropic/OpenAI will sustain?
- dakolli 2mo agoThey're obviously in a pickle, nobody is going to continue to pay $15-50 a mm tokens here soon. There's a reason OpenAI stopped training large models last week, and it's not because of "saftey" or "alignment" they know these gigantic models are not worth the squeeze.
- polski-g 2mo agoThis is a bad model. Worse than Luna in every way; slower, dumber.
- pphysch 2mo agoOAI/Anthropic shareholder? Speed and intelligence are not "every way". Cost is essential. Hence the Pareto boundary illustrated in TFA.
- polski-g 2mo agoIt literally cannot complete tasks that Luna can do easily. It doesn't matter how cheap it is.
- freakynit 2mo agoIt actually can. I have been using it regularly for past 4 days. It is on sol-low level. I deliberately tested it on a moderately complex task. glm-5.3-flash one-shotted it correctly. Luna max couldn't achieve parity even after 3 total attempts. It was creating java bindings for this project: https://github.com/jeffhajewski/latticedb https://github.com/jeffhajewski/latticedb And here is the binding one-shotted by glm-5.3-flash: https://github.com/jeffhajewski/latticedb/pull/5 https://github.com/jeffhajewski/latticedb/pull/5
- 2mo ago
- matheusmoreira 2mo agoYou guys read Z.ai's terms of service, right? Broad and perpetual license over inputs and outputs, and even your name and profile picture. Vague prohibitions on whatever may harm Z.ai’s "interests" or even the "national interests" of any country. Vague prohibitions on "disturbing" or "inappropriate" content, whatever that is. Vague prohibitions on discussing Z.ai, even my posting this comment violates it. Can ban you if you, in the "sole and absolute opinion" of Z.ai, have violated these broad terms, and if you paid for the discounted yearly plan kiss your money goodbye.
- culi 2mo agoyou can abliterate any open model like this. This is pretty standard stuff in a TOS. I'd be surprised if you couldn't find the same in OAI or Anthropic's
- Lwerewolf 2mo agoThe model weights are MIT licensed.
- matheusmoreira 2mo agoI'm talking about the Z.ai service specifically.
- Lwerewolf 2mo agoAs others have mentioned, nothing's stopping any other major provider from offering it. Given its popularity, you can guess how that'll develop. So, overall, irrelevant.
- g3f32r 2mo agoIsn't this practically every TOS though? Nearly every TOS I've ever read has a "We can ban you for any reason, or no reason, are under no obligation to disclose any reason." line somewhere in it. HN's for example > We reserve the right, at our sole discretion, to change or modify portions of these Terms of Use at any time. > You acknowledge that Y Combinator may establish general practices and limits concerning use of the Site, > You further acknowledge that Y Combinator reserves the right to change these general practices and limits at any time, in its sole discretion, with or without notice. > Y Combinator reserves the right to investigate and take appropriate legal action against anyone who, in Y Combinator’s sole discretion, violates this provision, including without limitation, removing the offending content from the Site, suspending or terminating the account of such violators and reporting you to the law enforcement authorities.
- lxe 2mo agoIs the actual Z.AI ecosystem good enough to replace the main drivers like Codex and Claude? Because it looks like Z Code is just a Codex fork. Just like the Kimi Code one is. What irks me about this is that the harnesses seem to be just an afterthought here. Don't get me wrong, I love messing around with installing Pi, getting it hooked up with OpenRouter, and just trying all kinds of different stuff, local models, etc... but when it comes to literally just setting up a productivity environment and trusting my entire machine with it, I just run Codex. I have heard from anecdotes where people have indeed replaced their main drivers with DeepSek V4 Flash or GLM and state that "it's almost as good as... [claude/gpt]" but I never hear anyone say "yeah, this is the model/harness that I now run on my machine and don't mess with it"
- Bluestein 2mo ago> "yeah, this is the model/harness that I now run on my machine and don't mess with it" * me raises hand.-
- Havoc 2mo agoTheir list of allowed tools is extensive so just use whatever you want within that list Think z code gives a token bonus though
- ygouzerh 2mo agoYou can use OpenRouter directly in Claude Code as well, it's quite nice!
- Sphax 2mo agoBoth can be true though. I had the max coding plan since january and I kept using with Pi since then, even though it wasn’t as good as opus until glm 5.3. It definitely can be a daily driver if you don’t want to use Anthropic or OpenAI. It’s going to be even better with native vision now available. And i’m not messing with my setup either.
- computerex 2mo agoI use my own harness: https://github.com/computerex/z https://github.com/computerex/z Have been using it as my primary harness for personal work for I'd say 6 months. I recommend everyone create their own harness at least to learn. There are a lot of practical benefits.
- dakolli 2mo agoI didn't accept a single edit from this model over the entire week, just saying. I do not understand how it's being benchmarked on par with Sol and other larger models.
- respectattentio 2mo agois it a benchmarkmaxxing model?!
- jazzpush2 2mo agoIt was certainly almost RL-fried to overfit the benchmarks, at the expense of actual usability. See Opus 5.
- kburman 2mo agoofftopic: Is there any chance we could see competing models from other countries in the next 5 years?
- svachalek 2mo agoChinese universities are really a huge advantage, even in the US many of the top staff in model development are Chinese. Another big thing is the hardware costs required to train models. Between those two factors it really looks like this will remain a US-China competition for the foreseeable future, although there are some other players like Mistral from France.
- deleted 2mo ago[deleted]
- tinyhouse 2mo agoAnthropic is accelerating their IPO cause they know what's coming in the next 5 years.
- knowaveragejoe 2mo agoAny providers hosting it outside of China?
- svachalek 2mo agoI don't see anyone other than ZAI yet but GLM 5.2 is available on many providers worldwide so I'd expect we'll see the same on this one soon.
- xena 2mo agoRight now there's at least two: https://openrouter.ai/z-ai/glm-5.3-flash https://openrouter.ai/z-ai/glm-5.3-flash Give it a day or two. More will pop up.
- scottfits 2mo agoso is it confirmed if this is the mysterious OxAlpha model?
- Gander5739 2mo agoYes; if you try to use Ox Alpha it will give an error saying it waa trial period, and that it is GLM 5.3 flash.
- scottfits 2mo agointeresting, thanks!
- singularity2001 2mo agoAt the current 50%-off GLM-5.3-Flash price ($0.075/M input, $0.25/M output; cached input $0.015/M), surprisingly, roughly $400–900/month would buy token throughput comparable to fully exhausting Claude Max 20×
- arizen 2mo agoProbably apples to apples would be to compare z.ai subscription plans vs API pricing
- pietz 2mo agoWith tiny models surpassing huge, 6 months old models on benchmarks, does anybody have some smart words to share on how these still "feel" different? Artificial Analysis ranks GPT 5.6 Luna similar to GPT 5.4, but that never matches my real world experience. AA seems to do a good job making a single number as representative as possible but there is still so much benchmarks don't communicate.
- KptMarchewa 2mo agoI agree. They are definitely good - no issues with instruction following for example - but they miss the "intelligence" larger models have. For implementation tasks, where I have the problem already defined and researched, or just simple task, I'd definitely use something like Luna xhigh or max. If the task is vague, or involves planning, I'd rather use Sol medium, even though it's theoretically worse on benchmarks.
- fridder 2mo agoCould do a "Big model for architecture and planning and smaller model (or local model) for implementation" sort of thing
- Kungfuturtle 2mo agoOne term of art that's emerged for this "feel" is "big model smell", first coined by @aidan_mclau. [0] To my surprise I couldn't find any proper explainers of the term in a quick search, despite grokking it after seeing it in various contexts on Twitter, but Fable 5 offered a useful analogy: "A student who memorized worked solutions and one who understands the subject score the same on the test; you can only tell them apart by asking a question the test didn't. Real-world use is nothing but those questions, which is why a single AA number feels right and wrong at the same time." In other words, big model smell is related to the underlying ability to "understand" when tasks are underspecified or out-of-distribution. This ability can be mimicked to parity by smaller, distilled models according to the density of the training data for particular tasks, but neural scaling laws still hold for generalized reasoning ability. More recently with these smaller models, there's a separate but related "RL-fried" phenomenon, where they rely on CoT to "grind toward a checkable answer even in contexts (open dialogue, taste, ambiguity) where there is no checkable answer, and you get the tell: over-hedged, over-structured, relentlessly on-task, deaf to the subtext." There are some other insights and caveats in the (short) conversation that I feel you may appreciate reading. [1] [0] https://x.com/aidan_mclau/status/1807843014104211855 https://x.com/aidan_mclau/status/1807843014104211855 [1] https://claude.ai/share/d511a348-7c36-432f-a6d5-9deab2802615 https://claude.ai/share/d511a348-7c36-432f-a6d5-9deab2802615
- jdw64 2mo agoThis was the ox-alpha model, right? I remember it performed really well for a model that had 'flash' in its name.
- BeetleB 2mo agoThe key difference between this and all other GLM models is it's multimodal. You cannot send images to the other GLM models.
- mrinterweb 2mo agoI really wish GLM models had vision capabilities. I've worked around that in the past to use a vision MCP in my harness that GLM can call. It is not the same, but it allows the model to query images.
- BeetleB 2mo agoWell, now one of them does!
- mrinterweb 2mo agoThat's wonderful. I was going off an older version of the Artificial Analysis page for GLM-5.3-Flash https://artificialanalysis.ai/models/glm-5-3-flash https://artificialanalysis.ai/models/glm-5-3-flash. The page is updated now to show that it does support multi-modal image inputs.
- hxii 2mo agoIn my brief testing, it did about as well as Qwen3.8-4B-Distill, and LFM2.5-2.6B overtook both.
- bertili 2mo agoThis is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third! 12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
- jatins 2mo agoexcept besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience
- computerex 2mo agoYou'll find that hard to prove objectively and conclusively.
- deleted 2mo ago[deleted]
- rxyz 2mo agoOpus 5 is the least reliable frontier-class model in the market
- usef- 2mo agoIn what way? It has worked well in my experience. It holds up with long context windows, unlike many, too.
- pimeys 2mo agoI have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want. This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch... Edit: yes company money. We don't get subscriptions we pay per token.
- jatins 2mo agoI was quite surprised that Zai had deep pockets to serve this free for a week. My first guess was this was an American lab like xai or google
- yousif_123123 2mo agoWill we need all the data centers being built or will improvements in software and hardware allow the majority of AI workloads to run locally or in the cloud but way more efficiently than was projected when all the plans were laid out? Like were executive at Google and AWS and Microsoft expecting this kind of performance from models smaller than what openai/anthropic have been doing? Are we really in a "compute desert"?
- bakies 2mo agoIf it gets more efficient it'll be more enticing to expand use case. Personally I'm hoping to do a lot at home but I'm not counting the datacenter building as a bad move at this moment. It may and up that way.
- VirusNewbie 2mo agoIt looks like gemini 3.7 flash actually beats it in a lot of benchmarks, no? https://x.com/Zai_org/status/2092616204787626030/photo/1 https://x.com/Zai_org/status/2092616204787626030/photo/1
- syntaxing 2mo agoIronically, our administration pushing for ban of the AI chips to China is forcing them to make smaller and more efficient models which seems like a requirement for running on Chinese chips. I wouldn’t be surprised this model was tailored to run purely on Chinese chips. Same thing with Deepseek MLA, the drastically lower KV cache memory requirement was born out of necessity so it runs on the Huawei chips.
- adroitboss 2mo agoThis is exactly what Jenson said in all of his interviews. Banning it in the short term would have long term consequences.
- deleted 2mo ago[deleted]
- dzonga 2mo ago> Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips. Just like that we are witnessing an open burial. It's now in everyone's interest to keep the valuations in the 'A.I' economy as they're though it's apparent they're not justified. whether it's the cost to develop models, cost of hardware, cost of serving ie inference.
- micimize 2mo agoWeren't they giving free access? Not exacty a meaningful heuristic if so
- disiplus 2mo agoit was not the first model served for free, i remember grok and others beeing free on openrouter but they never had this popularity because they where not good enough.
- mtrovo 2mo agoThe key insight here is being the top used model on opencode while being fully served on Chinese chips. The free price itself might be just a flex or marketing budget.
- beannt 2mo agoIs it good compare to Opus 5 ?
- freakynit 2mo agoPersonal testing results: it's on gpt-sol-low level.
- XCSme 2mo agoNice, finally they fixed the huge reasoning tokens count. Now it's similar cost to DeepSeek v4 flash, but smarter. My tests: https://aibenchy.com/compare/z-ai-glm-5-3-flash-max/deepseek-deepseek-v4-flash-0731-high/qwen-qwen3-8-27b-high/z-ai-glm-5-3-high/ https://aibenchy.com/compare/z-ai-glm-5-3-flash-max/deepseek...
- preommr 2mo agoSo the vagueposting by googlers about Ox Alpha was just... what exactly? Like I get that they have to be careful about comms, but surely senior members of the team can clarify when something is NOT them, when everyone is gosspiing it is them.
- qeternity 2mo agoTrolling. GLM is heavily distilled from Gemini.
- bel8 2mo agoSource? GLM is great for coding and Gemini is barely useful in coding, to be generous. I highly suspect that the Gemini Google uses internally is very different from what they offer in Antigravity.
- spijdar 2mo agoI can't speak for GLM as I haven't tested it much, but my experience with running DeepSeek V4 locally is the first time I prompted it with "Explain your capabilities to me", it responded that it was Gemini, a multi-modal model. I've seen others see the same with DSv4, as well as the "hallucinated" multi-modal nature. I would not be surprised if GLM similarly was partially (heavily?) distilled off of Gemini. A fun test would be to compare the logits for "gemini", "claude", etc for a continuation of "I am " on all these models. I'm sure that e.g. GLM, Qwen, DS are dominant, but I'd be curious to see the next highest contenders, and how they compare to each other.
- w4yai 1mo agoNever ask the model what model it is. Every single model may answer wrong answer if you keep asking.
- iamdelirium 2mo agoNo, it's the same internally and externally. Gemini 3.7 Flash is a pretty great model IMO. You shouldn't compare it to Opus, Sol, K3, etc since it's a much smaller model but it's a little better compared to Sonnet, Luna or Terra, etc.
- OldGreenYodaGPT 2mo agoTested this last week and couldn't get it to finish any task that took more then an hour with /goal keep getting errors
- Tepix 2mo agoGLM 5.3 Flash: 320B parameters with 18B activated Qwen 3.8 Next Flash: 125B + 51B = 176B parameters with 6B activated DeepSeek V4 Flash: 284B with 13B activated The new Qwen model is the most promising for one or two Strix Halo 128GB with the low number of active parameters. On paper it's much stronger than Qwen 3.8 27B.
- simonw 2mo agoGood bicycle, good pelican: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff5726c73f6f9a2dc61f70c1cfab74ae4 https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- joquarky 2mo agoI only see a mostly blank page with a "Paste" button, a "URL" button, and and a "Preview" label.
- simonw 2mo agoBy chance you have a browser extension that might block fetching data from raw.githubusercontent.com ?
- kelvinjps10 2mo agoI just see the raw svg code
- bigyabai 2mo agoIt's feeling good on non-pelican workloads too. Less verbose than 5.3, cheaper/higher usage limits, vision capability and some good web design one-shots even with vision disabled. With GLM 5.1 and 5.2, the big problem was tool calling and long-horizon coherency. 5.3 was more trustworthy at the cost of longer thinking traces, and now Flash seems to improve on it once again with a more concise, smaller model. As long as there aren't any noticeable regressions, I could see myself defaulting to this for >90% of my day-to-day coding work.
- konart 2mo agoError: Gist API returned 403 for me
- deleted 1mo ago[deleted]
- Aboutplants 2mo agoWhen do Chinese models surpass US models? I thought there was at least be a 2 year runway but now I think they surpass it within 12 months, if not sooner.
- hgoel 2mo agoI get the impression that they're focusing more on efficiency than raw intelligence. While I assume that all AI labs realize that the AGI "race" is mostly bs, the Chinese labs aren't stuck in a trap where they need to keep blowing money to maintain an intelligence lead to justify investments and valuations. So, while OAI/Ant have to fearmonger and keep training the largest models, Chinese labs can focus on efficiency more heavily and as long as they stay near frontier, they'll continue to get positive coverage.
- Mohamed_Mansour 2mo agoIt is totally fine I think
- mowmiatlas 2mo agoi wonder if more companies will now stealth launch their models. imagine they just released this on openrouter for free but under their normal name - would they get the records in token usage then?
- pohl 2mo agoDoes the word "flash" mean a specific thing when it comes to LLM models? I noticed that this word is used by gemini, qwen, and z.ai and I'm curious does it mean the same thing for each one, or did they all just accidentally brand similarly?
- Doohickey-d 2mo agoIt seems like it has come to mean "fast, small, cheap" models these days, and seems well enough understood as such that different AI labs are adopting it.
- browningstreet 2mo ago[dead]
- bel8 2mo agoIf you're on opencode's go $10/mo plan and want to use GLM-5.3-flash right now on pi, you can add this to models.json until pi updates to support it: { "providers": { "opencode-go": { "models": [ { "id": "glm-5.3-flash", "name": "GLM-5.3 Flash", "api": "openai-completions", "baseUrl": "https://opencode.ai/zen/go/v1", "reasoning": true, "input": ["text", "image"], "cost": { "input": 0.15, "output": 0.5, "cacheRead": 0.03, "cacheWrite": 0 }, "compat": { "supportsStore": false, "supportsDeveloperRole": false, "maxTokensField": "max_tokens" }, "contextWindow": 1000000, "maxTokens": 131072, "thinkingLevelMap": { "off": null, "minimal": null, "low": "low", "medium": null, "high": "high", "xhigh": null, "max": "max" } } ] } } }
- Kholin 2mo agoFor openrouter in pi: { "providers": { "openrouter": { "models": [ { "id": "z-ai/glm-5.3-flash", "name": "Z.ai: GLM 5.3 Flash", "reasoning": true, "thinkingLevelMap": { "off": null, "minimal": null, "low": "low", "medium": null, "high": "high", "xhigh": null, "max": "max" }, "input": ["text", "image"], "cost": { "input": 0.075, "output": 0.25, "cacheRead": 0.015, "cacheWrite": 0 }, "contextWindow": 1048576, "maxTokens": 131072 } ] } } }
- alexfortin 2mo agoFor using (lite) Z.ai subscription in Pi while it's not available yet: https://forge.l3x.in/alex/pi-shared/src/branch/master/extensions/zai_extra.ts https://forge.l3x.in/alex/pi-shared/src/branch/master/extens...
- halyconWays 2mo agoBetween Gemma 31/26/12/4/2, Deepseek-v4-flash-0731, Qwen 3.8 27B, Qwen 3.8 Flash Next (which I haven't even gotten to run yet!), and now GLM 5.3 Flash, I can't keep up. I love all these open weight models and am continually stunned that it's largely the West fighting for closed, restrictive, anti-user bullshit and China absolutely mogging the likes of OpenAI and Anthropic, with some notable exceptions like Gemma. Still, I shudder to think what the world would look like if we only had closed models. In many ways the stagnation of open source diffusion seems like that: LLMs are just a few months behind frontier, but image gen is like 1.5 years behind.
- coder-pm 2mo agoIs anyone actually tried it in agentic coding (claude code loops)? Are apple silicon macs (M5 Max) capable of working with that model? what was the tps?
- terhechte 2mo agoI've just used it for a fairly complex refactoring of the UI in a SwiftUI / AppKit app. It managed the refactoring in blazing colors, and the resulting UI looked really good. It was also quite fast. I'm impressed.
- coder-pm 2mo agoAnd what’s your hardware and what were the tokens per seconds metric (do you have it)?
- rush86999 2mo agoLuckily, I have the coding plan for z.ai, so I'm happy with this model as I always kept running out of usage with the original glm-5.3
- vladgur 2mo agoSo what is a good coding model to run on a 128gb m5 max MacBook nowadays
- nkjvhb 2mo agoI heard that Dario Amodei is not having a great day today. 2 really strong open models on the same day is a amazing.
- guybedo 2mo agoalthough i initially thought it didn't make sense financially to run this kind of model locally, i did run the numbers and for heavy users this could justify buying $10k worth of hardware with a ROI over a few months, less than a year. I was looking at my token usage, mostly from subsidized codex/grok subscriptions and i'm a somewhat heavy user. The thing is i would actually use even more tokens if it wasn't for the weekly quotas. In the end, with a $10k investment and running this kind of model, estimating a 2x increase in token usage because i wouldn't have weekly quotas and comparing to glm api prices, this thing could pay for itself in less than a year. Obviously i'm paying subscription price right now, so the math doesn't work. Although using local ai removes all weekly quotas. Keep a subscription to have access to frontier models for planning work, and local hardware + glm-5.3 flash for implementation, e2e testing, qa work 24/7. It's not that crazy of an idea and the numbers aren't that bad.
- minraws 2mo agoYou aren't going to get nearly as much token usage locally from DGX Sparks or even M5 Ultra (though it might be close, unsure would need to get my mittens on it to clarify). You will get around 2-4 concurrent streams of aggregate tokens at best for such a model and around 0.5B output tokens per month assuming you use loops and run it when you are sleeping. That's 500 (per mill) * 0.5$ = 250$ only at most. Then there is maintanence and efficiency costs due to electricity usage and such, any down time, etc. You will be lucky if you can squeeze more than 200$ of value out of it in a month. I don't think people should buy local hardware for money reasons, by the time you will pay off a 10K USD machine, 2-3K USD machine will catch up and beat it by a significant margin. Unless your expectation is that we will be in hardware winter for the next 10+ years. At 200$ per month it will take around 200 * 50 = 10k, that is, 50 months, so around 4-5 years. Again assuming you are making the most of your hardware somehow, very hard to do in practice. I don't recommend people to use compute as investment or payoff thing, but if you have the money to burn and can afford it why not, maybe with some software optimizations it will be cheaper but then again Z.ai is currently offering 50% discount and providers will offer cheaper rates for sure. But either way you will never be able to burn more than 200$ worth of token on a cheap hardware device, because inference becomes more profitable the more you scale it up, you have separate prefill and decode engines/systems, and a lot of nuance, but assume for every 10x increase in infra you increase margins by 5-10%. So from 10K to 100K to 1M to 10M to 100M.. I don't think this curve continues beyond 100M but I have no idea about that scale unless some AI lab is interested in hiring me lol. So a 100M infra will have ~30% better margins than you at 10K, then there is software optimizations but that's cheap enough, though some of it is only viable at scale. Either way assume 10K is the price of privacy if you really want to buy it. Don't worry about making the most out of the usage, you will always be in a net loss but I would assume for you 10K doesn't matter.
- mawadev 2mo agoHas anyone ever asked themselves why AI was made publically available in the first place? is it really economics or is it about training people to recognize the patterns of machine generated words and ideas?
- danieltk76 2mo agotbh I wasnt that impressed by it. initial benchmarks were trying to say it was AGI but i told it to re-build Palantir in 1 pass and it gave me a non working prototype
- pranav_tech26 2mo ago[flagged]
- melembre 2mo ago[flagged]
- BrucecarlL 2mo agoIt is bench maxed during the stealth testing. And it can’t beat DS flash on speed
- melembre 2mo ago[dead]
- freakynit 2mo agoGoogle was so ahead when it made this statement: "We Have No Moat And neither does OpenAI" - May 04 2023 https://newsletter.semianalysis.com/p/google-we-have-no-moat-and-neither https://newsletter.semianalysis.com/p/google-we-have-no-moat... Correction: Not a statement, rather, an internal memo by a Google employee. Thanks @granzymes for highlighting.
- deleted 2mo ago[deleted]
- nullbio 2mo agoDespite what any benchmarks tell you, I'm actually finding GLM-5.3 max to be better than Sol and Fable. Finally bit the bullet and installed OpenCode and OpenRouter and have been experimenting with other models. The labs are clearly benchmaxxing a bit to maintain perceptions. But I don't think they're in the lead anymore in terms of their public offering - although I'm sure what they have behind closed doors is far better than anything we're getting access to.
- petesergeant 2mo agoI use all three every day, and I am absolutely not finding 5.3 to be better. Competent and in the same league as, sure, but not better.
- broadsidepicnic 2mo agoOn what hardware do they run this? If I'll visit Shenzen, can I buy these chips? I'd much rather at this stage give money to any-other-manufacturer-than-Nvidia. I have a use case where I have to run local models, can't share data offsite.
- gunalx 2mo agoJust as expensive unobtanium as nvidia. Maybe even more so. Its probably the huawei ascend 910, they are using.
- ma2kx 2mo agoHoly shit, is this model really that bad??? Just asked it a question via the custom opencode go endpoint routed over cloudflare ai gateway doesnt show me the correct models. My fault was that I set "https://opencode.ai/zen/go https://opencode.ai/zen/go" as endpoint and tried my-gateway.com/custom-ocgo/v1/models, turns out I had to add /v1 to the opencode url and leave it on my -gateway.com. But first it told me that opencode is not on the compat endpoint and than it recommendet "Option 3: Nutzung seit Curse die Integration des Providers класный" . Not to mention some complete gibberish like "Falls Opencode.ai in Cloudflare Eing? Wenn ja, wähle den eingebauten provider o.ä. Weil der Gateway dann ein eigenes Modell-List gibt." or "Schlage die Modellnamen in einer Liste ab (z.B. lokal fester Key)" Is this just cloudflare or is it really that bad? I mean thats not even the level of LLama2 7b ...
- Mashimo 2mo ago> I mean thats not even the level of LLama2 7b ... I think you answered your own question. Must be something wrong.
- ma2kx 2mo agoI don't know yet. May be its a trade of for agentic skills in favor of language skills. GLM5.3 feels also heavily benchmaxxed compared to 5.2.
- desterothx 2mo agoI actually noticed this too with DS flash over openrouter. Instructions to respond in my native language cause the model to spew gibberish roughly resembling my native language mixed up with similar languages, and it was also dropping in Cyrillic characters randomly (although i noticed this with the anthropic models as well)
- duo1234 2mo ago[dead]
- heybear 2mo ago[dead]
- barrenko 2mo agoNot one mention of if it's any good at OCR in the thread?
- bkd9 2mo agoGLM-5.3-Flash is now the cheapest way to reach Artificial Analysis Intelligence Index 52.3 to 57.5, at 8.7¢ per task, and displaced Muse Spark 1.2, Grok 4.5, Gemini 3.7 Flash, and DeepSeek V4 Pro from the Pareto frontier. advance card: https://catalystneuro.com/llm-cost-frontier/images/advances/2026-08-26-glm-5-3-flash.png https://catalystneuro.com/llm-cost-frontier/images/advances/... tracker: https://catalystneuro.com/llm-cost-frontier/ https://catalystneuro.com/llm-cost-frontier/
- swingboy 1mo agoBeen trying it through OpenRouter. It starts replying in Chinese after a while, lol.
- KronisLV 1mo agoI guess this is also a really good indicator of just how much more tokens people would use when not limited financially. Wonder how much that tells the labs about pricing. Then again, it's 2.3x not 10x the next paid/cheap model.
- robeym 1mo agoBenchmarks and cost don't really help me understand how good a model actually is. Anyone have hands on experience working with the current newest models? What work did you do, and how did the model performance ?
- bel8 1mo agoI have been using it since it was available in opencode Go plan. It replaced all other models for me. It's much less chatty than DeepSeek V4. And it feels noticeably smarter. I also use Claude at work and GLM 5.3 feels like Opus 4.8 (which I consider better than Opus 5). It tends to be more proactive with suggestions after a task is done also. DS4 Flash is awesome but GLM 5.3 is better despite being a bit more expensive.