3 ms·
This will be roughly on pair with Kimi K3, but using a third of its parameters. Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in
by bertili 2mo ago
This will be roughly on pair with Kimi K3, but using a third of its parameters.
Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.
Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.
- cmrdporcupine 2mo agoCongrats def in order but as usual the proof will be in the pudding of actually running the thing. GLM 5.2 has token efficiency problems. It's not a stupid model, but it takes a lot of "thinking" to produce not-stupid results. ("But wait..."). Which makes its pricing deceptive. I tried to get by through the month of June on just GLM 5.2 and it was ... fine-ish for about two weeks. But the provider situation wasn't ideal.
- dannyw 2mo agoHave you counted your thinking tokens for say Opus or Fable? It wouldn’t surprise me if frontier closed models “over-reason” just as much, but you don’t see it thanks to the summariser. (We do know GPT5.6 have adopted the caveman shorthand, which explains its token efficiency).
- cmrdporcupine 2mo agoI have a $200 monthly Codex plan. I never run out of budget and it's... disturbingly smart. It's very hard for anything to compete with that right now. I do occasional experiments where I cancel or downgrade that and try to live on open models only and it just never works out financially or skills wise. There's nobody offering K3 etc at rates that end up being significantly cheaper. Yet.
- rammler 2mo agoKimi is a great model but it was clear from the start they achieved they brute forced that performance through scaling. The frontier models K3 compares to are rumoured to be smaller also. GLM on the other hand is way ahead in perf/parm but severly compute bound. Now once GLM can scale up or Kimi optimizes the training more, that gonna be fun times.