3 ms·
This is going so fast! What a time to be on hackernews: July 16th: The "Kimi K3 moment" - China has caught up to Opus! 4 weeks later: GLM 5.3 - Same performan
by bertili 1mo ago
This is going so fast! What a time to be on hackernews:
July 16th: The "Kimi K3 moment" - China has caught up to Opus!
4 weeks later: GLM 5.3 - Same performance, but cut the amount of parameters and cost to a third!
12 days later: GLM 5.3 Flash - Almost GLM5.3 performance but cut the parameters in half, cut prices to a fifth and serving on Chinese chips!
- jatins 1mo agoexcept besides benchmarks, most of these models don't meet reliability of Sol/Opus in coding work. Opus unfortunately talks very weirdly so not a great out of the box experience
- computerex 1mo agoYou'll find that hard to prove objectively and conclusively.
- deleted 1mo ago[deleted]
- rxyz 1mo agoOpus 5 is the least reliable frontier-class model in the market
- usef- 1mo agoIn what way? It has worked well in my experience. It holds up with long context windows, unlike many, too.
- pimeys 1mo agoI have been mainly using Kimi K3 on programming work for over a month now. It is so far the only language model that does not piss me off all the time and can deliver my daily tasks without any trouble. It does not talk annoyingly to me, it just answers and does what I want. This is from somebody who put thousands of dollars every month to Opus. Now it's 40% of that and I get as good or better results without having to turn the caps lock on before lunch... Edit: yes company money. We don't get subscriptions we pay per token.
- shusaku 1mo agoHonestly I really like GLM 5.2 a lot for coding. There’s some weird failure modes in Anthropic’s models where it just does absolutely idiotic things.
- unknownfuture 1mo agoEh, I use Opus professionally and DS v4 Flash for personal work. I honestly don't notice the difference too often other than Flash being twice as quick and an order of magnitude cheaper. The reality is most work people do doesn't need the very cutting edge and these open weight chinese models more than cut it most of the time.
- Alifatisk 1mo agoAnd don't forget the coolest part, DeepSeek, Qwen, Z.ai and Moonshot have almost caught up while being open about their research and their model weights. We can mostly speculate about OAI and Anthropic models, nothing else, how fun huh?
- kzrdude 1mo agoExactly, DeepSeek, Qwen etc are catching the attention because they put out their tech docs and papers, so we can read about how the models work and what they think their innovation was this time.
- johntarter 1mo agoDo they publish their distillation strategies on the private frontier models? Just curious.
- rolymath 1mo agoI'm not sure if they publicly admitted to doing that. Would be interesting though.
- nylonstrung 1mo agoThe next 12 months will see OAI and Anthropic spiral into into increasingly hyperbolic PR stunts, manufactured benchmarks and underhanded attempts at regulatory captures I'm sure they have nothing to rival this on a price/performance basis and have already given up on that
- skippyboxedhero 1mo agoThere is a massive price war going on. All of these Chinese companies are publicly listed and exist outside the hype bubble required to ship Dario's dogshit paper onto the pauper's pension fund.
- refulgentis 1mo agoDario has less space than a Nomad!
- pkulak 1mo agoWhat are you guys doing where cost is such a concern? I have a $20 codex subscription and I was able to use it to build a bespoke scheduling website for an acquaintance over three days without even going halfway through my quota. On Sol xhigh. I love hearing about new models, but every time I just don’t know why I should use something worse. I tried some random model on fireworks a week ago, and it immediately went it a thought loop for 10 minutes before I caught it. Blew through most of my $10 for no output. What’s the point, exactly?
- gloomyday 1mo agoLess powerful models are already extremely capable, so going for the best model is just like buying the most expensive hammer in the shop instead of the functional and well-priced one. Your experience is not representative of their usefulness.
- Alifatisk 1mo agoI bet Luna (xhigh) could’ve completed the same task.
- bertili 1mo agoThe point is open AI. "Open" as in open weights, open research, open future.
- puelocesar 1mo agoThat $20 price is heavily subsidized to addict people like you. Their plan is that once people are addicted to it, they can just increase prices. Which now won't be possible because we have chinese open models to use instead. I'm very curious about how long these US companies can keep on burning money like that.
- FluffyPancake 1mo agoIt's not really fair to compare API pricing to subscriptions. You can indeed get lots of usage from a codex subscription, but once you start paying per token it gets a lot more painful as you noticed. There a price decrease is a big deal.
- bertili 1mo ago0 days later: Qwen 3.8 Flash Next: Let's cut GLM 5.3 Flash parmeters in half and active parameters to a third! Chinese models had 94% reduction in parameters (from 2.8T/104B to 180B/6B) in 6 weeks, while staying close to the same quality.
- gravypod 1mo agoWhy are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?
- xyz100 1mo agoPresumably there is distillation or similar being used to transfer from a larger model to a smaller one.
- scotty79 1mo agoThey innovated a lot.
- maipen 1mo agoHardware constrains forced this?
- scotty79 1mo agoPossible. But by looking at other industries Chinese don't seem to need to be forced to innovate. They just can and do. Unlike the West they seem to be on the way up and it seems like sky is the limit. In the West, the interests of the shareholders and other types of rent seekers seems to be the hard limit. Chinese have no qualms about making the cow obsolete before they milk it dry.
- Tuna-Fish 1mo ago> Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc No? A large model obviously can be dumb, I don't think you can infer much other than by testing it. These small models are almost certainly worse at some things than the big models. They prize is making them dumber at things no-one cares about while retaining the capabilities people do care about. A model probably does not need to be able to give me a political treatise on the late 19th century "silver question" to be able to write me code.