3 ms·
Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models.
If $100 Claud Max subscription works for you, then great.
But
Except Fable won’t be costing $100 for enterprises that will be considering the Chinese models.
If $100 Claud Max subscription works for you, then great.
But you have to remember your pricing is subsidized by enterprises that pay hundreds of thousands of dollars each month, if not more, to Anthropic.
For those companies, a Chinese model that can cut their AI spend from $1M/month to $200k suddenly seems attractive.
And unfortunately for the American tech industry, the valuation is based off those enterprise deals, not your $100/month Claude Max subscription.
Yeah, fair. If we were talking $2,000/mo vs $200 then the maths starts looking very different.
> not your $100/month..
This is made brutally obvious by anthropics customer support for people with such accounts.
Didn’t deepseek recently announce prices will go up significantly?
Right now the US dominates everyone else in actual chips
in data centers. So even if deepseek etc tries to undercut, they’re very capacity limited.
Deepseek is open weight/source ie it will be running on US servers in the US maybe by on each companies own servers.
It's far from significant, it's partially doubled during peak hours. They could 16x it and it would still be two orders of magnitude better value than OAI's $200/mo plan.
It's that good. They are far from capacity limited, and even if they were, you can rent a single MI300X from somewhere like Hot Aisle and get more tk/s than you'll be able to use.
How does it compare to 5.6 Luna after the permanent 80% price cut?
That one is dirt cheap at API pricing, I can't imagine quota is going to be a concern on the $200 subscription, which in my opinion easily supports full time use of 5.6 Sol on xhigh.
It's still laughable. They blew it. I'm not sure they could even pay me to use their models at this point (and I don't mean via employment, meta and its refugees are permabanned to me). My experience over the past few months with open models has me seriously entertaining moving east.
A couple of very talented friends were uttering curses upon the entire bloodline of whoever convinced them to try letting sol xhigh do serious work. Deepseek cleaned it up for a fraction of $20.
Since you did not compare them, I now checked myself, and it looks like Luna benchmarks about the same as DeepSeek V4 Flash 0731.
The cost per task was $0.03 with DeepSeek, $0.05 with Luna. $1.23 for Sol.
Tokens per second 132, 202, 70 respectively.
> Didn’t deepseek recently announce prices will go up significantly?
They also previously said prices will go down significantly once they get a hold of the upcoming Huawei chips (later this year).
Prices are going up just because they can. It can easily come back down. They aren't strained by some IPO / VCs requiring them to 1000x their earnings.
I don’t think the industry knows how to price this stuff. Deepseek is great (I’m running it on a RTX 6000 pro setup) but it’s nothing like Fable. It’s still strongly human-in-the-loop which is fine, until you experience how good these models can be.
Think about it this way.
Let’s say you could buy an LLM that gets things right 98% of the time. But there’s another LLM that’s 100x the price but gets things right 99.9% of the time. To the lay person this sounds trivial but to a serious business this intelligence gap could represent millions, or billions of dollars.
If that’s the case businesses would be seeing millions to billions of profit gain (or cost reduction) in the past 4 months as they went from Opus 4.6 to Fable 5.
But that’s simply not the case. It’s very clear that vast majority of the business do not generate additional value from incremental intelligence gain from these models.
There is a reason why Chinese open weight models are now popular even in American enterprises, because CTOs realize that they are indeed good enough.
I work for a FAANG, and have my own personal projects for which I use the Chinese models. and the big models do indeed save/make us a lot of money.
The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
And when the big US models get better we will move with them. Until we stop seeing returns there is no "good enough", I don't know why this is so hard for HN to understand.
Unless you are very cash strapped, Fable is a very nominal fixed cost compared to the benefit of what it offers (fully autonomous agents, and no longer needing to pair program with one).
And even it isn't "enough". I can very clearly see myself using more advanced agents to move up the abstraction ladder.
For businesses that have actual problems to solve, I see them investing in the frontier for a good bit longer, probably until we have AGI that can replace employees, maybe even a bit after.
This is why I find the "good enough" arguments silly. Like, the usefulness of an AI tops out to you when you can pair program with it? Seriously? You cannot envision ways in which more advanced AI enables you to do more, better? That's bizarre to me. I don't ever see myself running out of problems to solve.
> The Chinese models are not good enough for anything other than pair programming, which is just a very last-gen way of using agents.
This matches my experience with DeepSeek V4 Pro at Max reasoning, the preview version of the model kept regularly messing things up. About 30-60% of additional time to fix the output was needed.
On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time, while it still definitely made noticeable mistakes, they were far fewer in total and less egregious.
Kimi K3 at Max reasoning drops that value to below 10%, it's about as good as Opus or approaches Fable in some tasks. At High reasoning it also seems to be pretty close to Opus 4.8, not sure about the latest Opus model yet, but it's up there.
Only problem is that K3 is nowhere near as cheap as DeepSeek models, despite me personally liking the writing tone more (less Anthropic slop) and finding that it doesn't block my cybersecurity prompts, recently reproduced SQLi with a proof of context so I could justify fixing it.
My overall thoughts (released over some time):
https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tragedy/ https://blog.kronis.dev/blog/ai-slop-is-a-self-inflicted-tra...
https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-done/ https://blog.kronis.dev/blog/kimi-k3-is-out-is-anthropic-don...
https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model-but-is-it-good-value/ https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates. Currently the main things keeping me with Anthropic are their performance (tokens/second) and the fact that their visualization abilities within the app are pretty good.
> This matches my experience with DeepSeek V4 Pro at Max reasoning,
That was ages ago (in LLM release timelines). DeepSeek V4 Flash beats it now and a lot cheaper.
> On similar tasks, GLM 5.2 at Max reasoning screwed up maybe 20-30% of the time,
GLM 5.3 bridges this gap.
> I'd say as Chinese models get better, whatever moat Anthropic and OpenAI have dissipates.
Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
> GLM 5.3 bridges this gap.
I’m sure the next models will only get better, when they’re released. Also super curious about what Moonshot will achieve and the full DeepSeek V4 Pro release!
> Their moat, especially OpenAI is funding and hardware resources. They gain train models 10x as large and also serve at large scale. That's it.
I’ve seen how much slower Kimi K3 can be and that part seems correct, their own GPU production still has ways to go and export restrictions definitely limit what they can do.
Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
> Not sure about the size part, if Kimi K3 achieves SOTA performance at 2.8T parameters, western models being >2x that size would be insanely bad in regards to efficiency. I bet they’re all within the same order of magnitude and below 10T and won’t really have a reason to go even that high for the foreseeable future.
They're a lot larger e.g. Fable. It is insanely bad. Do you know how much more resources "Western" companies have? Most in China don't have random GPUs to "play with" like every "frontier lab" employee does.
> As investors will start squeezing them for profitability, I suspect focusing more on efficiency will be commonplace.
They're born lucky though. Efficiency is "free". The next generation hardware e.g. Nvidia claims Blackwell -> Rubin is 10x efficiency (verified by Neoclouds apparently).
I use Chinese models, even smaller local ones, for much more than pair programming. If we are talking about deepseek v4 flash, which is basically a frontier model, it is much more capable than the local models I run on my MacBook Pro. The only issue really is finding the right harness.
I do have a way of correcting through redundancy, though. If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough.
> If you are just vibe coding, you need to use the most capable model you can find and even then it might not be good enough
I mean, this proves my point. Better models enable you to get more done. With Fable, 80% of the time, I no longer have chat with an agent over the details of a PR. I give it an outcome and it gets done. This means I can work on much more with the limited time I have.
And I don't see this ending. When better models come out that take that from 80% to 99.x%, I will have that better model manage teams of other models and move up the abstraction layer.
If models get even better than that, perhaps I stop reviewing PRs entirely. Maybe normies can start using agents to build real things.
Unless your business doesn't have many problems to solve and isn't in a competitive environment, it will benefit from using the best models.
If you don’t have a way to automatically check the results via redundancy, you need a really good model, since even 99.9% reliability is going to cause slot of headaches. If you do have a way to automatically check results, then you can use something that fits in your computer and was produced 3 years ago.
My point is you don’t need the best model if you just put in QA processes that can be done by models also. And if you don’t have that, the model is probably not going to be good enough.