4 ms·
The competition is real in pricing. Thanks for the Chinese open models, US big players have to cut their inference pricing. We've done a bunch of evals between
by pimeys 2mo ago
The competition is real in pricing. Thanks for the Chinese open models, US big players have to cut their inference pricing. We've done a bunch of evals between the models, and Kimi K3 was the first one that actually could compete or be even better than Opus or Sol in our use cases, with a fraction of the price. All our developers use K3 as their programming model, and it now powers a big part of our systems instead of Opus and GPT. Surprisingly the new Sol pricing is quite similar to K3...
Now DeepSeek v4 Flash 0731 is eating Gemini's lunch, and suddenly we saw a price cut (the "introductory price") for 3.7. DeepSeek is of same quality or sometimes better than Gemini for text, Google knows it and they have to compete. Too bad it's too little and too late, it's still 4-5x more expensive in our evals.
And these models are not going away, nor their prices going up because of competition in the inference providers and due to the fact that you can buy/rent the hardware and run them in your own premises.
- ywvcbk 2mo ago> Opus or Sol in our use cases, with a fraction of the price. I assume it's highly use case dependent, though? Even before the price cut seems like Sol was price competitive with Kimi https://artificialanalysis.ai/models?models=gpt-5-6-sol-xhigh%2Ckimi-k3%2Cgpt-5-6-sol https://artificialanalysis.ai/models?models=gpt-5-6-sol-xhig... And now it should be considerably cheaper
- pimeys 2mo agoLong-context agentic tasks and Rust engineering are our use cases where Kimi definitely is better than Sol. We can measure our own systems and the numbers say that Sol has no chance against K3 or Opus, and K3 is so so so much cheaper than Opus right now. You cannot just look at the price tags for these models, you must eval and see the price per task. In our previous eval rounds Sol was more expensive than Opus (with its original price), took much longer, and provided worse results. Kimi does not have these issues, it's just as good as Opus with a smaller price tag.
- mrngld 2mo agoChina's 50 Cent Party being a real and noticeable thing (and the two biggest things they like to shill is open weight Chinese models and the futility of resisting a Taiwan invasion), I have to take things like this with a healthy dose of skepticism without corroborating data, since independent evals didn't show the price per task lead you're showing. If there's independent data showing this feel free to share a link, I haven't seen it. DeepSWE has been most closely matching what I see in my own use.
- pimeys 2mo agoInternal reports from company? Maybe not. I'm just saying you have to eval eval eval if you are working in this industry. There's a ton of victories in price, and price is right now the key thing all the customers are talking about. It's not always Chinese models. For example GLM 5.2 just did not work for us at all. And Gemini is still the best cheap model for non-text agents. If you don't have a good eval set and if you don't check the models weekly, you are missing on things. And Opus 4.8 is still the absolute quality king for agentic tasks. Too bad it's so expensive. And the clearest thing here is that Fable, Opus, and Sol are all too expensive. I'd say a healthy 75% cut to token prices and they are back in competition.
- c16 2mo ago> I'd say a healthy 75% cut to token prices and they are back in competition. Surely you don't want them to be the reason the bubble bursts?
- pimeys 2mo agoYep. It's a bit scary also. There's a lot of opportunity in the market now, but the downfall of the big US inference labs is going to hurt here in EU too sadly...
- tkgally 2mo agoI was struck by a video ad that Google released yesterday with testimonials by three developers about using Gemini 3.7 Flash [1]. The point they emphasize most is price, followed by latency. The marketing strategy definitely seems to be shifting. [1] https://youtu.be/kacf2bib-X0 https://youtu.be/kacf2bib-X0
- theptip 2mo agoYou play to your outs. Gemini is far behind on quality.
- pimeys 2mo agoYes and no. It competes in the mid tear not in SOTA. It's a very valid model if you need things like computer use or image recognition. Especially with the 3.7 "introductory prices". It's multi-modal and better than GPT 5.6 Terra while only a bit more expensive.
- eru 2mo ago> And these models are not going away, nor their prices going up [...] Well, DeepSeek just raised prices.
- pimeys 2mo agoAnd Fireworks did not yet. They are still under the limit of not feasible to self host... Let's see if other providers follow DeepSeek with their flash pricing.
- deleted 2mo ago[deleted]
- pimeys 2mo agoThey did now. Landing somewhere between Terra and Luna now per task, with the quality of Gemini 3.7 flash.
- gagan2020 2mo agoI shifted from DeepSeek v4 Flash 0731 to Gemini 3.7 flash on openrouter and price shoots up almost double with no visible change in outcome. So, today I reverted back.
- Bayart 2mo agoTry DS through their own API if that's feasible, AFAIK they're much cheaper than through OS due to cache hit rates.
- pimeys 2mo agoThey just raised their prices sadly.
- bogdan 2mo agoIs there actually a difference in cache rates between OR and official API? I have a preset set up on OR so that I only send traffic to deepseek. The preset is important otherwise you will send traffic to different providers but if you weren't doing this already then what can I say, water is wet, of course cache rates will be awful. I get about 70% cache hit rate with the preset which is appropriate for what I'm doing. I haven't used the official API though.
- irthomasthomas 2mo agoZenmux say the cache hit rate is 98% for the deepseek flash API. I don't know why, but performance is definitely worse using openrouter. https://zenmux.ai/deepseek/deepseek-v4-flash https://zenmux.ai/deepseek/deepseek-v4-flash
- Bayart 2mo agoCould be compression or metadata in their pipeline adding noise and making requests less generic.
- guntribam 2mo agoI'm using hermes and keep close control over open router provider to DSv4 flash 0731 Use only the deepseek provider, you can configure openrouter to do that(but they create some obstacles, go to configurations and allow all providers) Right now i'm testing deepseek harness, don't wait to test it. The plugin architecture and self awareness of workflows really make you think about "what is a tool vs what is a project". It has been an amazing experience
- miki123211 2mo agoIt's funny how quickly we went from "the greedy US companies are subsidizing prices to keep competitors out of the market" to "the greedy US companies are overcharging because they are greedy."
- fluoridation 2mo agoThey are subsidizing the non-API use cases and overcharging on the API use cases.
- senordevnyc 2mo agoNo, almost everyone has been in agreement that they’re subsidizing subscriptions, but there have literally been dozens (hundreds?) of threads on HN in the past 12 months with people vehemently arguing that API prices are subsidized.
- jmuguy 2mo agoI think both can be true - they're losing money on the API and they're still charging too much. Which is really where the music stops for American AI investment. I also use K3 now and its perfectly capable for the development work I'm doing. I don't shed any tears for OpenAI or Anthropic.
- platinumrad 2mo agoAPI are not subsidized. We know this is true from the existence of independent inference providers (a lot of which are crypto companies that would otherwise just be mining if inference weren't actually profitable).
- jurgenburgen 2mo agoThose providers are VC-funded, they don’t have the leeway to pivot away from AI. I would be interested in some examples though.
- ee334y5rthsrth 2mo agoTell me why price adv not working this same on web pages. Why price od advertisment on portals, social media etc. not fall down?
- hacompanian 2mo agoNot including Grok/Cursor seems a major flaw in your model
- pimeys 2mo agoOh we did try to get Grok for evals but they had some weird EU limitations last time we checked. Which the open weight models don't have.
- gloryjulio 2mo agoWe know the price wars are coming. It's gonna be interesting to watch how it plays out in real time. Anthropic/Openai probably want to race to the IPO before the price wars start to have real impact.