4 ms·
Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.
by 542458 2mo ago
Kimi K3 was an interesting model only a month ago, and now we're looking at the same performance for 1/20th of the price. Wild how fast this is advancing.
- whinvik 2mo agoYeah either the benchmark isn't very useful anymore or V4 Flash is a really, really good model.
- ignoramous 2mo agoIn my use, DeepSeek v4 Flash (which replaced the quite excellent MiniMax M3) lags behind GLM 5.2 & Muse Spark 1.2 (let alone Kimi K3). Also, K3 is a much bigger multi-modal model, while Flash is text-only and likely optimised for coding tasks.
- nwienert 2mo agoYep, and the v4 flash final is about 2.5x slower than preview making it no longer a fast model, in fact slower than Luna and bigger models in many cases. Spark is actually the interesting one imo. It's significantly better, also significantly faster. If you are ok with letting Meta soak up your data (which DS does too) it's also the same price.
- fallingbananna 2mo agoGPT 5.6 Luna is an extremely cheap and still very capable model. A chinese model being in the same ballpark of capability at half the price sounds believable to me.
- nwienert 2mo agoIt's significantly worse than Luna and quite a bit slower in some fairly involved tests I run.
- jfaat 2mo agoThat's fascinating, it's WAY better than luna ime. What sort of things are you testing it for?
- debazel 2mo agoI've been using this DeepSeek model the whole day today after building with 5.6 Luna extensively over the last week and I would disagree, at least for Rust + OpenGL. DeepSeek just spend almost 2 hours trying to figure out why terrain textures were not working. It tried everything over and over again, it even had reference code for meshes on how to setup the rendering with materials, and it could just not do it. I finally gave up and gave it to GPT-5.6 Luna instead, and figure out in a single prompt after 20 seconds, that the terrain mesh was being initialized with None in the material slot. Other tasks it has managed to figure out at least, but it is significantly slower than GPT-5.6 Luna and it requires a lot more iterations. (Both were set to high reasoning)
- nwienert 2mo agoThat's roughly my experience. Luna is extremely efficient and at higher levels of reasoning and longer running tasks more capable. Reading the DS reasoning is wild, it's constantly going in circles. The most minor lack of clarity in your prompt and it will spend ages going back and forth on what you meant. It reasons 5x longer than the preview which makes it really slow now as well. We did a lot of work to nudge it to be decisive and improve our evaluation setup to there's more clarity, and it helped but only marginally. Ours is a full-stack app one shot test so it includes backend, frontend, design, and QA/testing. It's graded by Opus xhigh and Sol xhigh and the grades are averaged. DS4 preview would finish in 20 minutes flat on high reasoning and grades 6/10. Luna high gets 9/10 in about 30 minutes. DS4-final is crazy - at high thinking it's taking over an hour and getting ~8 but only had one successful run as I got tired of waiting so long after many early abort/retries trying to debug why thinking was so long. The lowest thinking still takes over 45 minutes, and with thinking off it actually is finally closer to preview in time but actually get's a much more varying result anywhere from incomplete to 6 it seems. Costs per run DS4 is best but not actually by a lot as it's spending 10x the tokens with all the reasoning and mistakes. It's a very brute force model and I really preferred preview in many ways for how predictably fast it was. Side note, Spark 1.2 is a nice model for this test, best in frontend design and fastest to get results together, though not nearly as efficient as Luna. Grok scores similarly to Spark but at like $50/run vs the contributor Spark costing $1.50. Edit: was curious to see and seems DeepSWE agrees at least: https://www.together.ai/blog/deepseek-v4-flash-0731-vs-gpt-5-6-luna-on-deepswe-cost-and-coding https://www.together.ai/blog/deepseek-v4-flash-0731-vs-gpt-5... Edit 2: btw it tests a team of agents working together in a special harness and stack. So 20 minutes is for 8 agents essentially. That said everything was built around DS as it was the cheapest to iterate against so even with that advantage the new one struggles.
- thehamkercat 2mo agoAnd now nobody seems interested in it because the price hasn't gone down it's still $3/$15 for all providers on openrouter because of some Kimi license https://openrouter.ai/moonshotai/kimi-k3#providers https://openrouter.ai/moonshotai/kimi-k3#providers
- johnnyApplePRNG 2mo agoMorph has it for a slight discount, apparently. Uptime looks crap, though.
- thehamkercat 2mo agoI believe it's because they are below $20 Million revenue limit (which Kimi K3's license has) So we won't see any price decrease unless Kimi changes the license of K3
- w4yai 2mo agoSynthetic is offering $7/month subscription for this weekend (which includes K3), insane value for this price ! https://synthetic.new/?referral=kwjqga9QYoUgpZV https://synthetic.new/?referral=kwjqga9QYoUgpZV
- MarkLowenstein 2mo agoReal question: is there anybody that is both maintaining alpha-dev capability by keeping abreast of all these daily changes, while also reserving enough time to actually work? Seems like we've reached the event horizon of whether AI advances are worth paying attention to.
- cyanydeez 2mo agoAre you saying we've reached peak Bike shedding?
- bee_rider 2mo agoHow about: The yaks have started shaving themselves, who can keep track of how good a job they are doing?
- becquerel 2mo agoI think the play now is to just try out whatever the best new model is every time you see a headline that fundamentally reorganizes your conception of what's possible.
- petesergeant 2mo agoI don't think you need to be keeping abreast of them really, you just need to be using the best model you can get enough tokens from, which for many people is Fable 5 @ $200ish, ideally fanning out implementation to cheaper models
- bronson 2mo ago> which for many people is Fable 5 Not for me, Fable refuses to debug Linux kernel bugs. Unless you say who you're speaking for, it sounds like you're just shilling for Anthropic.
- petesergeant 2mo agoI would love to be shilling for Anthropic, but I am not. I am part of a group of about 30 developers, and 80% of them are using Fable 5 and very sold on it, with the remainder being committed to Sol. Both are competent, but among our set (who will try anything), Fable 5 is definitely winning. The fucking refusals for security work are insane though, and I hate them. I use Sol and Grok 4.5 as my inline debuggers/reviewers, and both do well, and are decent at token save. DeepSeek V4 Flash 0731 found some interesting bugs when I tried it a few days ago, and I'm curious to see if that also joins the code-review line up
- dyauspitr 2mo agoNot for long, Deepseek is saying they will have a significant price jump soon. They really shouldn’t do it because they are on the cusp of capturing the scalable API market.
- telotortium 2mo agoThey need to be able to serve their market. The price increase is partly load shedding. If they improve their ability to serve their load, they can always drop it again, as OpenAI did with Luna recently.
- ignoramous 2mo ago> as OpenAI did with Luna recently My read is, OpenAI is neither able to claw b2b money (away from Ant) nor are they able to stave off open weights on the other. In short, they're struggling to hold onto their distant #2 position in the coding market, and these pricing changes reflect a (desperate) change in strategy.
- lukewarm707 2mo agoand i still won't use it, because they log and spy on your prompts XD. the private endpoint costs 10x (azure). private endpoints for deepseek (lots of providers) also cost about 10x more. but 10x more for deepseek is $0.028 cached input, and 10x more for luna is $0.10.