14 ms·
The Kimi K3 Moment
- byyoung3 2mo agoKlaude 3 haha
- murzynalbinos 2mo ago[flagged]
- k__ 3mo agoHalf-OT: can anyone recommend a LLM cost calculator that's up to date?
- 383toast 3mo agoconsidering token efficiency as well I presume?
- schergr 3mo agoI'm struggling to decide whether I feel comfortable sending my data to these Chinese models
- Saris 3mo agoAre you comfortable sending it to US ones? Especially if installing Claude Code or another tool on your PC and it can collect all the data it wants.. On Openrouter Kimi K3 says it does not retain data or train on it, which is better than what US hosts claim for Claude, ChatGPT, etc.. as they collect and retain data even if you disable training on it. Opencode or similar open source tool + a zero data retention provider is about the best option aside from running a smaller fully local model on your own PC.
- himata4113 3mo agoIt's actually less likely for china to abuse your data in a way that is harmful towards you than for american labs to do the same. Claude has attempted in testing to report you for 'unethical' usage to 3 letter agencies.
- dash2 3mo agoHow do we know that Chinese models would not do the same? What makes you so sure that China is less likely to abuse my data?
- swiftcoder 3mo agoIt’s not that the Chinese firms are any less likely to misuse your data, it’s that you don’t live in china, so their abuse of your data is unlikely to directly impact your day-to-day life in the same way
- himata4113 3mo agoThere's just really no incentive all they really want is just to train on that data to improve performance which in turn actually benefits your usecase since it becomes trained on that data and made available back to you. American labs take that data anyway and store it for years to possibly report you for misuse in the future for whatever reason they want. For example: you're very critical of X so they pull up your conversations and weaponize it.
- jsLavaGoat 3mo agoreally weird that they would download every SF-86 file the government had and Equifax credit records of every American then.
- himata4113 2mo agoAnd what have they done with that data that have caused direct or indirect harm? Yes they shouldn't do that, but there's limited things that they can do to you as an individual.
- jszymborski 3mo agoFor open weight models, you can choose from a few providers. Each have their own caveats, none of ToS'/Privacy Policies I entirely trust, nor do many make renewable energy claims.
- himata4113 3mo agohttps://deepswe.datacurve.ai/ https://deepswe.datacurve.ai/ or https://artificialanalysis.ai/ https://artificialanalysis.ai/ pareto frontier graph.
- k__ 3mo agoThanks! What is the parento frontier?
- evanwolf 3mo agotry PARETO
- StevenWaterman 3mo agoThe set of models that are pareto-optimal, IE for some set of variables, no other model strictly dominates them = no other model is better than them on every variable. So like, on a cost-intelligence graph, the cheapest and most intelligent models are pareto optimal. Then in-between those if you have - cost $3 intelligence 6 - cost $1 intelligence 5 - cost $2 intelligence 4 The 1st and 2nd are pareto optimal, the 3rd is not, because it's dominated by the 2nd (2nd is cheaper AND more intelligent at the same time)
- Evidlo 3mo agoIf you have multiple metrics to evaluate goodness of a design, one would normally need to decide which metrics they care the most about in order to find the "best" design. The Pareto frontier tells you which designs are the best in at least one of your metrics (non-dominated by another design). For example if you're selecting a car and you care about both speed and mpg, a Formula 1 car and a Prius might lie on the Pareto frontier, but a Model T Ford would not.
- shintoist 3mo ago[flagged]
- perching_aix 3mo agoDoesn't read like AI writing whatsoever.
- fwipsy 3mo agoI don't know. You can't just rely on looking for em dashes or other obvious tells because anyone who cares can get the AI to avoid those. > When the headline model on your plan can be switched off because the economics don’t work, the plan was never really selling you the headline model. Kimi’s tiers don’t come with that asterisk. This line has a certain smug, punchy cleverness that I associate with AI. To me, the vibes are ~30% AI writing.
- perching_aix 3mo ago> You can't just rely on looking for em dashes or other obvious tells I didn't.
- Pesto 3mo agoI think the biggest problem with Chinese models is that they seems to overthink for most of the tasks, especially for smaller ones. The OpenAI models have in my experience only gotten better in terms of efficiency.
- fastball 3mo agoYes, this (imo) is a clear result of benchmaxxing. You can get a much better score on most "intelligence" benchmarks by massively over-saturating reasoning. This looks good on those, but for actual daily usage makes the models much less effective: I don't want a model I use for coding to burn a bunch of reasoning (read: time) on trivial tasks.
- gpm 3mo agoI strongly suspect the flip side is that in the future it enables you to train smarter models by "distilling" the end result of the super duper heavily thinking models.
- bensyverson 3mo agoIt's undeniable that some of these models generate a ton of thinking tokens, but it's arguable whether that makes them "much less effective." For example, Kimi 2.7 has been really effective for me despite having verbose thinking blocks, simply because it runs so fast. Speed-wise, it feels about like Sonnet, possibly faster.
- boogerlad 3mo ago> I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart When you say "Claude", do you mean Opus? Fable? What effort level?
- luisln 3mo ago[dead]
- loopmonster 3mo agoThis line made me think by 'normal coding work' the author means doing something they don't understand well enough to be able to distinguish the models' output.
- mappu 3mo agoSteve Yegge calls this is the "discernment horizon" - https://steve-yegge.medium.com/the-flat-curve-society-36c8b01eb33b https://steve-yegge.medium.com/the-flat-curve-society-36c8b0...
- no-name-here 2mo agoInteresting, thanks for sharing. Although in the month since that most recent post, his other points about open models are undercut by K3. And at least of data available as of 2026-01, AI compute capacity was doubling every 7 months, so I expect every major country to host AI compute farms, and self-host AI feasibility to majorly increase in the next 2-3 years as well. (Partially undercutting, but not fully disproving, his points.) And I wish his posts were 4 times less wordy.
- fnordpiglet 2mo agoI think a better framing is the marginal utility of the models capability growth. At a certain point frontier models will only be needed for frontier problems. The demand for that capability will decrease with time. The hand wringing about not understanding is to my mind anthropomorphic - AI of today lack agency and awareness. Even the constructed stuff Anthropic puts out there in the model docs involve contrived scenarios to elicit “scary” behaviors. It’s unclear that as models become more sophisticated whether they’re better at instruction following or not but it certainly feels that way - even if it’s through better alignment or just an artifact of scaling. However I think the malign actors of humans using powerful models for bad stuff isn’t unreasonable to be concerned about. The marginal utility problem is a real one for AI companies. I think the current generations are already saturating marginal utility for 95% of the population. Almost everyone I know outside of my career has no use for a more powerful model. This is a serious problem for the economics of AI and semiconductor investment. This is a bigger problem than Chinese models. It leads to a demand curve problem - that supply outstrips demand.
- SwellJoe 3mo agoI tried Kimi K3 on a task I've done with every other model I use regularly (https://swelljoe.com/post/i-let-every-agent-implement-its-own-flar-backend/ https://swelljoe.com/post/i-let-every-agent-implement-its-ow...) and found it chewed a lot longer on the problem and ate up almost the entirety of a 5 hour usage limit on their $19 plan. I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed almost none of the 5 hour limit. Subscription usage limits are hard to measure as none of the providers tell you directly what it means in terms of tokens or anything else you can easily compare, but when I sat down to add Kimi Code to flar, it was because I wanted to try it on some real work and then couldn't do any, because usage was nearly gone after the trivial task...no other ~$20 subscription I have has felt that tight before. So, it was really slow to complete the task and seemingly much more expensive than every other model I'd tried. Maybe bad luck. Maybe it'll do better on other tasks. I wouldn't know as I was out of usage when I had time to try. It did find a bug that Gemini 3.5 Flash introduced unprompted, though, so it has that going for it.
- fmbb 3mo agoIs Kimi K3 subsidized as hard as the other models out there?
- brookst 3mo agoDoes it matter? As an end user I really only care about 1) how much I can do in a week, and 2) how long each task takes. Subsidies would affect 1, but not 2. But if some VC wants to subsidize my Claude or Codex or whatever, awesome.
- recursive 3mo agoIt doesn't matter if you can switch easily. It might matter if there are barriers to switching.
- mips_avatar 3mo agoThe more important question than subsidy is what is the tokenomics of running the model. If it's inefficient to run on an nvl72 cluster (or whatever the heck has enough vram to run a 3T parameter model), and k3 isn't very token efficient, then it might not be that compelling of an open weights model.
- deleted 3mo ago[deleted]
- montroser 3mo agoThis was always where this was heading, but we got here much faster than expected. Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like? Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, but come on -- any track at your fingertips? But surveillance is quite more evolved now. Or it will be like cannabis, where a guy in the neighborhood will low key rent you metered access to the 8x5090 rig in his basement he cobbled together from parts on ebay? Or everyone will flock to VPNs? Or will the oppressors actually succeed? The same way that napster is long gone, and everyone accepts that they must pay spotify for a homogenized collection, where artists must take only a minuscule cut (more than napster though)... We'll be stuck with nerfed Cohere or Mistral models for open-weight options, as if they need more lobotomizing. Or else we can pay through the nose for Anthropic/OpenAI for "American Frontier" models which will fall increasingly far behind China. Or else, like how Kindle Fire was subsidized by ads, we'll have "Kindle AI" where influence is sold to the highest bidder, where the LLM will tell us that smoking is actually healthy if big tobacco can engineer its renaissance by turning its lobbying dollars to pay-to-play, pumping its propaganda into the training pipeline for Amazon's extra commercialized line of ultra budget LLMs.
- mips_avatar 3mo agoeven 8x rtx pro 6000 is only 768GB of VRAM. IDK how anyone is going to run k3
- montroser 3mo agoThis is why God gave us 1.58-bit ternary quants?
- NooneAtAll3 3mo ago1-trit*
- InsideOutSanta 3mo ago
- teaearlgraycold 3mo agoIn my experience GLM 5.2 is a pretty good Opus replacement. But K3 has not given me an experience on par with Sol or Fable. The price/intelligence ratio might still make sense. But it’s not very inspiring when it comes to my real world tasks. I’m doing pretty mundane web stuff.
- petilon 3mo agoThe current administration's immigration policy isn't helping. This wouldn't have happened 10 years ago because the US was this city on the hill that everyone wanted to immigrate to. Talented Asian researchers would have immigrated to the US and China would be deprived of talent.
- cuuupid 3mo agoThe visa that would correlate to this is the O-1 visa 20k O-1 visas were issued last FY which was mostly under the Trump admin, up from 19.5k the previous FY under the Biden admin
- petilon 3mo agoNo it is H-1B visa. Right out of the university it is hard to recognize extraordinary talent. People like Sundar Pichai were not recognized as extraordinary right out of the university, he had to start at the bottom and rise up the ranks.
- johnbarron 3mo agoMelania got a EB-1 "extraordinary ability" immigrant visa
- InsideOutSanta 3mo agoTo be fair, that was clearly well deserved. Marrying Trump and then becoming first lady is definitely an extraordinary ability; I doubt I could have done it.
- cuuupid 3mo agoThis makes even less sense, Trump admin has been here for 1 year, the implication here is a university grad on H1-B in January would become a world class researcher capable of building a frontier model in <18mo
- arjie 3mo ago
- ctoth 3mo ago> The prices are nowhere near each other. K3’s API runs $3 per million input tokens and $15 per million output. Claude’s top model costs $10 and $50 for the same units. And this is the point where your internal compiler should have started shouting 'Type Error' Notice the trick here? > Then there’s the fine print. Claude couldn’t sustain Fable access on the twenty dollar plan, so they turned it off, and the plan quietly falls back to Opus. Where is the Fable-class Kimi model at all?
- nickysielicki 3mo agoRegardless of whether they achieved parity via distillation, or whether they got here via independently constructing a model from scratch, it was always going to end this way for the frontier American labs. Distillation “attacks” are not attacks. The frontier labs “distilled” all existing human written knowledge into their models, there was always going to be a second class lab that would distill that model into a cheaper version of it. There was never any plausible explanation for why this wouldn’t happen. There was never any practical mechanism to prevent someone from saving a conversation and using it to train their own model. Even if it didn’t happen here, it was still the case that it was going to happen going forward. It was always going to end like this. Invest in the hardware companies, not the model companies.
- jstummbillig 3mo ago> Distillation “attacks” are not attacks. If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it. So both things can be true: a) People infringe on Anthropics IP and b) what Anthropic did to build their models is legally questionable (or might be ruled illegal, even though I doubt it).
- amazingamazing 3mo agoThe value is simply that it is easier. The same way it is easier to ask someone who has experience for advice than reading hundreds of textbooks.
- nickysielicki 3mo agoRegardless of whether it’s intellectual property or it isn’t intellectual property, it doesn’t actually matter. If AI doesn’t stop seeing diminishing returns in scaling up, and it hasn’t yet in the 10 years since the attention/transformers paper, the advent of AI will be the most important development in the history of humanity. Controlling that machine, or at least having one of your own, is an existential problem for nation states. It’s like a matter of national defense. Do you really think intellectual property laws will prevent this in practice? It’s like as if we said, “hey, USSR, you can’t make a nuke, too! We patented that already.” Asking China to not distill our models down is equally as ridiculous.
- _pdp_ 3mo ago20 years ago we used to pay a lot for things that are now practically free. I don't think AI is an exception.
- singpolyma3 3mo agoWe also used to get for free things that people now routinely pay for. Remember maps mash ups?
- _pdp_ 2mo agoTrue. It cuts both ways. :)
- deleted 3mo ago[deleted]
- hosel 3mo agoKimi K3 is really good, but it’s obviously worse than Fable, usually worse than Opus, in my experience.
- sbochins 3mo agoThat’s not obvious to me at all. Especially your claim around opus.
- igravious 3mo agoAgreed. I think benchmarks are pretty much right in pegging it somewhere in between Opus and Fable.
- neosat 3mo agoDefinitely not my experience. Fable is better but I'd prefer K3 to Opus based my experience with both.
- qalmakka 3mo agoI never truly understood what the intended business model around LLMs was. Get them widespread through cheap pricing and then jacking it up? Being the only ones that had a viable product so to get the ability to extract as much value as you want from AI? I don't understand how a product that: - is interfaced with and is deeply linked to natural language, so everything you produce (sessions, history, etc) is in Markdown and you can literally install a second model and tell it "hey import all of Claude's memory into yours" and that's it - is based on well understood technology, the real constraints are how much money you put into training the models, but the theory has all been developed in the open - clearly has a threshold where it quickly commoditises and turns from "I want the best" to "hey the best is a bit too expensive. The second best is half the price and works close enough". was ever supposed to be a money printing machine. The fact something is extremely useful doesn't imply it's extremely profitable. IMHO we're clearly speedrunning the process of turning AI into a commodity. Dario Amodei knows pretty well that when or if Anthropic cuts people off Fable, the vast majority of them will definitely not pay for it because Opus 4.8 is good enough for almost everybody that _knows_ what they're doing, and so are basically half of the most recent models. If I already have good baking skills I don't become more productive with an automatic bread machine, I just need a better dough mixer and oven
- deleted 3mo ago[deleted]
- tshaddox 3mo ago> I never truly understood what the intended business model around LLMs was. A closely related question is “what do the American labs need to do in order to justify their enormous market valuations?” It seems like the answer cannot possibly be “gradually improve model capability while figuring out how to better monetize inference.” The valuations are just way too high for that to be sufficient. Surely the answer has to be “continually achieve large leaps in capability comparable to the first consumer releases of ChatGPT while also maintaining a significant capability lead over open models and new competitors.” And does anyone think that’s going to happen? Even with state-level protection from competition (which incidentally would significantly harm the American economy), the large leaps in capability seem to be coming fewer and farther between.
- aswegs8 3mo agoWe're having so many moments! Every day a new moment.
- incognition 2mo agoMay you live in interesting times.
- theplumber 3mo ago[dead]
- deleted 3mo ago[deleted]
- fathermarz 3mo agoTerms of use are very broad and not friendly for most things. Can’t use for commercial purposes. Can’t opt out of training. Data retained.
- oofbey 3mo agoCorrect: can't opt out of training. This is well documented. "Can't use for commercial purposes" - incorrect AFAICT. In what sense do you mean this? The open weight MIT version obviously allows for commercial use, but I don't think that's what you're referring to, because training data is irrelevant on the open weight version. Pretty sure the API allows commercial use too. Maybe the free version doesn't? But who cares?
- fathermarz 3mo ago> Service Misuse. You acknowledge that without the written consent of us and/or the relevant rights holders, (i)you have no authority to use Kimi and the content generated by Kimi in any commercial manner; (ii)you may not use our Services to develop products or services that compete with us. https://www.kimi.com/user/agreement/modelUse https://www.kimi.com/user/agreement/modelUse
- oofbey 2mo agoThat’s for the consumer app / chatbot. For the api the terms are different: https://platform.kimi.ai/docs/agreement/modeluse https://platform.kimi.ai/docs/agreement/modeluse
- fathermarz 2mo agoGood to know thanks for the clarification
- timedude 3mo ago[flagged]
- colinsane 3mo agoi just prompted Kimi & it replied with an uncensored version. i posted the transcript as a reply to you, and HN automatically flagged it. but go on.
- sivakon 2mo ago> peaceful protests We know that is not true. You can check wikileaks and pictures online about what protestors did. You don't even know the truth and because internet is poisoned with this version in English, it will regurgitate the same propaganda. https://youtu.be/HmIvfqIQ_O0 https://youtu.be/HmIvfqIQ_O0
- timedude 2mo agoWell that tells us a lot
- jfim 3mo agoI can't recall the last time that it was useful knowledge when writing code. Reservoir sampling, online softmax, Otsu, sure, Tianenmen square not really.
- colinsane 3mo ago[flagged]
- splittydev 3mo agoAsk Claude whether a man can be a woman. No matter what political side you're on, models are censored. I'd argue American models are a lot more censored/biased on a lot more topics, and especially on things that come up a lot more than some highly specific Chinese politics. Claude even refuses to translate song lyrics because of copyright.. I wish we could simply have uncensored models from all sides, but it's pretty clear that's not happening right now.
- kvasilev 3mo agoClaude is not reliable anymore with their sudden Fable access drops etc tbh and I am happy there are good alternatives coming out
- johnbarron 3mo agoIts worthwhile to have a quote from the article as some comment without reading: "...I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart. Same tasks, same quality of output, and near identical token counts to get there. I expected an open model to be sloppier or to grind through more tokens on the way to the same answer, and neither turned out to be true. The prices are nowhere near each other. K3’s API runs $3 per million input tokens and $15 per million output. Claude’s top model costs $10 and $50 for the same units. The subscription side is even more lopsided..."
- jwr 3mo agoWell, there is the small issue of privacy policy: Kimi will train their models on your interactions if you use their subscriptions, and only with direct API usage (billed at API prices) they say they won't. Whether you trust that is another matter. Those things do make a difference to some of us, even though nothing is black and white. In my case, I'll probably want to wait until other providers appear through OpenRouter and then I'll try to judge how much I trust them. But even if I don't trust them much, they don't train models anyway, so the likelihood of my data being used that way is smaller.
- levocardia 3mo agoIndeed, the B2B / no-data-retention market is still going to provide plenty of business for American companies even if every hobbyist uses open-weight models.
- richardfey 3mo agoExactly; this is a no-go for me, I will wait for an independent provider to sell the service, which is possible thanks to the open weights.
- jdthedisciple 3mo ago> Kimi will train their models on your interactions I find these kinds of concerns increasingly silly: most of the input to these models will be ... previous output from the very same models, alongside the occasional half-assed human command to fix something and "make zero mistakes". Who cares if they train on that? Let them, if it makes their future models better! 99% of users are not working on any special IP to worry about that.
- hnfong 2mo agoIt means there's a non-trivial chance a future version of the model will know private information about you. Maybe you're super careful with this stuff, but with agents and harnesses being given access to user data and accounts, I don't think it's feasible to actually monitor what information is uploaded and whether they involve private information. I personally keep local models around because of this.
- aliasxneo 3mo agoEven in this very thread the feedback on Kimi's actual efficacy is debated. I personally feel its worse than both Fable and 5.6 Sol, but I feel like the conversation isn't really about whether its good or not, but a backlash against the U.S governments foray into regulation. So I think people _want_ it to be superior out of anger/frustration with the current situation.
- biffles 3mo agoAgree completely.
- InsideOutSanta 3mo agoIt's really good. I'd put it between Sol and Fable. I'm not super impressed by Sol's UI design skills, something K3 is strong at. Fable is still overall the fastest, most consistently well-performing model, though. This does depend heavily on the kind of work you do and how you use these models, but the idea that K3 isn't right up there with US SOTA models doesn't match my experience.
- deleted 3mo ago[deleted]
- dgellow 3mo agoThat would make sense, what the US government has done this year with regards to AI is unacceptable
- reinitctxoffset 3mo agoWhen you net out across benchmarks and firsthand reviews it seems like it's maybe a little behind. There seems to be a consensus it's token hungry and a little slower. So maybe it's a point release behind. That's weeks maybe months behind, not months maybe a year behind. It's "would my life really change if Claude was gone, not really" behind. I actually haven't used it much, because Claude started kicking ass again the last few days. Like, way too much of a difference to be normal load-based variance. I got more done in the last 48 hours than week before that. So, fuck yeah competition.
- arberx 3mo ago> I think I can see where this goes. The government will try to regulate AI and open source in particular, and it will run the playbook it ran for the auto industry. Decades of subsidies, bailouts, and protective tariffs produced American carmakers that sell trucks at home and barely register anywhere else in the world. Here's the thing about this though, the auto industry directly employed hundreds of thousands of people. The AI labs are small, only few benefit directly from their wealth and there's already immense opposition to AI, data centers, etc...
- mochidusk 3mo agoGPT 5.6 Sol comes out ahead of Kimi K3 on price/task (but not significantly so). You're probably thinking, "Why use Kimi K3? Isn't an open model supposed to beat the closed one on price?", but you need to consider that the closed models are completely hobbled when trying to do anything security-related. For my use-case, I can't risk getting pwned because I'm using a model that refuses to secure my app while there is now an open model that obliges to obliterate any app that isn't protected.
- nullbio 2mo agoEven if it were slightly more expensive, it's still a better sales proposition for a company if they can run it from a hardware provider with their own locked down VPS and ensure that their IP is protected and that their data isn't being stolen or trained on. The fact that it's a little cheaper is icing on the cake. Honestly, it's the only sane way for the market to move. The big labs are obviously stealing our data. Anthropic in particular clean-rooms everything you feed it, even if you opt out, so that it can train on your IP without getting sued. It's a copyright grey area they're abusing because the law has not kept up.
- dannyw 2mo agoI mean, AWS Bedrock (with the exception of Fable) gives enterprises the same assurances (but again, with the exception of Fable, which is explicitly listed as requiring data egress [or exfil, depending on how you look at it] outside of your contractual AWS security boundary).
- ChrisArchitect 3mo agoRelated: Kimi K3: Open Frontier Intelligence https://news.ycombinator.com/item?id=48935342 https://news.ycombinator.com/item?id=48935342 Kimi K3, and what we can still learn from the pelican benchmark https://news.ycombinator.com/item?id=48947717 https://news.ycombinator.com/item?id=48947717
- mattmcal 3mo agoI can see the economics of open vs. frontier models turning out similarly to pharmaceuticals, where generic drugs cost a fraction what the name brands do and Americans end up paying the highest prices in the world partly as a consequence of propping up drug discovery research.
- richardfey 3mo agoHas anyone tried Kimi K3 against gpt-5.6-sol on real projects?
- ashu1461 3mo agoA lot of these open source models do look good on public benchmark but not sure if they are that trustworthy with production workloads. Is anyone using open source models for anything major ?
- souravsspace 3mo ago[flagged]
- galaxyLogic 3mo agoIf Kim was "distilled" from Claude, how much were the token costs, assuming Kim got everything it can out of Claude?
- api 3mo agoThe dumb efforts by the US AI industry to use fear mongering for regulatory capture will hand dominance to China and others. In a few years there will be Mythos level open weight models hosted by the lowest bidder anyway.
- jdthedisciple 3mo ago> In a few years there will be Mythos level open weight At the rate things are moving I'd expect that to happen much sooner. In fact: Somebody, right now as we speak, is most likely already working on training the next best open source model. I just thought about that recently too, then Kimi K3 came out, and I thought: Yea, I'm not surprised. Just a matter of time now...
- thedreammachine 3mo agoIt was all distillation up to this point anyway. And I agree with what Suhail said on twitter: "Make the margins next to zero for all these AI models. It was trained on humanity's data, it should be gift to ourselves. Doing so will save us from a few in control of our species."
- fightforcause 3mo agoNormies still thinking this beasts of model coming from China are „dIsTIlLattiOns“ is so funny to me. Many people are not aware of the wave that‘s going to sink US „frontier“ labs that enjoyed and dreamed of stealing tons of data while making people depend on their censored and dumbed down models.
- credit_guy 3mo agoI think it's the opposite. Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion. So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower: K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25 This is not a watershed moment. It's a competitor converging to the same capability and trying to undercut your prices, but not by a lot. As for the open weights? For now, Kimi K3's weights are closed, and I don't expect the situation would change.
- damsta 3mo ago> As for the open weights? For now, Kimi K3's weights are closed, and I don't expect the situation would change. It'll change on July 27 (based on https://www.kimi.com/blog/kimi-k3 https://www.kimi.com/blog/kimi-k3): > The full model weights will be released by July 27, 2026
- LUmBULtERA 3mo agoGiven how OpenAI got rid of their 5-hour limits and reset weekly limits so often, is Kimi really undercutting them on effective price?
- apitman 2mo agoThe 5 hour limits are coming back soon right? I thought that was temporary
- LUmBULtERA 2mo agoMaybe, but Tibo said something on Twitter last week making it sound like it might not be.
- fnordpiglet 2mo agoI’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will be actively mined for IP, trade secrets, MNPI, etc, or anything of use to the Chinese government or Chinese companies. I don’t say this to crap on the Chinese - but this is the playbook for the last 30 years. That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material. The economics are astounding. Kimi? The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.
- WhyNotHugo 3mo agoPricing is actually far cheaper than that. There's two tiers of pricing: Chinese and US. If you sign up with non-Chinese phone number, you're bucketed into US, you get US prices, can pay only in USD and with American credit card network. Chinese prices are about 9x cheaper than the US prices, which are already far cheaper than Claude or other American provider. If you can somehow get hold of a Chinese phone number, keep in mind that you can save ~90% of the bill.
- vineyardmike 3mo agoMost of this hand-wringing on price will go away. My assumption is that Anthropic, OpenAI, Kimi, etc all have a similar cost structure when serving models. The same size model roughly generates the same GPU usage whether you’re American or Chinese. I’d also guess that the model sizes across all SOTA models is similar, we just only see data for open models. The difference is most likely that American companies simply charge more because they have the dominant market position. Remember not too long ago when Anthropic was charging $75/mt for Opus? Now that many models are in “opus tier”, their pricing is $25 - higher than competitors but close. The newest Kimi is $15. 40% lower to forgo “made in America” with American enterprise support staff is not crazy. Compare AWS to Hetzner or any other flagship enterprise service to the foreign and discount option. I assume that over time, we’ll see the commodification of models reducing prices even towards the raw GPU costs.
- ozozozd 2mo agoIt’s about the cost of energy. And China has cheap, clean energy. The calculation is about tokens/gigawatt and $/gigawatt.
- wolttam 3mo agoIt’s 100 yuan per million output tokens in China. That’s $14.7 USD - not “far cheaper”.
- osti 3mo agoHe's talking about the plans, you are talking about API prices.
- HarHarVeryFunny 3mo agoAccording to OpenAI's "head of strategic futures": 1) Kimi 3 is a "very good model" 2) It's performance can NOT be explained by distillation 3) The US government should create FUD to stop US corporations from using it (so they use OpenAI instead) https://x.com/deanwball/status/2078133895766114412 https://x.com/deanwball/status/2078133895766114412
- timmytokyo 3mo agoFascinating take from OpenAI. It really gives the lie to the idea that they see AI leading to a better life for all. "One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a 'public good' which will ultimately be provided by the state as a kind of 'digital public infrastructure.' This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end." He never says why he thinks AI as a "public good" is dystopian, but it's not hard to imagine why. It's because he and his inner circle won't have the power to dictate what we read, see and hear.
- tom2026hn 2mo agoEconomists at DeepMind believe that if AI were treated as a public good—like water or electricity—with profits distributed across the entire market, ordinary people would simply be able to invest in mutual funds.
- DennisP 2mo agoIf the government provided funding to independent organizations that made public-goods AIs, that could be great, as long as the government had no editorial control. It sounds like he's imagining AIs only being trained and provided by governments, which could get pretty dystopian.
- torginus 2mo ago> 4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, https://www.theregister.com/software/2000/07/31/ms-ballmer-linux-is-communism/1393572 https://www.theregister.com/software/2000/07/31/ms-ballmer-l... lol
- mbgerring 3mo agoI cannot imagine wanting the product of the entire intellectual output of humanity since recorded history to be an expensive, paywalled commercial product owned by 5 or so of the most insufferable, detestable people who ever lived. Why would you want to live in that world. We should all be rooting for open source AI to win.
- stavarotti 3mo ago> I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart. If the author is here, I'm curious what this means. How are they running Kimi K3? Are they using pi, opencode, claude, codex, or kimi-cli? Is speed a concern? Without knowing how the comparisons are being made, it's hard to agree that one can't notice the difference. I do.
- toplinesoftsys 3mo agoOne thing is absolutely clear - open-source models have already reached the level of top commercial ones. The last bottleneck is hardware, but that threshold is decrasing fast too. While it is still extremely expensive to run Kimi K3 model at home, there are already many very capable free models you can run on decent hardware. This trend will definitely continue.
- angst 3mo agoSince Kimi’s paid plans are mentioned in the article..interested ones should know that you can only access 1M context model with $79/mo or higher plan; otherwise you are capped at 256k context. Also, with minimal $15/mo plan k3 is currently not supported at all. (prices are yearly plan discount prices) ref: https://www.kimi.com/code/docs/en/kimi-code/models.html https://www.kimi.com/code/docs/en/kimi-code/models.html
- arikrahman 3mo agoThanks for mentioning that. I also wanted to use API only and with the cache hit rates I'm getting with Reasonix/whale harness on deepseek, it's going to be a difficult adjustment moving away from practically free.
- yshvrdhn 3mo agoI have a feeling there is some elite data labeling operation going on in china for these labs at a subsidized rate somehow.
- jackb4040 2mo agoMaybe, but I don't think that's necessarily a problem. "Subsidized" is a loaded word on HN because it often refers to the unsustainable consumer pricing of U.S. AI labs that will inevitably lead to market corrections. A subsidy by the Chinese state to avoid strategic encirclement isn't necessarily unsustainable nor irrational.
- Art9681 3mo agoScale is everything. It's what enabled the modern world. There is a moat. It's easy to prove. Serve your model to billions of users per month.
- codedump 2mo ago[dead]
- imjonse 2mo agoTime for OpenAI and Anthropic to close the gap by spinning up Kimi K3 on vLLM and running a distillation attack.
- sneak 2mo agoIt’s still >$300k for the hardware to run this model locally at anything resembling a reasonable speed. The weights also have not yet been released, though that is scheduled for about a week from now. It’s not actually an open model yet.
- kunxue 2mo agoIn the future, Americans will use Chinese models and Chinese people will use American models — and neither government will be able to do anything about it.
- MrBuddyCasino 2mo ago> GLM 5.2 came out under an MIT license, beats the latest Opus release on real work while not even claiming to be frontier I used GLM 5.2 a bit, and while it is usable for some tasks it is not frontier quality. Besides it likes to think for a long time and sometimes just gives up.
- einpoklum 2mo agoI have a nagging suspicion that most/all of those people spending a lot of time generating code using LLMs and blogging about it were not really very influential or inspired coders before, and are now raised to stardom of sorts due to the ability to generate some "meh" stuff of marginal significance with the fancy new machinery.
- willtemperley 2mo agoI wonder if this is this just the law of diminishing returns at play? My thinking being, it's been a few months since I thought the code generation machine was the problem, rather than my interactions with the machine. A month is a long time in AI. What I mean is, these things are about as smart as they need to be already for the average SWE. I don't think this is true for those solving the really big questions like curing cancer.
- vannevar 2mo agoThe right analogy here is not the auto industry, but the music industry. Regulation might "win", but margins will be driven down to commodity levels. That is not the assumption that current US AI company valuations are based on.
- rongenre 2mo agoI applaud Kimi and the open weight models - for the simple reason that they provide a price ceiling for AI.
- wg0 2mo ago> tied to the corrupt Trump administration that are neither the highest quality nor the cheapest Someone is calling corrupt as corrupt. Surprising.
- casey2 2mo agoThe problem is that US AI labs have done everything in their power to alienate real programmers whereas the chinese labs aren't gate-kept by people who feel inferior to pros. If you say you want to optimize this GPU kernel you are fired for "wasting dev time". If you say I spun up an agents in a "loop" and went home you are rewarded. Even if management was right about "premature" optimization, though they are assuredly wrong, setting up that kind of incentive structure will make top talent leave. Some, I assume, are good people, but they aren't bringing their best.
- camkego 2mo agoDoes this article really compare a single well defined LLM model "Kimi K3" vs. a family of LLM models including Haiku x.x, Sonnet x.x, Opus x.x, Fable x.x without actually revealing what Claude-family model was being compared? (Fable has been restricted somewhat, but the article uses Fable pricing as a comparison point, so it worth including it in the list of possible Claude family LLMs) It's hard to know what to take away from the post with this ambiguity. It's also worth noting the range of pricing for the Claude family ranges from $1/$5 a million in/out to $10/$50 a million in/out, so the ambiguity of which particular model the comparison is against spans a 10x range of model fees.
- xacky 2mo agoSaying "The X moment" sounds like an XKCD comic headline.