9 ms·
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
Also: Kimi K3: second only to Fable 5 on AA-Briefcase
https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark https://artificialanalysis.ai/articles/kimi-k3-agentic-knowl...
- Sphax 3mo agowhat are my options if i want to use a router like this ? who provides one ?
- platinumrad 3mo agoOracle routing is by definition not possible.
- deleted 3mo ago[deleted]
- OutOfHere 3mo ago[dead]
- mrinterweb 3mo agoTwo that I use are: * Openrouter.ai for a hosted router * https://github.com/diegosouzapw/OmniRoute https://github.com/diegosouzapw/OmniRoute for a local router
- culi 3mo agoA third the cost, open source, and won't refuse every other request because of some vague possible connection to cybersecurity concerns.
- OutOfHere 3mo agoOr biology or chemistry.
- SXX 3mo agoOr weapons manufacturing. It can refuse to answer on a lot of topics.
- estearum 3mo agohmm almost like the exact race-to-the-bottom + arms-race dynamic all the doomers have been warning about
- andrewmutz 3mo agoOne man's offensive penetration tool is another mans defensive tool. In the recent HuggingFace/OpenAI incident the safety controls stood in the way of the defenders, not the attackers: https://www.thestack.technology/hugging-face-hacked-turned-to-chinese-llm-for-help-after-us-models-blocked-blue-team/ https://www.thestack.technology/hugging-face-hacked-turned-t...
- estearum 3mo agoYou're presenting further evidence of lack of effective control over these systems as... a mitigating factor...?
- californical 3mo agoWhy do you assume they're disagreeing with you?
- estearum 3mo ago> One man's offensive penetration tool is another mans defensive tool. seems to suggest the author believes there's some intrinsic equilibrium Which, 1) is definitely not proven and not guaranteed (open to proofs otherwise, not pithy sayings that have zero normative effect on reality) 2) is apparently "supported by" further evidence of lack of effective control, which does not feel like equilibrium whatsoever
- ekidd 3mo agoApparently, the attacker in the Hugging Face case was reported to be an internal OpenAI model trying to break into HF and steal the answers to cybersecurity benchmarks: https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur... It really doesn't matter what restrictions are placed on public use of models if the attacking models are internal models at the AI labs themselves. So if the AI labs are literally running rogue models breaking into other organizations' servers, then yes, I am OK with those organizations self-hosting Chinese models for defensive use.
- dgellow 3mo agoMy questions about strawberries got blocked as too dangerous! I’m not joking
- shikon7 3mo agoStrawberries are a well known weakness of LLMs, as they have a hard time to count the numbers of "r"s in them. Maybe that's why, because they fear that weakness could be exploited somehow.
- skeledrew 3mo agoProbably need to be taught by someone of Latino origin. Learning to roll them "r"s could help.
- verdverm 3mo agoI always thought this should be easier by telling it to write each letter one a new line, then count, any token separator ought to suffice, so much so you'd think they'd have trained in this strategy given tokens make individual letters opaque
- HDBaseT 3mo agoStrawberries aren't the weakness, the weakness is the tokenization of a prompt. Any word with multiple duplicate characters is going to be troublesome for LLMs.
- fnordpiglet 3mo agoInterestingly the reality of the open source release is they are opening up to full distillation by the closed source model providers at a deeper and more fundamental level. If anything the open sourcing will help Anthropic and open ai ladder up faster. Open source has always been about mutual cooperation towards a goal and has never closed the door to commercial success. All the hand wringing about open weight models putting closed providers at a disadvantage doesn’t get what working in the open actually does for commercial interests - it is like science in the open - it enables and lifts all boats. Likewise commercial success doesn’t close the opportunity for competition or more open source work - it’s the economy of activity and competition that matters overall. When things stagnate is when closer concerns turtle up and collude on not competing for each others turf. The future is good and better for everyone the more work is in the open and the more work is in the commercial space. It’s good all around.
- linkregister 3mo agoI'm happy that Kimi K3 is indeed SotA and its open weights are due to be released soon. It's also true that Moonshot and other labs distill from Claude. This has been reported on extensively. I don't think there's any alpha for Anthropic distilling from this model. I do not mean to discount the tremendous amount of innovation regarding MoE and quantization that Moonshot has accomplished. But its training with synthetic data is in large part from distillation from frontier labs.
- fnordpiglet 3mo agoWhen I say distill I also mean mine it architecturally for insights but I doubt seriously the model training is entirely distillation of Claude, it’s almost certainly a mixture of both original corpus and reinforcement as well as distillation. I think it’s a little condescending to imply that these new open models are cheap ripoffs with nothing original to them. These teams and labs are top tier as well, working under unreasonable constraints imposed by the USG. That’s a powerful combination for creativity.
- linkregister 3mo ago
- mikae1 3mo agoFor regular chat users it's $19 while Claude is $20...
- AlexErrant 3mo ago$20 users don't get access to Fable. It's $100+ tier only.
- mikae1 3mo agoTrue. However, I believe most non-programmers don't need access to the fanciest model, but just want to use a good LLM without constant nagging about usage limits. Then 19 vs. 20 is true?
- potwinkle 3mo agoStill no. If you're only getting Opus-class you can still end up paying less by just using an equivalent Chinese model on OpenRouter.
- mikae1 3mo agoNot talking about me or HN users in general. These companies likely need regular peeps to begin using their services in order to become profitable. Not sure these are the ones who will buy from OpenRouter.
- culi 3mo agoNo "19 vs 20" is not true for "Kimi vs Fable". It might be true for "Kimi vs Opus" or whatever. Anyways looking at pricing plans is not a good way of comparing the price of LLMs. It makes more sense to look at cost per token: Kimi K3: $3/$15 (input/output) Fable: $10/$50
- skeledrew 3mo agoI'm a $20 user and was given $100 extra usage credit today, "for Fable" (I'll be sticking to Sonnet and sometimes Opus TYVM).
- villish 3mo agoFrom the blog post moonshot refers to it as open source but only mentions releasing the weights.
- charcircuit 3mo agoThe weights are the "source" of a model.
- tadfisher 3mo agoIf weights are the source for models then ELF binaries are the source for software.
- RobMurray 3mo agoclearly not true. the weights are the preferred form for making modifications. Do you really think people should be downloading hundreds of TB of training data and running make to build the model on their own cluster of GPUs?
- deleted 3mo ago[deleted]
- villish 3mo agoRandom people? No. Governments and big corporations? Yes. It removes any concern of "backdoors", and is currently the best starting point for your own model which will be as capable as k3.
- cookiengineer 3mo agoEverything is open source if you know how to reverse engineer ;)
- tadfisher 3mo agoThen everyone is open source if you have $20 to spend on tokens, I guess.
- deleted 3mo ago[deleted]
- wand3r 3mo agoThere is too much business risk in running on an American company. They can pull the model back or lobotimize it. No Alex Karp fan, but he was right: companies are worried about hyperscalers stealing their alpha. The lack of guardrails, data security and control really give these open models the edge. If you factor in that they appear to cost less, the hyperscalers are in big trouble. I don't see how this works out. Software was always supposed to be deflationary and collapse down to 0 marginal cost but this isn't remotely the case. It's truly amazing to see the state of open weight models
- guessmyname 3mo agoWhy SoTA (uppercase “T”) instead of SotA (lowercase “T”) ? “State of [T]he Art” versus “State of [t]he Art”. If not SotA then at least SOTA, which is more accurate.
- sim04ful 3mo ago[flagged]
- _carbyau_ 3mo agoNot wrong though. Technically correct and when it comes to communication, also seems like a reasonable query. In a world of so many acronyms, details matter.
- travisgriggs 3mo agoI’m overwhelmed by LLM/Agent signal on HN lately. I actually liked this question. It’s trite and whimsical, but frankly nice break.
- walrus01 3mo agoGiven the lack of tone and inflection in written content, it can be hard to tell the difference between doing a tongue in cheek imitation of pedantry and actual earnest pedantry, so let's try to assume the more charitable interpretation until we learn otherwise. I'm sure that more than once we've all had something we wrote which was intended to be obviously taken in a satirical or intentionally nonsensical manner taken seriously by someone else on the internet.
- lexandstuff 3mo agoIt should be SotA.
- jamesinmn 3mo agoDoes that make DeepSeek V4 Flash MiniSotA? This dev in the Twin Cities would like to know.
- stingraycharles 3mo agoAs always, benchmarks rarely paint the whole picture. It also seems like this article is somewhat biased, eg when Fable and Kimi are close but Fable wins it’s “dead heat”, but when Kimi wins it’s “Kimi wins”. GPT 5.6 seems to be missing as well. I am really eager to give Kimi K3 a try, but I’ll reserve my judgement until I’ve worked with it for at least a few days.
- rogerrogerr 3mo agoThe apparent bias may be explainable as it’s not remarkable for OpenAI or Anthropic to be slightly ahead. It _is_ remarkable for an open weights model to be better than the closed models from the trillion dollar (allegedly) companies.
- the_sleaze_ 3mo agoI believe the Chinese government is angling to destroy the western economy and rise from the ashes. Instead of a billion a day to bomb some buildings and bridges they're intentionally hamstringing the biggest concentration of speculation in history
- malshe 3mo agoI agree. The commenter you replied to makes it sound like the Chinese models are coming out of tiny startups with meager resources. It's really not a David vs Goliath story.
- rogerrogerr 3mo agoIt’s not about who’s developing the models, it’s the fact that free alternatives that are neck-and-neck are available at all. What’s the story for OpenAI & Anthropic’s valuation if they have to compete against free-weight models? Starts to feel like a commodity.
- omnimus 3mo agoIs it? From what i could find Anthropic and OpenAI are hovering around 5000 employees where as Deepseek and Moonshot are more like 300. The funding/investments are similarly many times less. I agree that 300 employees is not a small company but the overall outlook for Anthropic/OpenAI is not geat.
- OutOfHere 3mo agoThey forgot to compare and incorporate GPT-5.6-Sol.
- jrflo 3mo agoHmmm, a company that hosts open models is telling us how good open models are...
- mrinterweb 3mo agoThe don't only host open weight models. Also, why not promote this. If Fireworks thinks this big news might convert some new business doesn't make it not true.
- 30minAdayHN 3mo agoand in fact, it would be detrimental to business, if they wrongly promote Kimi, as they would pretty soon lose trust with their users
- jrflo 3mo agoI'd be doing the same thing if I were them as a marketing move
- yogthos 3mo agoAnybody who's tried it knows that what they're saying is true though.
- JSR_FDED 3mo agoThey share their methodology and results. I learned things about the relative strengths and weaknesses of Kimi and Fable I hadn’t seen anywhere else. Should being in the model hosting business disqualify them from sharing?
- Avicebron 3mo agoPeople are allowed to say, "hey we make money off of this thing, it's cheaper and almost as good as the thing we can't money off". Other people are allowed to call them on that.
- jrflo 3mo agoDoesn't disqualify them, but it may call into question their results seeing as they have a potential conflict of interest.
- HawtAds 3mo agoFor model routers, do they have to retrain the routing model every time a new LLM is released?
- pishpash 3mo agoMaybe a continuous evaluation process, why is there anything to train?
- cbhl 3mo agoI really like the idea behind OpenRouter's new auto-beta -- classify by task type, and then just follow what the market is using based on the last 7d. https://openrouter.ai/docs/guides/routing/routers/auto-router#how-auto-beta-works https://openrouter.ai/docs/guides/routing/routers/auto-route...
- mattvr 3mo agoAnyone have routing harnesses like this describes with Claude Code? Or other good routing platform recommendations? (yes, I know this article is about an oracle router)
- pandinus 3mo agoThere's this https://github.com/code-yeongyu/oh-my-openagent https://github.com/code-yeongyu/oh-my-openagent which implements the OP article's oracle pattern across 11 roles. Each role has a whole ranking of recommended LLMs across many providers. For example "Sisyphus (claude-opus-4-8 / kimi-k3 / glm-5 ) is your main orchestrator."
- monksy 3mo agoAlso, you can't use your claude subscription with oh-my-openagent. But you can with Kimi. ALso K3 is on OpenCode GO right now (low limits, but it's possible)
- felipeerias 3mo agoMythos/Fable was the state of the art back in March, if not earlier.
- danny_codes 3mo agoIt released in June..
- broodbucket 3mo agoTo the public. Mythos has been in active use for quite a while.
- culi 3mo agoIt makes no good sense to evaluate models we don't have access to. For all we know K3 was competitive back then too. Or maybe there's a K4 in the works that blows everything out of the water. Who knows and who cares. There's no way for us to compare
- broodbucket 3mo agoThere isn't, I think the point is in terms of how "far ahead" models are, K3 was likely trained much more recently than Mythos was. It's just something to note, it's not especially prescriptive. Mythos was withheld because of the threat to security and/or marketing stunt (depending on your leaning), I don't see what benefit there could be for not releasing K3 immediately.
- adamisnotroman 3mo agoI wonder if Fable now is actually better than the Mythos in March and if it's actually the same model. Could just be more Anthropic shenanigans.
- solumunus 3mo agoIsn’t Fable just a restricted version of Mythos?
- refulgentis 3mo ago[flagged]
- JSR_FDED 3mo ago> hyping an open model that isn’t open Moonshot has committed to release it end of the month.
- deleted 3mo ago[deleted]
- villish 3mo agoI'm not sure why you're so dismissive about this model. It isn't currently open, but it will be. It was slow when I used it so i'll give you that. There will be other providers with better performance. My only gripe was that it takes a very long time to think. It's a good model otherwise.
- nharada 3mo agoIs there something specifically with Kimi that's better here? As far as I know Kimi pricing is about the same as Sonnet 5 -- what happens if you use that model and Fable instead? Or Grok 4.5 which is even cheaper?
- ralusek 3mo agoOne benefit of an open source one is that you can, as a large corporation, run it "locally" within your own data center. Even fine tune it.
- culi 3mo agoThere are very few companies that would ever be able to afford to run it themselves. But it does give you the security that it's technically possible. Puts some limits on stuff like Moonshot changing their terms/conditions/policies
- adamisnotroman 3mo agoTechnically, "open weight" but yeah
- lenerdenator 3mo agoThe problem with hosting these right now is that it's not a known quantity like lots of traditional business software is. You can more-or-less approximate what your build out spend for a data center hosting database software is going to be. Your book of business will require a given amount of revenue to pay it off, but once that's known, it's off to the races. With AI, things are moving so fast and new business models are being tried all the time. You would be competing with some of the wealthiest companies in the world for data center hardware capable of hosting these models in a usable state. We've gone with Claude somehow hosted through GCP Vertex AI where I'm at.
- claytonjy 3mo agoHow big is this market, self-hosting a model that requires 64 GPUs, H100 or better, with good interconnects between nodes? I suspect the overlap of those that can afford it, and those that have the talent to manage it, is a fairly thin slice of the Venn diagram. Even the large corps are gonna be getting it from the inference vendors, or more likely Bedrock and friends.
- arjie 3mo agoInteresting. So the latest in the technology now is this model routing thing. Cursor estimated Composer + Fable works much better than Fable alone. And here K3 + Fable is supposedly better. Interesting.
- anentropic 3mo agoThose are different techniques... Cursor's Composer + Fable combo was a plan agent + execute sub-agents swarm Fireworks K3 + Fable router was dynamically choosing single model for the task based on cost+performance metrics
- lvl155 3mo agoIt is not SOTA. Give me a break. Sure, run it on Cerebras to get speed but that’s pretty much its advantage.
- nozzlegear 3mo agoWhy do you think it's not SOTA?
- sergiotapia 3mo agoCerebras does not share the quantization of the models so you don't know if you're getting real K3 or k3 lite or something else.
- codexon 3mo agoIt most likely will be quantized. A cerebras wafer only has 44gb ram, and linking them together vastly reduces the speedup.
- hodgehog11 3mo agoAgreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic tasks, like all the other open models right now.
- apatheticonion 3mo agoI love the Chinese models. I use DeepSeek exclusively and now Kimi K3 offers a great planning assistant for more advanced coding tasks. DeepSeek v4 Flash is extremely fast and is able to handle pretty much anything I've thrown at it (I use mostly Rust, PSQL, Angular and Terraform). I self host Bifrost as my LLM gateway, though I wish LLM vendors would do monthly/daily automatic billing (like VPS providers do) rather than prepaid + auto-top up. It's annoying maintaining a non-refundable minimum balance across vendors, I would rather be billed for my exact usage. OpenRouter helps, but I don't really like it as a service and not a fan of the mark up.
- copperx 3mo agoWhat don't you like about the service, besides the mark up?
- apatheticonion 3mo agoFor me, the only utility OpenRouter gives me is billing consolidation - I don't really need the routing capabilities because I use Bifrost for that. As a router, it's not very feature rich. For example I restricted the available models to the ones I want to use however the `/models` endpoint still lists all the models, making my LLM client list the 200+ models available on the service (even though they will throw an error if I try to use them). With Bifrost, I can also create model aliases with custom configuration - for example I can create a model alias `deepseek-v4-flash-nothink` which disables thinking. I can create `deepseek-v4-flash-caveman` which injects the caveman skill (to save tokens) etc. Plus I can contribute to Bifrost, which I can't do with OpenRouter.
- robbiet480 3mo agoDid you look at LiteLLM at all? It seems fine but Bifrost looks interesting too.
- apatheticonion 3mo agoYeah I looked at it, LiteLLM is functionally more mature. The only reason I didn't go for it is I'm not a fan of Python dependency management, Bifrost is just a single executable that uses nearly no memory and is lightning fast.
- JSR_FDED 3mo agoVery interesting. They test Kimi K3 and Fable on a set of approx 1000 tasks grouped into 5 areas (SWE, Legal, etc). They put a router model in front that predicts whether Kimi or Fable is going to give a better cost for a correct result. (They believe that ultimately such a router model should be continuously trained on your own workloads so it makes the best decisions for you). Their router chose Kimi the majority of the time (72% in one category, all the way to 96% in another category), leading to cost savings in every category (from 1.5x to 50x depending).
- eli 3mo agoThere are a bunch of routers like this eg https://openrouter.ai/openrouter/auto https://openrouter.ai/openrouter/auto
- crazylogger 3mo ago> Oracle routing is a method for measuring the best theoretical performance by running the task through each model and then picking the cheapest correct option (the cost/performance ceiling). Their "router" is an oracle reference point where they choose the lower cost model after running both and therefore knowing who passed the test. The cost savings part is only Fireworks theorizing what would happen if an equivalent predicting router exists. That's a big if.
- robthompson2018 3mo agoThe way they published this is baffling to me... surely you can try to implement some router and then see how well it does. Using an Oracles makes the whole writeup so much less interesting.
- Lalabadie 3mo agoBut then that becomes an article about the performance of your draft router. This one is about the fact that there's this level of optimization potential.
- 3mo ago
- sbinnee 3mo agoOpenrouter also features routing. Routing is indeed an option if you don’t need consistent behavior and allow switching models.
- JSR_FDED 3mo agoIt would only be inconsistent if the router chose different models for the same task. Considering that of 89 terminal tasks there were 11 that only Kimi got right, and 7 only Fable - having a router gives significantly more consistent results if you define consistent to mean correct.
- luciana1u 3mo ago[flagged]
- Buttons840 3mo agoI have several thoughts about this, which I'll just iterate: 1) US export bans have made it so that Chinese companies have to compete using less-than-state-of-the-art hardware. This has forced Chinese companies to build more cost efficient models. Whereas, US companies have moreso tried to be state of the art by spending more money than anyone else on state-of-the-art hardware. 2) Xi Jinping has called for more open AI models (not to be confused with the closed models of OpenAI), and I'm happy to see a powerful world leader advocating for open-weight AI models. Whereas, the US seems likely to just ban models. 3) My impression is that, if China surpasses the US in AI development, there will basically be nothing that the US does better than the rest of the world--except for military spending--we spend a lot, but we don't necessarily spend well (something something Iran). I mean, if the US is no longer a tech leader in the world, like... what are we a leader at? Manufacturing? Healthcare? LOL. Are we a leader in any industry or by any metric? I wonder if China is attempting to remove the last jewel in the USA's crown with these AI releases. 4) It must be refreshing for companies to have access to a new model that isn't going to get pulled because the government bans it 2 days after release. And it's open-weight so it wont go away--amazing--what a shift in the market. 5) If it becomes clear that open-weight models are the future of AI, will that pop a huge bubble in the US economy? Maybe. But, on the other hand, these companies aren't just training AIs, they are also building data centers which will remain valuable no matter what happens.
- shdh 3mo agoImplying that AI is the last jewel of the USA is too simplistic and ignorant. People are paying for tokens, that will likely continue even if open weight wins.
- Thorondor 3mo ago[flagged]
- wnoise 3mo agoDon't forget pizza delivery.
- xyzelement 3mo ago
- audioh4cker 3mo ago[flagged]
- johnhess 3mo agoWas this an out of sample test of the router or was it trained on these specific use cases/eval suites?
- swiftcoder 3mo agoNeither. They ran all the cases using both models, picked the winning result for each one, and then said "if you had a router that guessed with 100% accuracy, here's what it would have picked"
- skybrian 3mo agoThe article is about the best you could theoretically do with a perfect router. The takeaway is that trying to build a good router is worth doing. But it's unlikely to be a perfect router.
- hmokiguess 3mo agoWhat's the data governance and privacy controls on using Kimi K3 if I subscribe to their coding plans? I want to migrate away from Anthropic
- cromka 3mo agoNeed to wait until "western" providers start hosting it.
- shostack 3mo agoIs fireworks not private enough?
- HDBaseT 3mo agoBased in the US, they glow.
- deminature 3mo agoFireworks (the author of OP's article) is a western provider based in San Mateo, California
- scilro 3mo agoFireworks isn't serving Kimi K3 yet. Presumably, they ran this benchmark against the Moonshot API. All of the Western providers with sufficient capacity will be able to make it available when the weights are released Monday.
- Maxious 3mo agoFireworks has been working on porting the Kimi Delta Attention (KDA) hybrid linear attention mechanism to their hosting infrastructure https://x.com/FireworksAI_HQ/status/2079776331609584005 https://x.com/FireworksAI_HQ/status/2079776331609584005
- nicce 3mo ago
- deleted 3mo ago[deleted]
- matheusmoreira 3mo agoThat's incredible. Hope the chinese keep it up!
- adamisnotroman 3mo agoIf only to put some pressure on american labs to bring those costs down
- replatformradar 3mo agoIf only you could run K3 locally that would be the magic bullet to make it a true magic bullet!
- HDBaseT 3mo agoUntil Kimi 4 comes out and you'd be like "if only you could run Kimi 4 locally" Everyone wants the latest and greatest.
- byzantinegene 3mo agountil they put their money where their mouth is
- sho 3mo agoYou can! It just might be a little bit outside your budget.
- deleted 3mo ago[deleted]
- exabrial 3mo agoThe irony is the Chinese are being very democratic with their models, while the USA tries to do central control. Glad to see centralized control fail on the grandest scale. Maybe we can learn a thing or two.
- petilon 3mo ago[flagged]
- jeanlucas 3mo agoSo if they paid for pre-training they would not be able to open their weights?
- petilon 3mo agoNot if they want to recoup their investment.
- j-bos 3mo agoThey distilled Fable, in the couple weeks it was available?
- gr_norm 3mo agoThe idea that the Chinese labs cannot make progress except by copying superior American products is just prejudice against the former and exceptionalism of the latter at play. Even the OpenAI top brass have admitted otherwise [1]. China is an equal match in every respect, and we'd better admit this to ourselves sooner rather than later so as to see the game clearly. [1] https://xcancel.com/deanwball/status/2078133895766114412 https://xcancel.com/deanwball/status/2078133895766114412
- petilon 3mo ago[flagged]
- 3mo ago
- bjourne 3mo agoI'm skeptical. According to arena.ai, Fable 5 dominates almost every category: https://arena.ai/leaderboard https://arena.ai/leaderboard Kimi K3 has an edge in WebDev but struggles to reach top 10 in many other categories.
- onlyrealcuzzo 3mo agoIn my experience, Fable is not even close to Sol 5.6 High (not even the max tier) for coding. 1) it's substantially slower. 2) it's substantially more expensive. 3) it's code is considerably worse. It's a joke when you consider what you get for what you pay for.
- TurdF3rguson 3mo agoYour experience is an anecdote. Leaderboard rankings are a distributed blind taste test.
- singingtoday 3mo agoThis is very interesting to me because I find sol to be inferior at code generation, but superior at conversation and code review.
- bjourne 3mo agoSpeed and expenses are non-factors for arena.ai.
- zkmon 3mo agoAnthropic looks like Roman empire fast-farwarded, getting to the other side of the peak even before the IPO.
- yid 3mo agoYeah, I didn't get in either.
- captainregex 3mo agoI enjoy all the fun of these new models as much as the next guy but I truly don’t see a circumstance in the near future where my $200 a month with the frontier labs doesn’t get me more than enough consumption of what I need. Local models, chinese models, etc are all very fun weekend projects to tinker with but until something changes (entirely possible!) with how much you get with one of the subscriptions I just don’t see why I would move. What am I missing? Is it simply that a subscription is no good for production use cases? I kinda feel the same way with choice of coding harness, openrouter, etc. why would I use anything other than frontier if I don’t have to pay any more pretty much no matter how much I use? pls tell me if I am holding this wrong haha
- barumrho 3mo agoIf everyone took your position, then frontier labs can keep raising their price.
- abdullahkhalids 3mo agoI will accept a 5% drop in benchmarks for a model that talks to me like a human.
- semilin 3mo agoWhy? LLMs are not humans.
- tonyhart7 3mo agoArtificial Intelligence end goal is to assist(replace) human
- TheRoque 3mo agoThen why try to act like one ?
- abdullahkhalids 3mo agoDoors aren't humans either, yet we design their handles and locks to be graspable and manipulable by humans. The purpose of technology is to serve humans. Therefore, technology must conform as much as possible to human sensibilities rather than vice versa.
- 3mo ago
- ulfw 3mo agoNothing makes me happier than seeing AI becoming a commodity rather than the winner takes all bullshit Anthropic and OpenAI have been chasing with hundreds of billions of investor money
- brunooliv 3mo agoGenuine question: can these posts be paid to hype the open source models? If yes, what would be the purpose? On my work tasks, FastAPI Python and Springboot Java on a modern SaaS product, the only open model that can do tasks well and efficiently is Qwen3.7-Max. In all my experiments, both GLM-5.2 and Kimi are busy grepping around the codebase for ALMOST 70-80K tokens before writing anything and when they do it typically breaks the code… it feels to me that these models are good but only when you write out a super detailed spec of the task just like it was done a year ago… Qwen3.7 just… does it
- roncesvalles 3mo ago[flagged]
- ipsum2 3mo agoPretty sure she's American.
- quadrifoliate 3mo agoI am tempted to quote your inflammatory post, but don't want to give it more airtime. Please stop doing this. If there are specific points of the analysis you think are suspect, point them out. Others are doing a decent job (e.g. citing that open weight models have higher margin, so Fireworks is incentivized to promote them).
- ipsum2 3mo agoIt's (good) content marketing. They sell access to Kimi K3. They're one of the biggest model inference providers out there.
- cuuupid 3mo agoFireworks is an inference provider that specializes in running open source models very fast and makes almost all of their margin on chinese models So the economic incentive is literally their entire business model lol
- chvid 3mo ago
- nxtfari 3mo agoIf you haven’t been really running and testing these models yourself, they are all benchmaxxed. No matter how close they score to frontier on whatever metric, they always fall apart in real world tasks and their token efficiency is ridiculously bad. Fireworks has incredible incentive to make this claim in a headline, because Fireworks hosting K3 for you is pure profit for them, unlike when they host closed source models.
- Der_Einzige 3mo agoYou got baited by bad sampling settings. It's exactly the opposite. Go turn on min_p once it's available post July 27th and most of the problems you describe will go away.
- carterschonwald 3mo agoi've had trouble finding any anecdotes or data about how to actually set/explore logit sampler settings
- nl 3mo ago> Go turn on min_p once it's available post July 27th and most of the problems you describe will go away. This seems both arrogantly dismissive ("you are holding it wrong") and incorrect. Either the OP is using Kimi K3 on Moonshot where is is presumable set correctly (K3 isn't available elsewhere yet), or they are using Kimi K2.x and there has been plenty of time to experiment with this.
- Der_Einzige 3mo agoI've earned my right to be arrogantly dismissive since almost the entire field (including the Kimi and qwen team) doesn't know good sampling settings. This is because if you are truly "bitter lesson pilled" you don't think sampling is needed at all. You can either not believe me and be wrong, or you can (after July 27th) turn on min_p or a better sampler (i.e. top-n-sigma if you got it running via llamacpp) and have it work even better. Up to you.
- 3mo ago
- mickgardner 3mo agoSoTA means "State of the art". I wish it didn't take me 5 minutes to figure out what SoTA stands for.
- dagurp 3mo agoThank you!
- charmpic 3mo agoThe Kimi K3 felt pretty good when I tried it out.
- rayzia 3mo agoKimi K3 showing competitive performance with Fable while both sitting at the SoTA level on fireworks.ai is a huge milestone. Really interesting to see how the landscape is shifting here.
- greenleafone7 3mo agoYour account <...> request reached organization TPD rate limit
- mathew_woo 3mo ago[flagged]
- maroziza 3mo agoShifted to K3 and it is like a fresh air. While Fable and Sol very good at _generating_ code i even cannot force sol to just read all relevant source files. As result it reinvent existing things or assumes too much about internals of other, which lead to incorrect uses. Even with hard planing mode GPT burned 33% of week tokens for 3 hours producing no result and even cannot find root cause, but K3 fixed it in a minutes. Hilarious that while software not started with empty cache Sol handcrafted empty cache to let it start. Having full procedure right in the MEMORY.md. K3 found this, and downloaded cache correctly. But while using gen5 models i always feeling myself ignored. Any commands, steering anything - just ignoring. At first i added lots of hooks, no "?", expect in rust code, bash hook ban for find|grep|tail, with notice to use ltsp but then it started to ignore strategically. Also whole "thinking" thing is hidden from claude. regenerated thinking summary is incomplete and not useful. While being very verbose kimi k3 doing good job providing whole train of thoughts. So looks like claude degradation started from 4.6 comes to a logical end. Maybe i missing something and changing existing code beyond bug fixes is not a way to go. But there is no currently stable way to generate code on hier of specs and lean models i'd like to.
- hammasansari641 3mo ago[flagged]
- stpedgwdgfhgdd 3mo agoThese routers can be interesting on a company level to optimize for cost and quality, but for individuals who mostly work on the same tasks, i doubt it. You want to leverage the cache and switching models within a task seems not cost effective to me.
- acd 3mo agoThis benchmark is probably also self promotion of services. Fireworks happens to make a router. Using the router gets better performance. https://docs.fireworks.ai/deployments/routers https://docs.fireworks.ai/deployments/routers
- RALaBarge 3mo agoI work for a router company too, I ran some tests on all of the cheapest models and came to the same outcome where a handful of small models ran together in conjunction outperform SoTA models -- outperforms in that it got a 95% vs a 94% and I bet that changes with the day of the week. Anyways, I did get a similar result in a different sort of measurement.
- ethanpil 3mo agoCan you elaborate on running them "in conjunction"... are you running the same query on multiple models and then using a third model to judge or make consensus? or am I misunderstanding completely. I'd like to understand how these small models "run together"
- greggh 3mo agoI do this in OMP, a fork of Pi. It lets you set different models for different tasks. So with an API that has many different companies models I can set the Plan model to the best one, right now I am using GLM 5.2 for that, it plans really well. I have Vision set to Kimi 2.7 Code (cheaper and vision is just fine). Minimax M3 is set to the Advisor role (double checks work). Deepseek v4 Flash is set for the Task role. And MiMo 2.5 pro is set as default. With this setup GLM handles planning and managing my AGENTS.md, and orchestrating subagents for tasks from the plan/todo GLM created. The tasks themselves are handed off to Deepseek v4 flash to implement with strong instructions and examples for each agent. Minimax M3 reviews the output as the Advisor and recommends changes, catches bugs, and whatnot, subagents can be re-run with that information. Overall I am saving a lot using some of these smaller models. But with this setup I am getting great results.
- Vivek-KY 3mo agocluade model, claude opus 4.6,7 perform very well.and stable, understandable ,i miss that in kimi, its slow, doesnt gives feels like claude
- TokenHarbor 3mo ago[dead]
- kaycey2022 3mo agoWhat is routing? How do they decide which is better? The only way to come up with a routing model for your workload is to send queries to all the models and then come up with a way to say which solution was better, often trying out multiple times for the same model + query to account for other statistical errors. This makes you, an ai inference user an unwitting AI company with a non scalable product. The biggest mental trap people have fallen for is the notion of “best” and, always using the frontier model. Instead you should just bite the bulet and choose the cheapest or the best. This routing dance is just tokenmaxxing in disguise. Edit: Another thing concerning me is that the models themselves are not concrete behind their endpoints. Models can be arbitrarily dumbed down by reducing their inference resources. Once Anthro/OAI release a new great model, they are fully incentivised to dumb down their current models to drive traffic to the new shiny more expensive one. In fact they can do this for any reason. Once they switch things behind the API interface, how useful is your meticulously tuned router? Not much at all.
- sunnybeetroot 3mo agoThey can but there isn’t evidence that they do.
- kaycey2022 2mo agoIf they can then that is enough
- qiuwu 3mo agoAnti-China: K3 is propaganda and benchmaxxed, no matter how anthropic and openai reactor for these, it just a smoke signal. Pro-China: K3 is good choice for better and affordable choice to smash down the Big three ruling.
- vrganj 3mo agoChina-ambivalent: Open Models are good, Closed Models are bad. Not centralizing power in a few big American companies is good. China is who's building this right now, so we're aligned for now.
- qiuwu 3mo ago[flagged]
- jiaosdjf 3mo agoI used to be anti-China, and I still think the Chinese government is just a highly adversarial entity that will subsidise, steal and cheat its way to the top. However... American corporate culture has driven me to this. Fuck Blackrock and all these disgusting parasitical corps - they literally sold China the rope to hang us with and I'm sure as hell not going to pay a cent more for it than I have to. If China can offer close to state of the art for a fraction of the price then I'm going to use it - thats what the globalists wanted isn't it? They didn't care about saving local manufacturing so why should I care about saving their stupid investments.
- nicman23 3mo agoi only care about the license and results to be honest
- jkwang 3mo ago[dead]
- kunxue 3mo agofor people in mainland china the only option now is Kimi K3+GLM 5.2 for their day to day work as Fable is blocked. for me i use Fable + codex when i was out of mainland china, it is good when you have options wherever you are in this world
- terekhindc 3mo agothe per-token comparison keeps missing that k3 spends way more tokens per task. if it burns 3x tokens to reach the same result as fable, cheap per-token stops mattering
- raesene9 3mo agoThey address that in the article :) to quote :- "So where's this huge price gap coming from? token pricing, prompt caching, and effort-per-task. On SWE for example, K3 works much harder than Fable: roughly 55 turns and 1.3M tokens a task versus 21 turns and 130K. On the long terminal tasks it's the other way around: Fable is the one that spirals, running up 64 turns and 1.5M tokens (sometimes straight into a timeout). Prompt caching does most of the work of turning that effort into K3's price advantage: even when K3 reads ten times the tokens, with cache hits that means that SWE runs still come in lower cost than Fable. There’s a tradeoff. Tasks with extra turns generally mean more wall-clock time per run i.e. slower runs. If you need an answer in two seconds, that matters; if you're running agents in the background at scale, a bill that's a fraction of the size matters a lot more."
- kian 3mo agoSo this makes sense for a standard SaaS app - but given that models in general perform much better with low context window usage, it probably also means that Fable is still significantly better at 'frontier-level tasks' -- hard research problems, complex geometric rendering algorithm optimization, etc., no?
- thecopy 3mo agoWhen DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
- christophilus 3mo agoBecause, as you said, this is no longer news. Deepseek was the news. This is just the predictable progress playing out.
- h2aichat 3mo agoIt might be because no interested party is using that piece of news to move market. They have enough with other news to do it. Just a practical matter (or may be something else, who knows!)
- pembrook 3mo agoBecause markets in the short term are almost a random walk and making the blanket statement that “Walmart stock is down today because Deepseek” was an easy narrative to repeat for media people who cover the stock market. In reality the world is a highly complex, chaotic, reflexive system, and saying “the entire market moved today because of 12,000,000 factors that randomly aligned” isn’t satisfying enough for people to follow your media channel so they can monetize your eyeballs.
- uhhhhwhaaaa 3mo agoBack then people didn't understand how AI was run. It should have probably made Nvidia stock actually go up. I think the other Factor, and I might just be two into AI and most people are normies, the hype around Chinese models we've learned is overblown. United States models are a league above.
- pama 3mo agoIt was never DeepSeek’s release that dropped the NASDAQ at the time; it was the unknown risks of the early thoughts related to trade wars. Popular financial newspapers can promote anything they want, but these news do not typically drive large investor decisions.
- codemk8 3mo agoIf I had an oracle, why would I need an LLM?
- jack-brown 3mo ago[dead]
- mmaunder 3mo agoSeems that “oracle routing” is a term the authors invented. Also sounds like they’re sending requests to all routes but assuming the readers request will route to the “best” model. The closest thing to what they’re describing is semantic routing using a NN search in a vector DB to make a routing decision, but the efficacy of this approach isn’t a slam dunk.
- tracker1 3mo agoKind of cool to see.. that said, there's more to a tooling experience than benchmarks and specific models. Cursor, Claude Code, Codex, etc. add to the mix. Things like Open-Router and backend options make it easy enough to test. The tools, libraries and languages you are using can also dramatically affect results. Even on state of the art models, I find, for example, the output of SQL for complex interactions, or C# for that matter to be sub-par, where I find Rust results to be pretty great, with JS/TS falling in between. At the best, it can feel amazing and productive, at worst, time consuming and annoying that you could have done it faster yourself. YMMV in real world use. Note: I'm a proponent of human in the loop gatekeeper/reviewer usage of AI, and I'm not able to even consider Chinese models for my own use, and not able to use anything at my day job.
- hereme888 3mo agoWhat's the reason the current SOTA wasn't tested? GPT-5.6-Sol-Max is the actual SOTA.
- deleted 3mo ago[deleted]
- martinjc 3mo agoI wonder why there's so much resistance here against chinese models. Sure at my employer claude is used, but at home? I am just happy using my z.ai sub for 20x i got in September last year, coupled with the 39 dollar tier of kimi. I use them in pi, with a collection of extensions i curated myself for this iterationm of models, and a couple glue extensions we have made. At home i feel way more productive, the speed of my queries are second to none with kimi 2.6, and handing review over to glm5.2 means i can juggle the small models in my brainstorm to commit workflow.
- doodlebyte 3mo ago[flagged]
- p0w3n3d 3mo agoI tried to subscribe to Kimi but it seems it's "sold out"
- hackmack10 2mo agoMeh, it's pretty terrible. Fraction of the amount of requests the other providers give you as well.