9 ms·
Kimi K3, Qwen 3.8, and Anthropic's (Potential) Unravelling
- drob518 3mo ago> More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google. Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.
- cl42 3mo agoThey are trying to diversify into consumer hardware and also are further along in owning data centers. The consumer hardware product, assuming it launches and does decently well, puts them in a very unique position relative to pretty much every other frontier lab. It's more speculative, but I'd argue it can change things quite a bit for them if it works out.
- NitpickLawyer 3mo agoThey've been doing hardware stuff as well. Both "consumer facing" bs like wearables / portables, but also more importantly chips for inference. Having a good model is one thing, being able to serve that model at good speeds and match demand is another. See Anthropic ~6months ago. Or Moonshot, they've already suspended subscriptions to their coding plans, because they can't meet demand.
- eikenberry 3mo agoDidn't the Apple lawsuit throw a monkey wrench into their hardware plans?
- georgemcbay 3mo agoThe article does address what they see as the difference ("its investments in product, consumer experience, site publishing, voice, and hardware are all directions that have clearer moats") That said I think it is pretty easy to make a case that these would-be differentiators are either currently underwhelming or completely unproven (as in the case of hardware).
- drob518 3mo agoRight. Arguing with the original article, not you, my counterpoint would be, “Sure, if they’re able to do that. But so could any of the other model-only competitors.” The best you can say today is that they have announced an intention to do those things, but in no way have they established themselves as being successful, yet. And the model-only Chinese providers have access to lots of e.g. wearable and consumer tech.
- treis 3mo agoChatGPT is synonymous with non technical/work related LLMs. They're amassing a ton of user history. That history improves the product for the user because it has more context into the person. They can feed it back into model improvements and for advertising. You can see a future where a user types in "plan a vacation for me" and ChatGPT coordinates everything from there. Those sorts of users aren't going to switch because model X is 10% cheaper or better.
- drob518 3mo agoI see a future where lots of models can plan a vacation for me, not just ChatGPT. Saying that there is brand equity in the ChatGPT brand feels a lot like saying there’s a brand equity in the MySpace and AOL 25 years ago. Google in particular, via Android and its relationship with Apple to power Siri, has a much better shot at grabbing the “plan a vacation for me” consumer market, IMO. I could easily see OpenAI become irrelevant in 2 years if they stumble at all and don’t keep up with the other frontier models.
- m_ke 3mo agoAnthropic will get squeezed by open models for 80% of the use cases that don't require frontier capabilities and by vertical specific labs for the high value tasks that would (bio, finance, math, etc.), where smaller use case specific models will beat them on cost and speed while matching or exceeding the performance of their largest general models. Even their hail mary of being first to "AGI" will never happen because all it takes is China blockading Taiwan or Nvidia cutting them off to stop them from eating up a large chunk of the economy. There is no scenario where the rest of the world will sit on their toes and let OpenAI or Anthropic monopolize "AI". Too many countries, large well capitalized players and partners / suppliers who could never let that happen. Kimi K3 allows all existing players to restart at the frontier and keep competing with OpenAI/Ant. It also gives employees at these labs a better more lucrative path of starting new labs with fresh books and clean cap tables, building on top of K3 without needing to spend all the capex on pretraining their own models. Plenty of them already vested their stock and would have 0 problems raising 100s of millions of dollars for new labs, making them paper billionaires over night.
- energy123 3mo agoNone of my use cases require frontier capabilities but I still pay $200/month to a frontier lab. I value the additional time saved at more than $200/month. If I had to pay actual API rates, then I'm not sure what I would do, but it would not be an easy decision.
- m_ke 3mo agoSure, but anthropic is charging businesses based on usage now and tried hard to pull Fable from the consumer subscriptions before Sol and K3 dropped. Even now on the $200 plan I use up my Fable credits in a single day and had to start using codex and openrouter for more usage because Fable burns $100s an hour when billed on usage.
- cmrdporcupine 3mo agoYes, the reckoning here will happen in a year or two when the (probably subsidized, maybe?) coding plans become either unavailable or much more costly. It's also already the case that in larger companies people do not have access to these plans as employees and must use API rates. There's also the political angle. When Anthropic and OpenAI held back their top of the line models because of Bessent & Trump's bullshit, and threatened to deny unwashed foreigners like me access... I dropped my Codex plan and made do purely with GLM 5.2 for three weeks before OpenAI finally released 5.6 Sol. Feels inevitable that this will happen again. Or, somebody will come up with a way to serve e.g. Kimi K3 or the new Qwen model in an extremely cheap way. Or DeepSeek releases a competitive model at their cut-throat rates. And then the cost argument just wins.
- bko 3mo agoI think the risk is overstated. For one, on the margin people are willing to pay a lot for slightly better models. I know personally the value the LLM adds to my workflow is considerably more than the $200/m I pay the frontier labs. I have no interest in optimizing that to get it slightly lower. There are a very vocal minority that optimizes this or companies whose LLM expense is marginal, but I think that's the minority (correct me if I'm wrong, curious what their customer base looks like) Also the actual LLM is a tiny portion of the value added. Anyone that tried to build agentic solutions from LLM apis quickly realizes that a huge value is the Claude Code / Codex harness. There are open source implementations like OpenCode but they're not nearly as good. Think about it another way. Consider how much money Microsoft spends on maintaining Excel. There are open source alternatives that have >90% of the functionality, they'll even work w/ Excel files and generate them. Google sheets is probably 99% and available to everyone and better in a lot of regards. But the immense value spreadsheet software produces workers above the $100 or whatever a year makes it so that there is a real moat and no one bothers exploring alternatives.
- uncivilized 3mo agoWhat is your workflow such that frontier models increases productivity considerably?
- cmrdporcupine 3mo agoFor people cranking out SaaS / web services / web pages / "full stack" work I don't see a huge difference. If you're building an optimizing compiler, a CUDA kernel, a database, a high performance concurrent data structure with tricky locking, etc. etc. it's still not really close. Sol 5.6 on high just slays e.g. GLM 5.2 for this kind of work for me. I'm sure K3 is fine for these things too, but I can't afford it at its API rates compared to a Codex coding plan. For now.
- uncivilized 3mo agoRegarding the latter, are you able to obtain good results from an LLM? With graphics programming work, I find LLMs only help in cases where the task would take me 5 mins or so.
- behnamoh 3mo agoNope, I'll still buy Claude because the overall XP is better than Kimi and Qwen who literally copied basic harnesses to make kimi-cli and qwen-cli, respectively. Also, you can tell if a model is genuinely powerful and well-thought-out vs a model that acts like it. It's like Apple vs Xiaomi/Huawei. Sure, you can get a Huawei with bells and whistles, but most people learnt the hard way that those companies just copy the iPhone, so might as well get the real deal.
- palata 3mo ago> so might as well get the real deal. Almost the same thing for twice the price, just for the pleasure of saying that you believe Apple was first?
- Alpha3031 3mo agoTo be fair, the flagship Xiaomi, Vivo, Oppo etc are comparable in price to other flagship devices. You very much pay a premium for the large camera sensors they put in these devices, amongst other things.
- osti 3mo agoLol yet I've used Apple and Android phones extensively and would choose Android every single time.
- InsideOutSanta 3mo agoAs somebody who bought an iPhone the day it became available and has been using iPhones for years, my favorite phone I've ever owned is my Huawei Mate XT Ultimate. People keep pretending that Chinese companies only make second-class copies of American products until it's too late.
- daadx 3mo agopeople buy Apple because of the brand - in particular trust. the vast majority of people do not care about like-for-like phones based purely on features.
- overgard 3mo agoI keep thinking about the Figma thing. If you're unaware, here's the google summary: ---- The Board Departure: Mike Krieger, Anthropic’s CPO and a co-founder of Instagram, sat on Figma’s board of directors. He resigned on April 14, just days before news of Claude Design broke. This sparked speculation over conflict of interest and the use of proprietary product strategy information. Betrayal of Partnership: The launch aggravated the tech industry because Figma relied on Anthropic's models to power its own AI features, and even announced a joint "Code to Canvas" integration. Reports indicate Figma was blindsided by the depth and scope of Claude Design. Market Reaction: The "SaaSpocalypse" thesis—fears that major AI foundation models will rapidly build application layers and cannibalize their own SaaS partners—was realized when the news broke. Figma’s stock saw an immediate 7% drop upon the announcement. ---- I would suggest to people using LLMs: you should be cautious about giving these companies data or relying on them. If you're building an AI startup, there's a very good chance they could decide to directly compete with you if your idea has traction. You're also at their mercy for API pricing etc.
- intrasight 3mo agoThen run your own fine tuned models for your AI startup.
- muldvarp 3mo agoDoesn't that mean you bet against the bitter lesson?
- halfmatthalfcat 3mo agoThis has been a story, age as old as time. How many times have startup founders been told "you're a feature" or "Google/Meta/etc can copy you so fast". Operators just need to stay vigilant of their value add, or they never really had a moat in the first place and were living on borrowed time anyway.
- sieabahlpark 3mo ago[dead]
- 3mo ago
- simianwords 3mo agoTo everyone praising Open weight models, could you answer a simple question? If Anthropic doesn't make money because of distillation attacks, how would they convince investors to invest in them, such that it makes financial sense for Anthropic to train even bigger models? Assuming it is preferable for everyone that we get better models in the future. Distillation attacks remove the financial incentive.
- __bjoernd 3mo agoThis assumes all the open models are just a result of distilling Anthropic models. Which remains to be proven. And if they are, the point remains that Anthropic has a brittle product advantage that users and investors should be cautious about.
- simianwords 3mo ago> investors should be cautious about if they are cautious, what would make them invest in newer bigger models without the expected return? generosity?
- drakythe 3mo agoSunk cost fallacy. So much money has been invested in Anthropic and OpenAI at this point that to declare it a loss and walk away could potentially destroy a lot of VC firms, and a non-trivial chunk of the US Economy.
- simianwords 3mo agoSunk cost fallacy is not relevant to future investment.
- daadx 3mo agowhat an absolute plonker do you know how ROIC is calculated? Good luck hiding your huge sunk cost of bilions and billions in there bro. I wish people who had zero understanding of actual finance would never comment about it. Moreover if they declare their existing assets are bunk, the valuation is marked down, especially after now pricing in failure risk - this is catastrophic for VC's. Again, bro, just be quiet.
- joshstrange 3mo agoI think the jury is still out on K3/Q3.8 and if they are equivalent to Opus or Fable. Benchmarks have gotten incredibly murky and I've tried models that are "the same as X model" and been unimpressed. I've tried other models with Claude Code and with tools like OpenCode or Pi and nothing has really come close to Claude Code using Anthropic's models (mostly Opus 4.8). I'm not saying these other models are trash, just that I'm not quite ready to put them on equal footing to Anthropic or OpenAI models. I think the future of LLM coding (and more) is probably routers that decide on a per-task basis which model to route to. Know when to use Fable/Opus and when to fall back to DeepSeek/Kimi/Qwen, even all the way to local models. Of course that's not in the frontier lab's interest but it feels like there is a lot of low-hanging fruit there. I hate that right now it's pretty much "just use the same model for planning/execution/etc" (without standing on your head).
- joey64 3mo agoSo, what's the most affordable way for a pleb who doesn't own 17 H100s to use Kimi K3 or Qwen 3.8?
- vanillax 3mo agoyou cant. The best you can do is Qwen 3.6 27b with a 24gig ( or cumaltive gpus ) to get to 24gb vram. ala 3090, mac with 36gb ram, amd cards, halo strix amd, dgx spark etc. Lots of youtube videos out there.
- Alpha3031 3mo agoWell, if you're happy with around (as in within an order of magnitude or two of) 0.1 tokens per second... I believe that's around what people are getting when loading MoE weights from NVMe.
- svachalek 3mo agoYou've got to consider how much power that uses though. Depending where you live, some of these providers can serve it for less than you pay for power.
- drnick1 3mo agoThere will be smaller versions in the 10-30B parameters range that can run on consumer GPUs.
- svachalek 3mo agoI haven't seen either of these running outside their creator's services yet, but typically you can watch services like openrouter or nano-gpt for it to show up at a (usually small) discount.
- torginus 3mo agoUpper bound of AI progress - recursive self improvement. In this case AI will be responsible for building better models, making people who own datacenters the winners. Anthropic/OAI is cooked. Lower bound of AI progress - plateau. Progess is slowing, focus is on serving a meaningful peak capability at the lowest possible price. There's been news today that Google is building a Gemini chip with weights baked into silicon. Considering a chip's lifetime of 2-3 years at minimum, and that a 2-3 year model today would be useless today, they're expecting they wont make a similar amount of progress in the next 3. Game is about selling at the lowest margin. Anthropic/OAI is cooked. So their survival rests on the presumption that AI progress will fall between these two extremes.
- in_a_society 3mo agoBaking weights in makes a lot of sense for inference speed and power efficiency and has the added benefit of putting many end-users on the hardware refresh treadmill.
- torginus 3mo agoYes, but how much would be a static Sonnet 3.5 be worth today? Its just about 2 years old. I'm not even going to ask about 3yo models like GPT4.
- amaranth 3mo agoThe first one might not make sense but it gets them the pipeline for when it does make sense in the future. Plus for a lot of things using LLMs they'd be fine with older tech, especially if running it was even faster and cheaper.
- himata4113 3mo agoOkay, but the models today will be useful for a lot longer than sonnet 3.5. They're already more than capable to do nearly anything you throw at them given enough time and human assistance. The next step up is faster, cheaper and better user experience. I have only had two instances where I needed to reach for 5.6 sol and that only totalled around $2.7 in api costs. I would imagine it would look something like this: Ground breaking/novel research -> SWE -> day to day assistant conversations -> chat support bot...
- philipkglass 3mo agoIn November of 2025, I would have said that Anthropic's Opus 4.5 model together with their Claude Code harness was the first and only system where a well-specified software feature could be implemented correctly for me in one shot. Today, I'm about equally happy to use Claude or Codex. And if both of those start to squeeze customers for money or get too zealous about safety it looks like there are going to be plenty of capable open weight models and true open source harnesses to use with them. Even Google might eventually deliver a capable model + agent combination (Gemini 3.1 Pro still seems pretty strong, but Antigravity was inept the last time I tried to use it.) I'm happy if Anthropic's business remains viable as one of several strong competitors. The company's safety-first ethos is driving them to increase refusals and deliberately-built-in ignorance with their newer models. In the long run, Anthropic may be best remembered for accelerating the development of software in general so that other people could build less timid tools.
- himata4113 3mo agoI would rather have anthropic fizzle out and somebody else take their place.
- sim04ful 3mo agoIf there's any hope for AI sovereignty and equality, we would have to either make expensive models cheap to run or make cheaper models do less work. Making the latter happen involves either reformulating work in ways less intelligent LLMs can work better with. Or condensing intelligence into smaller models.
- curious_cat_163 3mo agoI think condensing intelligence into smaller models is the way to go. Also, “smaller” can mean many different things. The cost is not in storing the weights on disk. It is in the power required to do the inference with the “active parameters”. There is increasingly more evidence that those two can be decoupled and more power to those who are pushing on that lever!
- LarsDu88 3mo agoThe open weight, open architecture releases of the past several days has me more convinced that ultimately, the winner will be whoever burns their models to ASICs fastest. The LLMs themselves are capable of doing some aspects of chip design as evinced by the K3 press release. Furthermore, the frontier models are "good enough" for a wide swathe of tasks and will soon hit that threshold for a good amount of software engineering (if not already). Does anyone think we need a Mythos level model to plan a road trip, or give someone tips on making a cake recipe? A Fable 5 model running at 9,000 tokens/s on an ASIC rather than 150 tokens/s on electricity chugging Nvidia GPUs, or even giant SRAM Cerebras or Groq chips could be good enough to meet the majority of demand. Furthermore, if you're an enterprise the risk of data exfiltration and feeding data to a potential competitor like OpenAI or Anthropic is greatly reduced if you could shift to on-prem ASIC deployments. A handful of chips could cover a wide variety of use cases and cover them more securely. There are a lot of corporate use-cases for LLMs that are not frontier math research or coding.
- misiti3780 3mo agois anyone doing this ?
- LarsDu88 3mo agoThe closest example I've seen is ChatJimmy: https://chatjimmy.ai/ https://chatjimmy.ai/ a prototype from Taalas running Llama 8B Scaling this up to 2.8 Trillion (350X increase), will certainly be challenging. If I was younger and had the right background, I'd love to dive into attempting somethign like this
- selectodude 3mo agoAlso a 3-bit quant. Useable but a long long way away from useful.
- carterschonwald 3mo agonot sure about that, but im actively working on designing ultra sparse models that i want to have perform competitively with stuff 100-10_000 times larger. ehich does yield similar throughput. time will tell id it works out
- port3000 3mo agoIt's amazing how quickly Fable went from 'Game-changing model that needs to be banned' to 'Yeah it's alright, but OpenAI is also just as good and there are a couple of good open weight alternatives that are equivalent for almost everything' The hype cycles are shortening, perhaps we really are reaching some kind of plateau this time (famous last words)
- potsandpans 3mo agoAlmost like the company has no credibility with respect to its safety claims.
- rockinghigh 3mo agoOpen-weight models were lagging 4 months behind OpenAI/Anthropic at the beginning of the year. They are now just 4-6 weeks behind.
- yogthos 3mo agoAnd given that Chinese models are closing the gap there are basically two thing that could be happening. One is that they are moving faster than US companies developing closed models, and two that we're starting to hit a plateau for model capabilities where all the easy gains have been plucked, and now it's not really possible to move forward at the same rate on the frontier. Of course, both things could be happening at the same time.
- sdfefcxv 3mo agoThis happened ages ago. But OAI and Anthropic are trying to cash in ahead of their IPO window. I think that window is pretty much gone now.
- yogthos 3mo agoMy prediction is that they're going to angle to become a vendor of record for the government and get bailed out. That's the only path at this point because there won't be any competition from China in this niche.
- davidpapermill 3mo agoI think a big question is whether any of these labs can produce a model that is _ahead_ of Anthropic and OpenAI. A related question is how much they're dependent on the APIs of Anthropic and OpenAI to achieve their results - whether through distillation or other uses. If these models are derivative of Anthropic/OpenAI I would expect performance to be more narrow and progress to be limited.
- sdfefcxv 3mo agoThey dont need to be ahead on performance alone. Its value per unit of currency spent. Financials will ultimately drive decision making. We are already seeing that more intelligence does not correlate with more revenue, for the firm purchasing tokens. If I was OAI/Anthropic I'd be brown and yellow in the boxers.
- zkmon 3mo agoAdd another dimension - users tightening their purse on AI spend as the reality of returns hits them. Large companies might go for locally hosted models.
- chihwei 3mo agoAnthropic needs to stop fearmongering and stop gatekeeping legitimate frontier AI usage from the general public! The AI built by distillation of the civilizational intelligence should not be exclusive to the privileged researchers.
- IshKebab 3mo agoSo many words to say so little... Don't bother reading this.
- warm_soup 3mo agoWell written article. I liked how the author categorizes companies and their strategies into different buckets (not authors choice of words) and/or combination of harness, data centers, electricity, foundational models. I would have loved to read how companies mentioned are pivoting to build their moat
- piazz 3mo ago> This is where OpenAI has an advantage over Anthropic. While its models are trailing Anthropic's in recent months, its investments in product, consumer experience, site publishing, voice, and hardware are all directions that have clearer moats. I was with you up until this point. I don’t think OpenAI has any more of a substantial product moat than Anthropic; if anything, the Claude / mythos etc brand is a valuable asset that OpenAI lacks. Yes, many of the elite HN engineer always online types have come to prefer Codex, and but if you actually talk to regular engineers in industry, agentic coding is simply still synonymous with Claude Code. And for the non-engineering uses, Claude is so much more pleasant of a conversational companion than any of the GPT line, and I suspect is this baked deeply into the model, otherwise OpenAI would have closed this gap by now.
- SubiculumCode 3mo agoHow much is Anthropic's price due to inference cost or extra margin they can get away with by having the best model?
- spaceman_2020 3mo agoDario can reverse this by going on the podcast circuit again and threatening everyone with 75% job losses due to (his) AI this time Ramp the number up to 85% if that doesn’t work If it still doesn’t work, go nuclear and target 100% job losses language
- athrowaway3z 3mo ago> Once a model is built, the biggest cost is inference Something i cant find any reliable data for, but would help for a sense of scale: How much use before its equal to training? I.e. assuming you have the training data and setup, and we only care for compute - How hours of using eg Kimi K3 / Fable, for it to equal the compute required to train it?
- philipportner 3mo agoif you assume that training requires about 3x the compute of inference (one forward pass, one backward pass, parameter updates), and we take DeepSeek-V3 since their numbers are public. they used ~14.8 trillion tokens with about 2.66 million GPU hours. 14.8 * 3 = 44.4 t inference tokens. obviously, this is back of the envelope math, but at 100t/s you would need like ~14k years. scale this to >100k GPUs and your in the hours to a couple days range.
- yalogin 3mo agoThe bigger question for me is , at what point does investing in higher capability general models will stop showing the ROI? For example - What percentage of workflows require this new highly capable model? How much of it can be replaced with the software tooling around it? What I mean is if the software tooling can optimize the query over a few iterations does it get the same output as from a single shot high capability model query?
- dgellow 3mo agoWhat ROI? Nobody is currently making money outside of the hardware manufacturers and hyperscalers
- deleted 3mo ago[deleted]
- a13n 3mo agoDefinitely feels like there’s a particular narrative being pushed on HN today.
- Havoc 3mo agoOAI feels far more cooked than anthropic. One of them heading to IPO and the other opting to not show their books should tell you everything
- sinuhe69 3mo agoI have a problem with the cost per task metrics of Artificial Analysis. We don’t know how they calculate it exactly. But recently, cost per task has become the most discussed topic. The logic is basically: if model A achieves 55% on benchmark X and model B 60%, but the cost per task of A is 50% cheaper, people would choose A instead of B. But that implies that all output of the less intelligent model A is usable, perhaps only a bit worse than the output of B. But what if the output of A is unusable, or it can only deliver usable results in 1 out of 5 tries? In such cases, the user will have to rerun the task and it will very quickly double or triple the cost and makes the old average number misleading! I would argue the retry and flaky cost will be many times bigger than the average token cost and that is the true cost the users have to bear. AI-Benchy [0] (admittedly a one man benchmark) shows a much different figure than the numbers of Artificial Analysis. Opus 4.8 cost per task according to AA is $1.80 and Kimi K3 is $0.94$. According to AI Benchy, however, the *cost per successful task* of Opus 4.8 is 10.7 cents vs 19.4 cents of K3. The number of correct tests and pass rate of Opus 4.8 is also higher than Kimi K3. So on a cost-per-usable-result basis, Kimi K3 is actually pricier than Opus 4.8 — the opposite of what AA’s headline number suggests. Thus, I don’t know if I can believe the numbers of AA or we need to track the cost ourselves. [0] https://aibenchy.com/compare/anthropic-claude-opus-4-8-medium/moonshotai-kimi-k3-max/ https://aibenchy.com/compare/anthropic-claude-opus-4-8-mediu...
- vsilent 3mo agoMiMo 2.5 Pro works just fine too