5 ms·
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the a
by cmiles8 1mo ago
Every larger company I talk to these days has an active project on moving away from OpenAI and Anthropic to open models. And they’re actively shifting, as the article says, so the threat is far from theoretical.
Unless they both dramatically slash prices then they’re in big trouble. Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO.
However the cold reality for both is that there is zero moat to a model anymore. It’s a pure commodity. Those selling compute and access to open models are gearing up to wipe the floor with Open AI and Anthropic.
- r_lee 1mo agoI knew that there was no real moat from the very start, I mean, these things were close enough from the very start, how could it not result in a race to the bottom, especially as you can't really prevent distillation reliably?
- 0xbadcafebee 1mo agoA race to the bottom is where you lower standards, wages, or regulations to cut costs and attract business. What's actually happening is the opposite: a race to the top. Every model is trying to get better. Simultaneously they also happen to be getting more cost effective, but it's sort of a coincidence. Companies still require very good models, but they are not picking the "absolute best at any cost" anymore, because it turns out "any cost" isn't worth it.
- calebkaiser 1mo agoI support and use open models as much as possible, but I'm not totally convinced that OAI or Anthropic have no moat, even as open models catch up to the frontier. Serving and inference are still hard problems when you're talking about a 2 trillion parameter model. Fine-tuning, if that remains a realistic need for businesses, is also a difficult infra problem at that scale. In the most bearish case, where there is no competitive advantage to using their models, big labs still have an advantage in this area. Maybe there is some threshold where the price/quality math for your standard business tips in favor of smaller models and self-hosting the entire stack. I'd certainly love that.
- jopsen 1mo agoAlso people are people and they will get emotionally attached to claude :)
- epistasis 1mo agoOr ready to get a divorce... The load-bearing seam of the relationship is affection versus annoyance.
- r_lee 1mo agoI think this is happening right now. It's beyond ridiculous and many seem to be praising 5.6 Sol and others, even Grok 4.6 seems to be a breath of fresh air
- epistasis 1mo agoYep, I cut out Claude, and though ChatGPT seems less capable, at least I haven't had it decoherr into unintelligible babble that it can not explain. But the "open" models are right there too, and I'll be using them more in the next weeks as I get better chat interfaces.
- bluefirebrand 1mo agoI bet many people who use Claude at their workplaces still call it "chatgpt" because chatgpt is getting the kleenex treatment and they have no idea what Claude is
- ijidak 1mo agoTrue. The problem is serving is a skill readily mastered by the hyperscalers. That's their MO. All they need is weights to serve. And the open models provide that. OpenAI is relatively well placed in that they have inference chips they've designed and they own compute.
- nutjob2 1mo agoOK, but somehow there won't be companies who will sell you appropriate hardware and a turnkey system to serve inference? Or companies that will help you fine tune popular models? It's not about money, it's about control. Companies have lots of money and want control over their key technology.
- intrasight 1mo agoThere are most definitely is a moat - but it works both ways. The railguards in the models create moats keeping customers out. And the cost to build a modern agentic model is in the 10 figure range and growing. This is an expensive arms race that is going to create moats.
- cmiles8 1mo agoBut most commodities are the same way. It’s super expensive to drill for oil. I need oil and I’m in no position to mine my own because of the massive capital investment. But it doesn’t stop it from being a pure commodity. I couldn’t care less which company drilled for the oil… it’s all the same to me. Models are increasingly no different. OpenAI and Anthropic are a gas station saying “buy our gas for 10x the price!” When the world is looking at them saying it’s just gas, we’ll take the cheaper brand. We’ve tested your gas and it’s really no better than the stuff that’s 1/10th the price. Thats why their present business plan is screwed.
- euroderf 1mo agoSo the business challenge is to balance sizable investments and relatively small marginal costs. Not so different from other digital goods.
- cmiles8 1mo agoFair assessment. The challenge for OpenAI and Anthropic is that they need sizeable margins to pay for the massive costs incurred. Market forces are driving things in the opposite direction and fast. When your competition has a tiny cost base compared to yours and lacks the bonkers future capital commits you made then that’s a terrible position to be in… hence their conundrum.
- hungryhobbit 1mo agoEven "massive costs" is an understatement: these companies have astronomically massive costs! Take Open AI for instance: it has "zero debt" ... and $665 billion to $1.4 trillion in "long-term commitments".
- motbus3 1mo agoThey have good friends in the big ballroom to not allow you to use something cheaper and be locked in on them for your own safety
- nxobject 1mo agoAs a taxpayer, the only silver lining here is that we won't be paying for the ballroom... /s
- simianwords 1mo agoI know what happens in many big companies and not a single one is moving away from Anthropic/OpenAI/SpaceXAI. > Unless they both dramatically slash prices then they’re in big trouble False, they have already done so many times. > Neither of them can afford to do that and both desperately need to convince the street that the opposite will happen if they want any hope at a successful IPO. False, margins are higher and I can have a formal bet that prices will go lower. > However the cold reality for both is that there is zero moat to a model anymore False, LLMs are not fungible and there exists a natural moat. I like the behaviour of Fable, not the behaviour of Opus - the fact that many people speak about this is evidence.
- chrinic630 1mo ago[flagged]
- harrymunro 1mo agoBS - Grok 4.6 is a legit model.
- throwitaway222 1mo ago> lower-middle class Classism is the new racism.
- 23asgh1 1mo agoNo it isn't. The racism culture wars were meant to distract from classes and people are simply becoming aware of the real issues again. Trying to steer them away won't work any more.
- simianwords 1mo agoIts not that complicated, people use Cursor and that comes with the option to use other models but also Grok. Grok Bot is not used within enterprises.
- fatcatsbestcats 1mo agoI don’t know what your sources are, but I work on AI at a Fortune 100 and open models now make up >90% of our internal token spend. Used to be 100% closed before this summer.
- throwitaway222 1mo agoIf companies are really doing this, then we're saying they have no problems spending tens of millions to get somewhat decent TPS and then having their employees complain they are timesliced and getting lots of timeouts because their org has 500 employees?
- usrnm 1mo agoOpen weights != self hosted
- vohk 1mo agoThis is all spitballing, but I'd wager it's a third the type of workloads, a third hedging against your business depending on a single external provider, and a third trust. Not everybody is coding or doing work that lends itself to burning tokens for warmth. Reuters for example seems to be more interested in using it for research, editing and formatting citations and the like. There's only so much of that work that needs doing, it doesn't always need to be real-time, and they probably don't see it scaling exponentially. They also need to be very aware and in control of their model's biases, or they risk it compromising their work output. It's widely expected that all of the major providers will need to - and surely want to - drastically raise prices to justify the ludicrous amount of capital they're burning. Multiple companies have already talked about how their AI costs have exploded, and from what I understand that scale of enterprise is paying API rates. I would be disappointed if big business wasn't having a think about what that liability could look like. It's one thing to be reliant on a relatively "stable" vendor like Microsoft for Windows and Office, another to get AWS sticker shock, and then this is promising to be an order of magnitude worse. Then just plain trust. What if ChatGPT starts recommending your competitors products, or the USA bars export of Anthropic's latest model (again, but for real this time), or they stop serving a model your business now depends on, and so on... That's a lot of risk to leave outside of your control.
- kittikitti 1mo agoI'm not sure why someone hasn't developed a company offering services that distributes AI across all idle or under-utilized VM's and PC's for enterprises in order to serve open sourced models. Outside of the electricity bill, there's no additional expenditure and you get the AI. We've all seen the office spaces where there's 200 empty computers on a floor. Combined, it's something like 500 cores at ~3 Ghz each and around 3 TB of RAM. The networking is already there and software like exo already exists.
- pianopatrick 1mo agoMaybe open AI and Anthropic could just license their models to run on your own hardware. So a fixed cost instead of per token pricing or subscription with limits
- SSLy 1mo agoBut think about the safety!! (/s)
- phoghed 1mo agoFixed costs are already available via PTU reservation and afaik most serious enterprise projects are using this
- apefulsin 1mo agoThis would require them to give you a copy of their models, which would quickly be leaked on the internet, and they know that, so it won't happen.
- Mkengin 1mo agoGoogle is already doing that since last year and the models where never leaked, so I don't think that would even be a problem. https://cloud.google.com/blog/topics/hybrid-cloud/gemini-is-now-available-anywhere https://cloud.google.com/blog/topics/hybrid-cloud/gemini-is-...
- nutjob2 1mo agoThis is the future, because the market will demand it. I think there will be separate and huge markets for models, hardware and compute. That will maximize competition and innovation. Why? Because even Blind Freddy can see the huge usefulness and power of these (and future non-LLM) models and no-one in their right mind is interested in becoming OpenAI's or Anthropic's bitch. Those companies have tickets on themselves. Given the recent behavior of tech companies and the US administration, no one trusts either anymore.
- epistasis 1mo agoThis is exactly true, I get annoyed by Claude one day and switch to something else, and the only thing that's ever keeping me tied towards Claude is the ability to search my old chats easily. But Claude also makes it really hard to do that, so what am I even really paying for? Time to extract all my data, put it into a sqlite with FTS5 and make sure I never rely on the overly-opinionated, low-thinking PMs from these giant orgs again. Of course, that "easy" step has lots of partial solutions like CTK (Conversation Toolkit) or MyChatArchive and I haven't found the perfect one yet, ideally it'd be something that dumped everything into Obsidian or an Obsidian-alike, but surely somebody is working on that? I'd pay $5/month for somebody to solve that problem for me, as long as I still owned the data...
- rpastuszak 1mo ago> the ability to search my old chats easily. I’d try: 1 exporting my data (I imagine it’s common outside of GDPR?) 2 asking Claude to convert it to an easily digestible format :)
- chrisweekly 1mo agoJust use a proxy and log everything.
- epistasis 1mo agoA proxy only catches stuff going forward and it doesn't provide me chat search and chat viewing. What I'm doing instead: syncing coding agent sessions to a central backup location, and for cloud LLM chat providers I'm occasionally exporting data. Still need something to automate that syncing, and provide search and viewing.
- jjav 1mo ago> search my old chats easily Don't rely on chat history. Have it write and maintain summary files that you can import into different sessions, at least for anything important.
- unrented7977 1mo agoClaude by default deletes old chats after a few months. I installed a custom end-of-session hook that throws chat transcripts into a database so that any model can read any other model's chat history. Super easy
- piva00 1mo agoSame experience, all my other friends in the industry report the same; at work we went from a huge push for ChatGPT last year to switching to Claude, back to ChatGPT when it became cheaper than Claude, and in parallel a deployment of open models being trialed with mechanisms to route to other models when needed (and based on pricing). Over time I can imagine us becoming mostly open models on our deployments when hardware is more accessible and the need for expensive frontier models is constrained to very few use-cases that might demand their capabilities.
- Karrot_Kream 1mo agoChatGPT or ChatGPT Work or Codex? Are folks at your company still just using a chatbot to do work? (If so, this is quite surprising, as I find harness-based agent usage to be much much better than chatbotbot agent usage.)
- piva00 1mo agoNo, the full suite: ChatGPT Work, Codex, chat, Deep Research, etc., same for Claude with Code, Cowork, etc. And any new release of model or tool is trialed and assessed. Just too many products to exhaustively list, it's a big tech company, we're trialing a lot under the sun to find workflows, tools, integrations, including a lot of bespoke internal research, there's a large ML department since almost the inception of the company.
- Karrot_Kream 1mo agoMakes sense. If you have the resources, trialing out various models to find the pareto optimal point is worth it.
- disdegeneration 1mo ago[dead]
- novok 1mo agoThe open models are good because of distillation, which the US labs are actively working against via not revealing CoT ever and now you can see with OpenAI Astra 6 not even having a lot of CoT equivalents being emitted as tokens. Once the anti-distillation stuff is in place the open distillation models will probably start having larger and larger gaps. If the companies survive the next few years, which they probably will because they represent too much of US economic growth to allow them to fail, this gap will keep on expanding. Starting from zero without distillation is a lot harder, a lot more expensive and a lot more work. OSS models is what a laggard does to get adoption. China's gov't might keep on sponsoring it as a counter GPU embargo thing, but when gov't get involved, usually the other side gets involved too. As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors.
- achrono 1mo agoYes, evidence is needed but especially for the claim that distillation is what makes these open models good. Serious citation needed. Think about it: even if they distill the shit out of frontier models, the model still gotta learn, right? If anything, as you can see from the K2 Horizon release, aggressive (self-proclaimed) reliance on distillation does not result in a model that has remotely any frontier capability. Try asking K2 Horizon to write iambic pentameter for instance, or even give it the car wash prompt. I tried both these on the Q8 quant for the 7B model and the results were depressing.
- achrono 1mo agoTo whoever downvoted this, it would be helpful if you actually reply with something substantive. The post I'm responding to makes sweeping characterizations, and I challenge it with a relatively good heuristic and indirect evidence, and only get downvoted?
- novok 1mo agoBecause it's evident you didn't engage with the last sentence I wrote properly. "Citation needed" in more words is not a sufficient response. > As for people asking where is the evidence for half of this, you will never have public evidence for most of this, but deduce what the partly hidden parts reveal about the whole and it is fairly obvious, especially if you look at the past behaviors of the governments and other actors. To help you further understand, the Chinese labs will never admit they were distilling until they are better than US labs with their own non-distilling process or some sort of espionage-like public reveal shows it. So you need to look at secondary indicators. Much like how the chinese government lies about their economic stats so 3rd parties use secondary indicators to figure it out.
- lenerdenator 1mo agoThe schadenfreude is that, at least for OpenAI, they were originally set up to make open models. They were set up as a public benefit company and their returns were capped at 100x. They (well, Sam) went out of their way to put themselves in this death march to the IPO. If they had just done what Mark Zuckerberg did with Muse, they're not in this position. Do you know how bad you have to be at the tech business to make Mark Zuckerberg look like a prudent-yet-visionary leader?
- alexashka 1mo agoYes, but also remember what happened to Mark trying to do a currency? You're giving him too much credit if you think he figured anything on his own - this is the guy who thought 'metaverse' was a good idea.
- lenerdenator 1mo agoExactly. Even that guy could read the situation.
- deleted 1mo ago[deleted]
- TitaRusell 1mo agoZuckerberg knows how to wear a suit and answer questions from politicians. You know the actual job of a CEO.
- lenerdenator 1mo agoIf he knew how to wear a suit and answer questions from politicians, there are several regulatory penalties over the last decade that Meta doesn't have to pay. The job of a CEO is to have a vision that is compelling for the market and an ability to execute upon it, thus the "executive" in "Chief Executive Officer". Zuckerberg doesn't have this. What he does have is an absolute majority of voting control over Meta's shares, so his actual ability to do his job doesn't matter. It's almost as if creating incentivizing systems that don't hold people to account is a bad idea.
- Gareth321 1mo ago> However the cold reality for both is that there is zero moat to a model anymore. The moat right now is a) the hardware, b) the electricity, c) the intelligence, and d) scalability. On hardware, it's very expensive to purchase anything which can provide a fraction of the performance of a subscription. Traditional accounting depreciation would imply that purchasing local hardware is a terrible financial decision. On electricity, this is a surprising cost center depending on location. A system with just one 5090 can easily pull 1kW, and to achieve usable performance for a workplace is going to require dozens of machines. This can represent an extra $10-20k in electricity in cheap places. In California or Europe this could be $30-60k per year. As for intelligence, the frontier models from OpenAI and Anthropic are still superior, and they have at least a 3-6 month head start. Distilled models are closing the gap on some metrics, but they still can't compete. That's why they cost so much less. The last major moat is the ability for subscriptions to scale with need. This means easily adding and removing licenses. This is far easier than purchasing extremely expensive hardware (and managing it), and selling it if/when internal demand changes. It's the same reason companies use contractors. The ramp up/down costs are very high. The only real moat that local LLMs have right now is privacy.
- sedansesame 1mo agoPrivacy is non-negotiable for corporate. Even without considering costs or country of origin, we've seen from OpenAI that claims of AI safety are worth less than the (virtual) paper they're printed on. All it takes is one incident, and all your company's internal data will start showing up in public users' chats. You can rely on a contract to prevent this, or you can guarantee it by using a locally hosted model you fully control. When combined with the cost savings and good enough performance mentioned in the article, this can become a huge selling point.
- robrorcroptrer 1mo agoIs this not true for everything then? All cloud services, all Internet providers, everything within M365, OneDrive and Databricks?
- alfiedotwtf 1mo agoI know a few Australian devs who work at places also moving, or already have, from Anthropic… and not because of cost but because of Trump’s edicts to ban non-nationals using AI. Think about how crazy it is for non-US companies to use American AI providers - their marketing boasts that you can treat their models as co-workers, assign tasks, invite them to slack annd video calls, etc. Taken at face value, would you hire someone remote who lived in a country that commonly does random shit like deciding whether or not remote workers aren’t allowed to go to work?
- __rito__ 1mo agoEven if I pay for a true large model, I am going to prefer an open model hosted in Europe/my country, not them.
- fittingopposite 1mo agoThe only moat lives at the Pareto frontier. If you are on the Pareto frontier you are good and can charge money. But the frontier is moving every week so it's super competitive. If you are the quickest innovator, I still believe there is a chance for a working business model for them
- dzonga 1mo agoto me the biggest event more than deepseek launch was when zAI served their latest model on all Chinese chips.
- genxy 1mo agoModel serving is trivial, and inference is just memory bandwidth. The cost of serving will be asymptotic to flash read energy. Having trained on your own chips, that is the impressive part.
- g8oz 1mo agoModel serving and inference will be trivial in the long term, right now it's a real choke point. Being able to do that with domestic chips is an important win.
- siruncledrew 1mo agoIf anything, I’m worried they’ll both increase prices after they IPO and are pressured by investors.