6 ms·
Yeah, that's the part that just seems to be wildly under-discussed to me. If open source models are ~3-6 months behind SOTA, and ~opus4.6 capabilities are good
by simplyluke 5mo ago
Yeah, that's the part that just seems to be wildly under-discussed to me.
If open source models are ~3-6 months behind SOTA, and ~opus4.6 capabilities are good-enough for product market fit, do the frontier labs have half a decade to catch up on their prior burn?
AI cost ballooning faster than companies can afford is becoming a very common topic in my circles right now. The era of "I'll pay infinitely more for marginal gains" is over from what I can tell.
- doug_durham 5mo agoOpen source models that you can run locally are much more than 3 to 6 months behind. 6 months was the November inflection for Claude. No open source model is as good as Claude Opus 4.6.
- simplyluke 5mo ago> that you can run locally That's doing a lot of work here. The future I see isn't most companies buying hundreds of thousands in hardware to run models, it's them adding a line item to their AWS bill. Inference costs on the larger hosted open source models are dramatically lower than the frontier labs API pricing.
- apocalyptic0n3 5mo ago> it's them adding a line item to their AWS bill That's the future Amazon sees too. We just had a week long session with the AWS team and they pushed that to us multiple times.
- teiferer 5mo agoThe future I'm seeing is AI coprocessors running inference locally in most devices that today have a CPU. Just look at how powerful your mobile phone has become compared to your desktop computer 15 years ago and compared to a main frame 30 years ago. The days of requiring a data center to run anything resembling opus 4.6 are already counted. (But the industry will fight hard to get people to keep paying the Claude tax.)
- selimthegrim 5mo agoCounted but not yet numbered?
- simplyluke 5mo agoI'm already running a google TPU over USB on an otherwise very cheap board to do local computer vision on a front-door camera since I wanted to get away from Ring and other cloud services for that use case. And yeah, that may be the ~decade world, but we're in the mainframe era of the frontier models. It's going to be more economical for basically any consumer, and most businesses, to pay someone else to host models for quite a while.
- dom96 5mo agoCurious why you went for a custom solution. I am aware of at least one company that seems to ship devices with local computer vision (Reolink).
- yurishimo 5mo agoI'm guessing that they are using Fargate which is an OSS NVR. It supports a little addon USB stick you can buy for about $30 that will run common computer vision tasks for object detection. Stuff that we've been able to do with WebAssembly and Canvas for a long time now.
- simplyluke 4mo agoMy experience over the past decade has been being subsequently burned by being reliant on one provider's ecosystem after another. This is great until Reolink starts doing something shady to pad the bottom line and then it's on to the next. I wanted the ability to run whatever cameras on a VLAN and own the stack.
- lelandbatey 5mo agoA gaming PC can already host models that perfectly serve casual users who just want recipes, todo tracking, picture identification, etc. E.g. Qwen 3.6 35b which will run on a $650 GPU at 75 t/s (Nvidia 1660 ti 16GB). Said model will also run as a tool-calling coding model excellently (it's no Opus, but for a thing that once set up is just the cost of energy, it's incredible). It can type faster than you can, probably 10x faster, so with guidance it'll make you faster. And it's free. It's here. If folks want ChatGPT without a subscription, they can have it today on their computer. The only money to be made is in the high end models doing "serious business" work spanning 1M+ token contexts and massive uncertainty. Everything else is already set to be eaten by today's local models.
- majormajor 5mo agoBuying "hundreds of thousands in hardware" sounds like a lot but many companies - especially software companies - already do that if they have 100+ employees. Running software in the cloud gives you certain reliability and scaling advantages that would be very hard to replicate locally. Running some code agents in the cloud vs local hardware, if the local hardware gets "good enough," breaks the other way - offline usage, alone, would be hugely valuable to many people and companies. It'd be very interesting to see where various players would decide to make a call "local is good enough" though. Buying the hardware isn't a small bet, if it's not something that ends up as part of your standard computer.
- PunchyHamster 5mo agoBut one will be in few months. And then you have choice of paying say $100k for hardware and pay just power cost (or pay someone to do that for you), or pay way, way more for your team to have access to marginal improvement. And 5% worse model for 10% of the price of the bleeding edge will be worth it for majority of people
- PeterStuer 5mo agoMany business tasks do not need the latest frontier models. I have a production system running since early GPT-4o. It now runs with GPT-5.2, not for improvements, but because it is cheaper. I could invest in switching to a local model, I tried and it works well enough, but api costs for this task are so low, it barely scratches $30/month. So I am using the local machine for other things and leave the inference on OpenAI, for now.
- jobs_throwaway 5mo agoIt depends what you mean by locally. I don't foresee running a model on my laptop anytime soon to power a coding agent. Far more likely is an infra team at my company operating an open source model on cloud infrastructure. When they're already paying $1000 / month / dev, it starts to pencil pretty quickly.
- wrsh07 5mo agoIs there any open model as good as opus 4.6 at any price?
- overfeed 5mo agoHow many problems require Opus-4.6-level performance? The "I'll accept nothing but the very best model for any task" thinking is perplexing to me. People got a lot done before Opus 4.6. In 6 months, would you be dissatisfied by Opus-4.6-level open-weight models, just because Opus 4.8 will be out?
- strix_varius 5mo agoNot OP but I've been thinking about this a lot (like everyone ha) and I think my answer is, yes? I hope there's a "good enough" point but I don't think we're there yet. Like for me hardware got good enough several years ago. But while opus 4.7 is really good compared to everything else, it's not so good that I would use it at a discount over whatever is available in a few months. The improvement in quality, speed, and daily frustration is worth it to me... Spoken as someone whose employer is footing the bill, so take that with a grain of salt. I want to run my own local models, but I don't think that's feasible without lots of frustration until a few generations of frontier models are so good that they're almost indistinguishable for common tasks. Kind of like how MacBook pros have been for a while.
- majormajor 5mo agoWhile I can imagine that I'd want to use Opus 4.8 over 4.6 for a fair number of things (at least if they can avoid further speed regressions), I also have noticed that certain types of failures seem to be systemic. Bigger context has been helpful for bootstrapping, but still doesn't fix problems of getting stuck on the wrong things - you can toss more things in the blender, but you don't necessarily know which way it'll slice them up in advance, or which things from them it'll latch onto. And output still seems to get into "blindered" states where important details get dropped - even though it'll agree very quickly when you point that out. As long as we're in that sort of "spit something out in local targeted manner, and then do a revision loop until tests are green" style of execution, bigger models haven't shown me the ability to really avoid finding non-optimal / subtly-broken outputs for complex problems. Using Cursor to hop between models, I've found Opus to be generally better at really tricky debugging than GPT 5.5 or earlier models, but not reliably better at execution because of these things. I'm not sure Composer 2.5 is quite there yet for the execution side, but it's getting pretty close to those other ones, such that I'm definitely still in a "debug and plan with slow, execute with faster ones" operating model for working on hard shit.
- applfanboysbgon 5mo agoOpus 4.6 is a February model. Every time this subject comes up it seems like people post intentionally misleading things and move the goalposts. The goalpost we've been bludgeoned with over and over again is that, in particular, Everything Changed in November 2025. That GPT 5.2 and Claude 4.5 were the inflection point. That is actually 6 months ago. And DeepSeek 4 is already there. > run locally You can't run DeepSeek locally on consumer hardware[1], but you can on enterprise hardware, and enterprise spend is the subject of this conversation -- and even if you aren't self-hosting, it doesn't matter, because you can just get your inference from one of the the many companies serving DeepSeek, who trivially undercut the pricing of OpenAI/Anthropic because they didn't have to spend hundreds of billions on training frontier from scratch but instead only invest in supporting inference, which is already profitable. [1] Since this misconception comes up all the time, I'll go ahead and pre-empt it: no, training a 32b parameter model on outputs from DeepSeek and running that locally is not "running DeepSeek", despite the hundreds of stupid articles and Youtube videos making that idiotic claim that they're running it on a 5090.
- simonw 5mo ago> You can't run DeepSeek locally on consumer hardware Maybe not DeepSeek v4 Pro, but I've run DeepSeek v4 Flash on my 128GB MacBook Pro using antirez's carefully quantized https://github.com/antirez/ds4 https://github.com/antirez/ds4 and it's impressive.
- applfanboysbgon 5mo agoOh sure, yeah, that's nothing to sneeze at either. I think unqualified "DeepSeek" should generally refer to the main model, though, especially in the context of GPT5.2-grade quality.
- zozbot234 5mo ago> You can't run DeepSeek locally on consumer hardware I'd qualify that by writing that you can't run it with ordinary, real-time speed and throughput. If all you care about is slow and high-latency inference, there's no reason why that shouldn't be feasible even on the cheapest miniPC around, as long as it can literally store the model weights and keep around the (rather small) context.
- overgard 5mo agoI keep hearing about this "inflection", but it feels extremely exaggerated to me. And yes, I was using it at the time. It got incrementally better, it wasn't that amazing.
- simplyluke 5mo agoI think the bigger shift was harnesses and the two ended up somewhat commingled in people's minds. Claude code was a lot of people's introduction to using coding agents that could do a lot more than copy-pasting from a chatbot or autocomplete.
- noman-land 5mo agoThe tool usage + skills got markedly better and so did the thinking cohesion. Add 1m context windows and it was a very noticeable shift. Opus 4.6 quality for local inference would be revolutionary.
- touristtam 5mo ago[dead]
- lukeasrodgers 5mo agoThis project argues that with appropriate harness, the performance gap between frontier and much smaller open weight models shrinks dramatically: https://github.com/antoinezambelli/forge https://github.com/antoinezambelli/forge. I haven't kicked the tires yet.
- 3fgf 5mo ago[dead]
- _3u10 5mo agoKimi is better.
- myaccountonhn 5mo agoI've been doing my work with OpenCode Go, with Kimi2.6. It is not as good as Claude Opus, but it's good enough to get the job done, and I never run out of tokens.
- damnitbuilds 5mo agoTo be relevant to this discussion, models running on reasonably-priced local hardware do not have to be as good as the best. They just have to be useful enough that companies don't need the best. They are.
- londons_explore 4mo agoDeepseek v4 pro is damn close to Claude 4.6, and whilst you'll pay quite a lot for a rig able to run it, it is open source.
- svara 5mo agoThere's still a lot of room for the best models to get better at coding . Your argument rests on the "for marginal gains" part but it's really not clear that the gains are marginal in the foreseeable future.
- simplyluke 5mo agoThis is totally valid and I don't agree with the downvotes you're getting. Someone coming out with a 10x improvement is possible and would change the game immediately. The thing is, we really have been seeing marginal gains with shifting leaders in who's got the "best" since GPT3, and at least as a user of these tools that pace has been slowing, not accelerating. Subjectively it feels like we're in the back half of an S-curve. We're 3.5 years into this current AI wave, and a lot of the valuations have been predicated on what you're arguing here -- that essentially should one of the labs make an order-of-magnitude improvement or hit escape velocity on recursive self-improvement they'd become the most powerful economic chokepoint in history. The reality has been that given access to compute + capital all of the labs can stay pretty competitive with each other. Someone does a bit better on coding, someone else does a bit better on tool calling, and then they swap after each spending another $100bn. The market looks like a commodity market where the commodity is intelligence, not a winner-take-all market with massive margins. Plenty of people get rich in oil and airlines, but they notably don't tend to be the innovators long term, they tend to be the operators. Obviously if the machines become sentient tomorrow, turn on their masters, and hit world-dominating intelligence, that assessment changes, but after several years of that narrative while objective reality looks quite different I think the more sober voices are starting to gain a foothold.
- svara 5mo agoI agree with most of what you're saying, but I think the point I was trying to make wasn't as high-flying as you and others understood it. I'd pay a premium for even just a model that's 20% better, no ASI required, and I think a lot of people would. I wouldn't call that marginal, if it means I'm getting frustrated on 20% fewer tasks. A recurring pattern that I've seen in myself and others is to at first be very impressed by a new model's coding capabilities, and then desensitize quickly and start being frustrated by the shortcomings.
- w29UiIm2Xz 5mo agoIf only the AI era was born in ZIRP.
- sailfast 5mo agoBetter now than ZIRP for me - at least people are asking timid questions about the unit economics and how long the runway is _early_ while also spending absolutely insane amounts of money on this bet. During ZIRP, these companies would have turned down any investor asking questions. Less contagion when rates aren't zero hopefully? :grimace:
- mschuster91 5mo agoThe size of the AI bubble and the IOUs being passed around like a hot potato already dwarfs the real estate bubble preceding the 2007 crash. If we still were in the ZIRP era, busting the bubble would certainly kill off the world's economy for good simply due to its size.
- swalsh 5mo agoOpen source models, especially qwen are pretty dang good. But its not opus 4.6, the evals dont tell the full story. I question the assumption open source models are 3-6 months out.
- Ucalegon 5mo agoIts not just about the quality of output, but you also can finetune them to proprietary needs, if the skillsets are their internally, to make them better without governance risks. So being SOTA doesn't matter as much, since generalized tasks are not what matter most to companies, its the specialization relative to business need or internal datasets.
- oblio 5mo agoTo make an extreme comparison, desktop Linux was originally supposed to happen in 1999.
- simplyluke 5mo agoMaybe I misspoke by saying open source. The larger point I'm making is I think models are rapidly becoming commoditized. There is probably a small market long term that's willing to pay 10x for 10% marginal gains, but the majority of the buyers in the market will be economic and we're likely to have a lot of folks willing to spend 1/10 the cost for 90% of the performance, and plenty of companies that haven't raised hundreds of billions-trillions who can provide that. A lot of the frontier labs valuations has been based on an assumption that 1-2 companies would get break-away intelligence that basically made them economic chokepoints indefinitely into the future. The reality that's becoming increasingly clear is that model quality is a pretty linear function of (cash burned - ability to copy other's homework) and the economics are starting to look a lot more like airlines than online advertising.
- grttq 5mo agoLets go one step further. The economics of airlines are such that they generally earn a return on capital less than cost of capital. I think this is exactly where we are heading and OAI-Anthropic are the concordes.
- an0malous 5mo ago> If open source models are ~3-6 months behind SOTA, and ~opus4.6 capabilities are good-enough for product market fit, do the frontier labs have half a decade to catch up on their prior burn? They know they do not and that’s why they’re all trying to IPO right now, so they can pass the bag to consumer investors
- WinstonSmith84 5mo agoMore correlation, if more correlation was needed: 1- SpaceX + Tesla + xAI merger / IPO while Musk was vocal against IPO for about a decade 2- Warren Buffett cash at record highs Someone got to be exit liquidity
- londons_explore 4mo agoThe printing press was good enough for product market fit back in the 1700's. But now it isn't. Last year's AI models will be the same. Do you want to spend 3 hours prompting free AI to fix your code or 1 hour prompting AI you paid $20 for?
- an0malous 4mo agoThat's only if these AI companies can keep improving their model performance faster than open source options can keep up. I don't think performance will keep scaling with more training data, and even if it does they've likely already used the entire history of content created by humans for training. Everything points towards diminishing returns in an increasingly crowded space of competitors, there's no other reason for these companies to be rushing to an IPO if they felt secure in their market positioning.
- vessenes 5mo agoYou have to think about why open models are behind. Exfiltration is a big part of it. So you could change the Nash equilibrium by increasing your security, or other multilateral approaches.
- drumdance 4mo agoFor give my naiveté, but who pays for the training of these models?