4 ms·
I love LLMs too, but I am concerned about their cost. They are all still very subsidised. Is there any guarantee that I'll be able to run a Opus 4.8-level model
by dom96 3mo ago
I love LLMs too, but I am concerned about their cost. They are all still very subsidised. Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices?
- EgregiousCube 3mo agoGuarantee is too strong a thing to seek, but healthy competition makes it highly likely that the supply/demand curve will meet at a healthy place. You're always guaranteed that you can stash away the open models!
- talloaktrees 3mo agoCurrently, because of the subsidies from the frontier models, demand is mostly for higher intelligence. If subsidies do end, demand for price efficiency per unit of intelligence will go way up.And because there's many players in the market, this demand should be met by at least some of them.
- byzantinegene 3mo agowould i hire a phd grad if they cost the same as a degree holder? sure, why not. but if they cost twice as much?
- Aurornis 3mo ago> They are all still very subsidised. I think the opposite: I think the frontier labs have good margins on their inference unit costs. We can already see what it costs to run near frontier-size models. There are independent business pivoting to serving these models at reasonable prices and they're competing on OpenRouter for costs much lower than frontier labs. > Is there any guarantee that I'll be able to run a Opus 4.8-level model on my personal computer before the big AI labs decide to hike up the prices? I would bet good money on prices going down significantly, not up. If we get to the point where you can run an Opus 4.8 model on your local computer, it's going to be even cheaper for a datacenter to serve it on their hardware. That means prices crash, not that they're going to rise.
- dgellow 3mo agoInterestingly enough, geohot also has an article covering this: https://geohot.github.io//blog/jekyll/update/2026/06/18/prices-cant-go-down.html https://geohot.github.io//blog/jekyll/update/2026/06/18/pric...
- Aurornis 3mo agoThat's commentary on company valuations. Token prices are going down. Competition is global. A company could choose to keep their API prices high, but if another company comes in at 1/10th the price for 95% of the performance then they won't have many customers.
- dgellow 3mo agoYou’re right, my bad, I read that too quickly
- helloplanets 3mo agoThe subscription based plans are heavily subsidized, but the direct API inference pricing (which larger companies need to pay) is profitable. Using a full Claude Max 20x plan to 100% of weekly usage would easily cost you 2k through the API. While the Claude Max 20x plan is 200 a month.
- Aurornis 3mo ago> Using a full Claude Max 20x plan to 100% of weekly usage I doubt many of their customers are on the 20X plan. Of those, I doubt many of them are using 100% of their weekly usage regularly. Comparing the 100% maximum usage scenario of their most discounted plan against the API cost has been a trap in this conversation since it came out. I bet if we saw their financials it would be a tiny sliver in a pie chart somewhere.
- user43928 3mo ago[dead]
- ekidd 3mo agoYou can maybe run a local Sonnet-4.5-ish-level model (sort of) for less than the price of a new car, even at current massively inflated prices for fast RAM. This is probably not what you were looking for. But it's there. You could share one server between multiple developers. Maybe make a little AI co-op or something, with a pair of RTX Pro 6000 cards? Also, DeepSeek V4 Pro is cheap via any commodity API, and DeepSeek V4 Flash is essentially free at API prices like $0.09/M, $0.18/M out. This is generally not subsidized. For a more practical local setup, Qwen3.6 27B on a used Nvidia 3090 (US$1300) or two is surprisingly nice. It needs clear instructions and you can't use it for hands-off vibecoding, but it's actually quite reasonable for hands-on programmers.
- arjie 3mo agoI’ve got a pair of those cards and DS V4F is incredibly good. I’m happy I did what I did because I like this stuff but if you just want stuff then you are absolutely better off not spending $20k on two of these cards and using the API. This guy is absolutely correct.
- fragmede 3mo agoGLM-5.2 is runnable and downloadable today on a MacBook studio that costs a stupid amount of money. No one can take that away from you except by force though, if you want to set it up today.
- int_19h 3mo agoWe're supposedly getting Mac Studio with 1.5Tb RAM in 2 years. That would be enough to run an Opus-level model. Of course, it will also probably cost somewhere around $50k... But if local AI really does become pervasive, maybe it'll be one of the things people buy on credit, like cars.
- 4ffad 3mo ago"Of course, it will also probably cost somewhere around $50k..." Whats to stop people remotely accesssing this? People already do this when working remotely in finance - they connect to a virtual environment that does their work in spreadsheets lmao. nobody cares about the lag, managers certainly dont care about sub-ordinates complaining about it - the same way nobody will care about a slight loss of quality if the economics make sense. frontier labs are screwed really.