4 ms·
What would happen to Nvidia, Anthropic, OpenAI, if tomorrow someone released an open weights model on HuggingFace that matched performance and accuracy of Opus
by gymbeaux 2mo ago
What would happen to Nvidia, Anthropic, OpenAI, if tomorrow someone released an open weights model on HuggingFace that matched performance and accuracy of Opus 5 running locally on an RTX 5070? That won’t happen tomorrow, but it will likely happen someday… what’s the plan beyond “don’t be the one holding the bags?”
- d_sem 2mo agoI guess I'd like to understand the technical reasoning on how you think an how an Opus 5 could over time fit on an RTX 5070.
- nl 2mo agoIt seems very very unlikely that an Opus 5 matching local model that runs on a 5070 will be released within the next 5 years (I don't want to say "ever"). If it does happen then NVidia will sell a lot of 5070s though!
- martinald 2mo agoNothing would really change IMO? 99% of users don't have anything like a RTX5070 (mobile especially). Even if it did, it still doesn't make much economic sense running a model locally vs on a datacentre. For example, I managed to just about squeeze a Q2 quant of Qwen 3.7 27b on my 9070XT. I get around 60tps decode (slightly faster prefill). _but_ it uses 300W of power to do so. At UK electricity rates of 30c/kWh this works out at something like 42c/MTok. I can get far far better models on openrouter cheaper than that, plus I'm not horrendously constrained on context length.
- protocolture 2mo agoI dunno a lot of things said about AI economics sound like an IBM executive making reassuring statements about their terminal/mainframe business before the personal computer took off. Like even if you run it in a datacenter in this scenario, you could do it on a cheap GPU instance in Azure, you still wouldnt need OpenAI or Anthropic specific clouds. >uses 300W of power to do so. There are plenty of people with phat electricity pipes in their on prem server rooms that have been vacated for cloud. Companies who want the benefits of AI but dont want the risk of sending their data to foreign API endpoints.
- JacobAsmuth 2mo agoYou're suggesting that if a very good and cheap AI model came out tomorrow everyone would rush out to rent Azure instances to run batch size 1 inference on their model?
- protocolture 2mo agoI am suggesting that Azure and AWS would change course and push corporate customers towards more expensive, but more private options.
- Jlagreen 2mo agoThe analogy with IBM mainframe completely ignores Murphy's law which came up and lead to the small and fast chips we have today. But Murphy's law is dead. No future chip will leapfrog easily current chips because we have reached hard phyical limits in chip density and downsizing. Huang's law by Jensen Huang focuses on something else and that is token performance per Watt at scale. Blackwell needs double TDP than Hopper and Rubin again needs almost double TDP on a rack but in the end Rubin will be like 100x token performance per watt on a scaled data center. This means you have more energy need but you get multiples of token performance because you start scaling in the data center. The local chip will never be able to keep up with the data center scaling economics. This is why everyone is so crazy about building data centers because they can see the economocs behind it. What people don't seem to understand if tokens become more available and cheaper then not only more people can use them but a single person can use more as well. Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you? This is why demand will grow exponentially with the growth of token economics. We have seen it for the last few years and much more is yet to come.
- protocolture 2mo agoDo... do you mean Moores Law? >The local chip will never be able to keep up with the data center scaling economics. Assumes the software has been completely solved. >Why should you be limited to 1 AI agent? Why can't have you have multiple agents running on multiple devices daily for you? At some point we cap out the bandwidth of the human to keep up with their mistakes.
- milkshakes 2mo agoinference is the cheap part; training is expensive. what compute infrastructure would train this mythical magic model?
- ElProlactin 2mo agoExactly. If OpenAI and Anthropic didn't have to train new models, they'd (probably) be instantly profitable and with good margins.
- ColdStream 2mo agoThose companies will be quick to copy the tech, inference cost would plummet and there is a greater chance that these companies could make it to solvency. At least in the short term. Long term it might not be so great as consume hardware catches up.
- drivebyhooting 2mo agoInference time scaling means whoever had the most compute has the highest intelligence model.
- lisplist 2mo agoIf you could run Opus 5 on a 5070 then the labs must have achieved RSI at that point
- jimbo808 2mo agoThere’s no reason to assume frontier-level intelligence eventually collapses all the way onto a midrange consumer GPU. In fact, there are quite a few reasons not to assume that (information-theoretic constraints, etc).
- christophilus 2mo agoBut, it could happen for a coding-focused model, or an accounting-focused model, etc. most tasks only need a subset of the total model to be done effectively.
- fooker 2mo agoThere's no information theoretic constraint we know of that prevents this. You will almost surely win a Turing award if you can prove this. It's almost a given that whatever is frontier intelligence today will run on a potato in a few years.
- jimbo808 2mo agoKinda silly to follow your “prove it” challenge with an absurd claim you most certainly cannot prove, much less support with evidence.
- fooker 2mo agoIt was not a "prove it" challenge. I'm pointing out that there's no known information theoretic constraint about the impossibility of frontier AI models being improved to fit/run on a small GPU. Please do not make up plausible sounding science facts.
- jimbo808 2mo agoPlease do not assert I am making a claim I’m not making. Information-theoretic constraints exist. My comment does not require some specific, hard constraint to have been clearly defined, for my point to be valid. If I were to say you could put a motorcycle in my car’s trunk, it would be perfect valid for me to say there are space constraints that make your idea unlikely. The same is true in this discussion, even though I have not computed the exact dimensions of the motorcycle and my car’s trunk.
- fooker 2mo ago> on an RTX 5070 RTX 5070 prices go up ~N times. Nvidia makes more money because it's easier to make these things than it's to make a GB300.
- jkahrs595 2mo agoWorkloads will inflate just as they have been. Remember when llm assisted development used to be good only for a function, then a whole file, then a handful of files, then a code base, then a full stack, etc etc etc. People will claim to have “enough” even though they already have the equivalent of last years capabilities locally.
- notatoad 2mo agoprobably not all that much... the market would dip, just like every time a new open weights model gets announced. but hundreds of millions of people aren't going to immediately self-hosting their own models. the biggest winner in that scenario would be ai providers, who suddenly have a capable model that they can serve much more efficiently. and the incumbents have a whole lot of compute. wouldn't anthropic and openAI just start offering that open weights model at prices that nobody else could compete with?
- zhivota 2mo agoThey could but then their valuation is no longer justifiable, which breaks a lot of things downstream (loans being the biggie). They'd rather lose money than start making money in a non defensible way.
- PaulRobinson 2mo agoI think this might be the core signal that it’s a bubble.
- trollbridge 2mo agoIf I could have shown up somewhere in 2022 with a Mac Studio M1 Max w/ 64GB of RAM running Qwen-3.6-27B or 35B-A3B, I would have pretty much been a demigod - to a degree far more impressive than being able to run Opus 5 locally today. So yes, I think your scenario is likely to eventually happen, but there will be a much more powerful, capable frontier model then.
- JacobAsmuth 2mo ago"eventually" is actually a function of frontier model capabilities. You only get Qwen6-27B when you have Opus 7 producing extremely high quality tokens for them to train on. So the market for local models is always significantly behind the frontier, by definition.