5 ms·
Every AI subscription is a ticking time bomb for the frontier provider; within a few years we will be running local models as good as today’s frontier models wi
by evo_9 5mo ago
Every AI subscription is a ticking time bomb for the frontier provider; within a few years we will be running local models as good as today’s frontier models with almost no cost burden. The floor will fall out of the enterprise market for all the frontier companies.
- adamgordonbell 5mo agoOr put another way, the frontier models are very quickly deprecating assets, because of the competition in the market. They have to keep getting better to stay ahead of each other and open weight. Which means it's the opposite of a timebomb, the article has it completely backwards, tokens at current level of reasoning will continue to get cheaper. I'm not sure 'local' will be the end state, as hardware needs are high. But certainly competitive forces tend to push profit margins toward zero. Extended discussion on this topic: https://corecursive.com/the-pre-training-wall-and-the-treadmill-after-it/ https://corecursive.com/the-pre-training-wall-and-the-treadm...
- airstrike 5mo agoWell, it's a timebomb for the companies who get paid per token, so the parent is right and TFA is probably wrong
- guesswho_ 5mo ago[dead]
- nijave 5mo ago>within a few years Eventually, we'll see. Frontier models still need some pretty serious hardware which will slowly come down in cost. Smaller models are becoming more capable, which will presumably continue to improve. I think there's still a pretty big gap, though. Claude estimates Opus 4.6 and GLM-5 need about 1.5Ti VRAM. It puts gpt-5.5 around 3-6Ti of VRAM. That's 8x Nvidia H200 @ ~$30k USD each. Still need some big efficiency improvements and big hardware cost reduction.
- throw1234567891 5mo agoOr a single mlx cluster if one can find second hand machines somewhere. Difficult to get your hands on today, certainly, but not impossible.
- snovv_crash 5mo agoQwen 3.6 27b is somewhere around Opus 4. It runs on a 5090, a $2k desktop GPU, at reasonable speeds.
- crazygringo 5mo ago> within a few years we will be running local models as good as today’s frontier models with almost no cost burden Based on what? The RAM requirements alone are extraordinary. No, running large models on shared, dedicated hosted hardware at full utilization is going to be vastly more cost-efficient for the foreseeable future.
- alsetmusic 5mo agoLocal modals are 6 months to 18 months behind frontier. Even if the performance of a cloud model is faster, it's clear that local is catching up.
- greesil 5mo agoHow do you know this? I'm not trying to attack your statement, I am genuinely curious how anyone knows anything about model performance outside of benchmarks that are already in the training set.
- scragz 5mo agousing them you kind of get a feeling for skill level and can extrapolate that better than juiced benchmarks.
- calvinmorrison 5mo agoif that's true - and in 6 or 12 months i can get what i have today, it might not be worth paying anthropic.
- lukeschlather 5mo agoIt is not getting easier to obtain hardware that can run models which are sufficiently useful to undercut frontier models, if anything the cost of such hardware has gone up by 25% or more just in the past 6 months.
- aleqs 5mo ago
- wolttam 5mo agoThere's still going to be plenty of use-case and demand for frontier models running across hundreds or thousands of GPUs. It's just not going to be in the current shape - certainly not accessed by the general public for rote business tasks.
- vb-8448 5mo ago> within a few years we will be running local models as good as today’s frontier Unless there isn't some important breakthrough in hw production or in models architecture, it's quite the opposite: bigger, more expensive and more energy-intensive hw is needed today compared to 1 or 2 years ago.
- ls612 5mo agoAs good as today’s frontier. Gemma 4 today is roughly equivalent to the frontier a year and a half ago at gpt 4o tier.
- antisthenes 5mo agoWhat's the cheapest PC you can buy today that will comfortably run Gemma 4 and everything else you want it to run at the same time? And how many tokens would that buy?
- ls612 5mo agoI run it on my 4 year old MBP and get 10 tok/s. With the RAM shortage buying anything new today is a nightmare but anyone with a reasonably modern Mac could run it at q6 probably. It is mostly a toy as 4o models weren’t really suitable for real work IMO but at least it won’t ever give me a refusal.
- jazzyjackson 5mo agoAt 10toks, are you using it interactively or do you submit a prompt and come back to it later? I always thought it would make sense to just do conversations over email, asynchronously, the model can take all the time it needs and get back to me when it has an answer.
- ls612 5mo ago10 tok/s is around the borderline of interactive being good. I did the math and it is mostly bottlenecked by memory bandwidth, so in the future I can expect to run a similarly sized model on my 4090 once it gets retired from gaming service and get ~25 tok/s which will be very usable.
- YesBox 5mo agoYou'd have a point if Cloud ^tm didnt take off into a multi billion dollar industry.
- slashdave 5mo ago> within a few years we will be running local models as good as today’s frontier models I seriously doubt it. Scaling is already strained (don't buy into the "exponential" hype). And, in any case, the competition will be against the frontier models that will exist in two years.
- christopherwxyz 5mo agoI would readjust your convictions. We are only 2-4 years away from consumer grade immutable-weight ASICs.
- slashdave 5mo agoWe are discussing how rapid development has been, and now you want to freeze your model in silicon?
- rogerrogerr 5mo agoGenuine question from a place of ignorance: what in the silicon pipeline makes it take 2-4years to produce chips with a new model on them? Curious what the process bottleneck is.
- jazzyjackson 5mo agoWithout being an insider, I imagine that most global fab capacity is contracted out several years in advance. You might be interested in the tiny tape out project, which guides you through the process of getting your own design etched on silicon. If you only need larger features and not the next gen single digit nanometer stuff, you may not be so supply constrained. https://tinytapeout.com/ https://tinytapeout.com/
- pjc50 5mo agoI think you could get it down to three months between weight changes, if you can encode it in metal layers only. The remaining limits are the fab lead time, and the cost of a metal respin (hundreds of thousands to millions of dollars depending on process).
- stingraycharles 5mo agoThe economics of local AI just doesn’t make sense. A model like Opus is - supposedly - something like 5T parameters, which is likely something like 3TB of GPU memory. Local models never reach the % utilization that cloud providers have (80%+), and they’re always going to be much better than local models for this reason.
- lumost 5mo agoCapex, opex, quality, and volume are tricky things to balance. On balance, pc/mobile are cheaper to operate than equivalent cloud and on prem deployments. It’s not unreasonable to suppose that in 2 years time an opus 5 quality model will be etched into silicon for high performance local inference. Then you just upgrade your model every 2-3 years by upgrading your hardware.
- jazzyjackson 5mo agoI haven't been following anyone baking models into ASICs, is it not still necessary to pack just as many transistors onto a chip, whether it's an NPU or GPU, ASIC or not you still need to hold hundreds of gigabytes in memory, so how is it cheaper to bake it onto custom silicon than running it on commodity VRAM? (Asking because I don't know!)
- lumost 5mo agoNot my area either! But my understanding is that there are more efficient methods of representing static numbers when you can skip the vram lookup. https://taalas.com/ https://taalas.com/ Is an example startup in this area claiming 16k tok/s on an asic for llama 8b. Qwen has a 27b model at opus 4.5 quality.
- jazzyjackson 5mo agoNeat, thanks for the link
- majormajor 5mo ago
- otterley 5mo agoPeople who are this certain of their predictions should be forced to put real money on them on Kalshi or Polymarket instead of drive-by blowharding on HN.
- watwut 5mo agoMeh, having opinioms should imply necessity to gamble on gambling site. Not even when that site calls itself "market" to create plausible deniality.
- whackernews 5mo agoOooh. You’re hard.
- planb 5mo agoIf that’s true, then it will be even cheaper to provide them as a subscription. Following your logic, every company would be running their own data centers instead of using cloud providers.
- adrithmetiqa 5mo agoI disagree. No one will want to use second rate models when the frontier models reach a specific level of capability. Enterprise will keep paying.
- malfist 5mo agoNo one? When free means I get 95% of the capabilities of something very very expensive, you bet your bottom dollar many many people will choose free.
- xboxnolifes 5mo agoBut its not free.
- upcoming-sesame 5mo agonot every company can pay for the best engineers in the market, some can only afford to pay for cheaper engineers and it's fine same with models.
- intothemild 5mo agoI've spent the last month bringing in a small demo of what the future could be like, running Qwen, Gemma, and Deepseek, behind LiteLLM so we can monitor token usage, and instead of some dumb ass "tokenmaxxing" we're actively trying to get the cost of inference both down, and in-house. Boss is happy, very happy. We're rolling it out more widely now. But this is the future.
- jmount 5mo agoI think this is a good under-represented point. Again and again things that could only run on a mainframe get ported to the personal device level. However it looks like the campaign to eliminate the PC (by pre-buying all RAM) is the counter-stroke.
- himata4113 5mo agoThis is wrong because local models are very expensive, just as expensive as the frontier. It would cost me $300 in normal deepseek v4 pricing (non discounted) PER DAY, but I get it all for $500 worth of subscriptions.
- nozzlegear 5mo agoWhy are you paying $300/day to run a local model? The whole point is that you run them on a machine you already own.
- himata4113 5mo agoNone of the models advanced enough to replace frontier will be able to run on your machine for any forseeable future or at a reasonable speed. 5tok/s is not acceptable. To run deepseek v4 class model, you would need to spend $120k just in gpus.
- aleqs 5mo agoHard agree - the benefits of local/self-hosted models are not just hardware/cost (it might be more expensive at the moment), but what you get in exchange is unnerfed/unstupified models, full cost/usage transparency, optimized/specialized models, privacy/security, etc.
- WarmWash 5mo agoLinux in year 2000 vibes...still waiting to get off windows 26 years later
- claysmithr 5mo agoI agree. The AI bubble is going to pop, people will move to local models, and the datacenters will be abandoned
- voxleone 5mo agoI can only hope that you'll be right someday. As of now, an RTX 3090 struggles to run most of the good local models.
- czep 5mo agoThe economic question is whether the average company will have the time or talent to roll their own models instead of eating the cost increases. The firms in question are exactly the same that have already decimated their teams. Can they so quickly pivot to self-hosted models if their AI workloads suddenly cost them 10x more? I bet most will simply start shoveling themselves deeper.
- tsycho 5mo agoA lot of home computers are capable (with a large margin) to run a large amount of self-hosted services (eg: jellyfin, immich, minecraft, plex, karakeep, ... whatever people want to use). And yet, less than 0.01% of the population (made up number, but I am more likely to be overestimating than underestimating) do so. Running local models to do real work is likely to be another niche hobby.
- throaway198234 5mo ago^^^^^^^^^^^ AI is the future operating system of every computer everywhere
- Ferret7446 5mo agoThat's why cloud died out and everyone is running their own servers right?