9 ms·
They will be, and that moment is not that far off. We've got the progression in place already: first, large data centers could have performant LLMs, we are now
by pronik 5mo ago
They will be, and that moment is not that far off. We've got the progression in place already: first, large data centers could have performant LLMs, we are now firmly in "a bunch of servers with a couple of H100s each" territory, slowly going into "128 GB VRAM on a MacBook Pro or a Strix Halo". Within the next year, the pattern of "expensive remote LLM for planning, local slow-but-faster-than-human LLM for execution" will become the norm for companies, slowly moving to "using local LLM for everything is good enough". And then we'll have the equilibrium we already have with the "classic cloud": you either self-host or pay for flexibility and speed. The question will be: how much of the current compute capacity craze will local hosting give the kiss of death to and what that means for the market.
- dakolli 5mo agoThis is simply delusional, It cost 20-30k a month to run Kimi 2.6. The tokens are sold for $3 per mm. To sell tokens profitably you'd need to be able to run inference at 150 tokens per second for less than $1,000 USD a month. I don't think people realize how expensive it is to host decently capable models and how much their use of capable models is subsidized. You can only squeeze so many parameters on consumer grade hardware(that's actually affordable, two 4090s is not consumer grade and neither is 128gb macbooks, this is incredibly expensive for the average person, and the models you can still run are not "good enough" they are still essentially useless). People are betting their competency on a future where billionaires are forever generous, subsidizing inference at a 10-1 20-1 loss ratio. Guess what, that WILL end and probably soon. This idea that companies can afford to give you access to 2mm in GPUs for 5 hours a day at a rate of $200.00 a month is simply unsustainable. Right now they are trying to get you hooked, DON'T FALL FOR IT. Study, work hard, sweat and you'll reap the benefits. The guy making handmade watches, one a month in Switzerland makes a whole lot more than the guy running a manufacturing line make 50k in China. Just write your own fkin code people. Don't bet your future on having access to some billionaire's thinking machine. Intelligence, knowledge and competency isn't fungible, the llm hype is a lie to convince you that it is.
- hparadiz 5mo agoPosts like this are so funny to me. I'm staring at a mountain of old hardware right now that cost about $20k ten years ago. I have to pay someone now to come haul it away. What makes you think the current new hardware won't end up with the same fate. > Just write your own fkin code people Bro is nostalgic for googling random stack overflow threads for 10 days to figure out a bug the agent fixes in an hour.
- cindyllm 5mo ago[dead]
- dakolli 5mo agoI'm just saying that agent that can fix your bugs actually cost $100-150 an hour to run and you're getting it essentially for $200.00 a month. The cost of cloud compute actually hasn't gone down for old hardware all that much, it still costs $500.00 a year rent 4 core i7700k that's 10 years old. Don't expect much more valuable hardware, like modern GPUs to deflate in price all that quickly. There's 3 fabs in the world that make ddr7 and they aren't going to be selling their stock to consumers going forward, it will be purchased by datacenters almost entirely and stay in them until EOL. Your brain is going to atrophy (this is proven), they'll raise the price to something thats closer to break even and you'll be forced to pay it because you no longer have those muscles.
- hparadiz 5mo agoThe architectural problems I deal with day in day out leave no room for atrophy. This is just cope.
- platevoltage 5mo agoYou're going to see major cope once that bargain $200/month plan goes away, and every person or company that has embedded these services into their workflows gets to see their actual costs.
- RataNova 5mo agoThe biggest impact of local models may simply be that they prevent remote inference from becoming the only game in town
- reisse 5mo ago> They will be, and that moment is not that far off. It's here, right now. I'm running quantized Qwen and Gemma on a decent, but three years old gaming rig (think RTX 3080 12GB and 32 GB RAM). Yes, it's slow, it has a small context window. But it can (given a proper harness) run through my trip photos and categorize them. It can OCR receipts and summarize spendings. It can answer simple questions, analyze code and even write code when little context is required. Probably I could get a half-decent autocomplete out of it, if I bother with VS Code integration. "128 GB VRAM on a MacBook Pro or a Strix Halo" is already a minimum viable setup for agentic coding, I think. > And then we'll have the equilibrium we already have with the "classic cloud": you either self-host or pay for flexibility and speed. Currently, it works exactly the other way. The cloud versions are orders of magnitude cheaper than self hosting, because sharing can utilize servers much more efficiently. Company can spend half a million bucks on a rig running GLM 5.1, and get data security, flexibility and lack of censorship, but oh it's so expensive compared to Anthropic per-seat plans.
- datadrivenangel 5mo agoIn my experience once you get to ~30 gigs of ram for a model like Gemma4, the rest of the 128g of memory is simply nice to have. The speed and costs are what make it tough though, because its slower and more expensive than the same model served on a big accelerator card, and is going to be worse than a frontier model.
- digitaltrees 5mo agoI wonder if it really needs to be worse. I am playing with the idea of fine tuning a model on my exact stack and coding patterns. I suspect I could get better performance by training “taste” into a model rather than breadth.
- epicureanideal 5mo agoI also wonder about JS only, Python only, etc models. Maybe the future is a selection of local, specific stack trained models?
- pier25 5mo agoHow fast do you reckon most people will be able to afford 128-256GB of RAM?
- Schiendelman 5mo agoOther than this recent spike, it's been trending cheaper continuously for decades. In a few years 128GB will be as affordable as 12GB (what flagship phones have now) is today.
- pier25 5mo agoI'm sure it will happen but I don't think it will be soon. 10 years ago I was using 16GB in my MBP and today it's 48GB. It's just a 3x increase during mostly a bonanza period.
- DennisP 5mo agoFor most of that time, I don't think many people had much use for more ram than that. If demand picks up, companies will provide it. And the Mac Studio was available with 512GB until ram got scarce and they cut the max in half recently.
- pier25 5mo agoThe Mac Studio is a high end computer that the majority can't afford or justify its expense. There's plenty of demand for RAM right now. We'll see how this turns out.
- numpad0 5mo agoIMO that was a really weird choice that everyone seemed to make. DDR5 2x64GB before the spike was like $250. I had not much justification to NOT go with 64GB for my pre-COVID build. It seems that a lot of PC building people are confused too deeply by Intel marketing and fixated on getting the flashiest CPU attainable within budget. Similar things happened with previous AI hype, and some people were using HDD boot drives on GPU rigs and asking others whether low end i7 would cut it. They acted very confused when told that they need SSD and Pentium is plentium. I mean, there is a shortage going on, but when it'll be over anyhow - whether due to all the last three standing filing bankruptcy or CXMT-Huawei starts delivering in shiploads or Kioxia enters the market - and it comes back down to $2/GB, or even $5/GB, just max it out and forget about it for 10 years. Why not.
- elbasti 5mo ago> The question will be: how much of the current compute capacity craze will local hosting give the kiss of death to and what that means for the market. This will depend on how much inference happens for consumer (desktop, local) vs enterprise ("cloud"), vs consumer mobile (probably also cloud). I would assume that the proportion of "consumer, local" is small relative to enterprise and mobile.
- stubish 5mo agoI think the proportion is small because someone has to pay for the cloud services. When phones, PCs and Desktops ship with NPUs whole new markets open up for all that stuff people want but not enough to pay for.
- root_axis 5mo agoYou are greatly underestimating the hardware requirements for productive local LLMs. Research consistently shows that parameter count sets the practical ceiling for a model's reliability. Quantized models with double digit param counts will never be reliable enough to achieve results in the realm of something like Opus 4.6.
- byzantinegene 5mo agoi would argue we don't need anything near Opus to be productive. Sonnet is plenty productive enough
- JumpCrisscross 5mo ago> we don't need anything near Opus to be productive. Sonnet is plenty productive enough For niche applications, sure. For general use, I think the tendency towards the best model being used for everything will–to the model publishers' delight–continue. It's just much easier to get a feel for Opus and then do everything with it, versus switch back and forth and keep track of how Haiku came up with novel ways to dumbfuck this Sunday evening.
- root_axis 5mo agoI use Opus 4.6 as an example because it's the LLM that has been widely recognized by the public as being reliably capable of doing real work across many domains. However, the same logic applies to Opus 4.5 and even previous generations. These models have huge parameter counts and large context sizes, there's no training technique that can compensate for those qualities in small and quantized models.
- wincy 5mo agoWon’t these H100s drop in price in a few years? With the data center build out surely these will become 1/10th the price and you’ll be able to set up a local LLM as good as opus 4.7. Even if the frontier model become more advanced, and memory hungry, you could use the same power usage as your oven to run a current day frontier model as needed? If I could drop $10,000 to have an effectively permanent opus 4.7 subscription today, I would.
- inf3cti0n95 5mo agoCertainly, I don't think Data centers are the way here. I guess, it'll most likely be an AI processing and everything else becoming API. In case of GPTs and Claudes of the world. They'll be just using an Indexing APIs and KB on top of their LLMs.
- dnnddidiej 5mo agoExcept you will want the frontier to compete. Local models are useful but you will always need $$$ to be in the same order of magintude as frontier. And also $$$ for same token speed. The question is would you choose to save $10 a day if it causes your inference to slow down 10x and waste 2 hours a day waiting on stuff.
- emadb 5mo agoDo you think small models will arrive? I mean if I need to write a web application in typescript why should I use a model that knows all the programming languages and it is able to reply to any questions about almost everything? I just a need a small performant model that knows how to write web applications in typescript. That could be very helpful and easy to run on my laptop.
- thot_experiment 5mo agoDepending on your laptop, if your laptop is a Strix Halo or a Macbook with a decent amount of ram, that day they arrived is about 6 months ago, and today if you can run Gemma 31b, you're golden for your basic workslop code. You can do most of it with local models. Heck, for a lot of the tier of programming you might encounter in the average job Qwen 35b MoE is good enough and it can hit 100tok/s on decent hardware.
- driese 5mo agoFor the same reason that a human who is fluent in five languages can probably express themselves better in either one compared to human that only speaks one, while also having a more nuanced understanding of general grammar. From what I know, learning on a more diverse set makes a model better overall.
- amelius 5mo agoThis might be an interesting research question: can you train a model on many languages, and then extract a much smaller model that knows only one language without much loss of quality?
- kelnos 5mo agoHumans brains and LLMs are not the same, though. I don't think your analogy is remotely applicable, even if your conclusion may be correct.
- DrScientist 5mo agoI think it's inevitable that access to good enough LLM models will be democratised. However that's not the real battle here. The real battle is control of information to operate over. While I might have access to a decent model - I don't have the huge integrated databases of everything that companies like Google have, and increasingly governments will accumulate. As a citizen AI operating of these large datasets is where the concern should be.
- xnx 5mo ago> how much of the current compute capacity craze will local hosting give the kiss of death to and what that means for the market. Nvidia and other hardware sellers would love if they could sell a bunch of chips to individual consumers that would sit idle for 95% of its life.
- simooooo 5mo agoEven on a 5090 qwen is really impressive. Felt as good as Claude for little projects.