4 ms·
You are greatly underestimating the hardware requirements for productive local LLMs. Research consistently shows that parameter count sets the practical ceiling
by root_axis 5mo ago
You are greatly underestimating the hardware requirements for productive local LLMs. Research consistently shows that parameter count sets the practical ceiling for a model's reliability. Quantized models with double digit param counts will never be reliable enough to achieve results in the realm of something like Opus 4.6.
- byzantinegene 5mo agoi would argue we don't need anything near Opus to be productive. Sonnet is plenty productive enough
- JumpCrisscross 5mo ago> we don't need anything near Opus to be productive. Sonnet is plenty productive enough For niche applications, sure. For general use, I think the tendency towards the best model being used for everything will–to the model publishers' delight–continue. It's just much easier to get a feel for Opus and then do everything with it, versus switch back and forth and keep track of how Haiku came up with novel ways to dumbfuck this Sunday evening.
- root_axis 5mo agoI use Opus 4.6 as an example because it's the LLM that has been widely recognized by the public as being reliably capable of doing real work across many domains. However, the same logic applies to Opus 4.5 and even previous generations. These models have huge parameter counts and large context sizes, there's no training technique that can compensate for those qualities in small and quantized models.
- wincy 5mo agoWon’t these H100s drop in price in a few years? With the data center build out surely these will become 1/10th the price and you’ll be able to set up a local LLM as good as opus 4.7. Even if the frontier model become more advanced, and memory hungry, you could use the same power usage as your oven to run a current day frontier model as needed? If I could drop $10,000 to have an effectively permanent opus 4.7 subscription today, I would.
- root_axis 5mo ago> Won’t these H100s drop in price in a few years Doubtful. The increase in demand is greatly outpacing supply, and all signs point to a continued acceleration in demand > If I could drop $10,000 to have an effectively permanent opus 4.7 subscription today, I would. lol well obviously, but realistically that price point is going to be closer to $100k, with a perpetual $1k a month in power costs.
- wincy 5mo agoCool, thanks for the information. I guess they drive prices down by massively parallelizing requests on say an H100 X8 array? So this is spread across. So if I say, wanted to use it for 8 hours a day in my theoretical world it’d be too expensive. My work definitely wouldn’t pay $100,000 for a server farm even if it’d give an AI to all our employees, you’d have to have engineers, a colocation space, basically all the problems that companies didn’t like and went to AWS for.
- root_axis 5mo agoWell $100k was a generous guesstimate for some time in the future where something like an Opus 4.7 is old news. If we think about the near future, something like Kimi2.6 is within the realm of Opus 4.6 today, but requires closer to $700k in hardware to run.
- Galanwe 5mo agoKimi 2.6 is very close to the Opus family from my experience. Also it does absolutely not require $700k to be able to run locally in an interactive fashion. We are talking more in the range of $10k for a slow Q2 with degraded perplexity, to ~$35k for an acceptably fast 200k context Q4 (quasi lossless perplexity).
- aaronblohowiak 5mo agotaalas!!!
- 5mo ago
- CuriouslyC 5mo agoParameter size gets you world knowledge and better persistence of behavior as context grows. Both of those things can be engineered around to a large degree, and the latest Qwen models show that small models can be quite smart in narrow domains and short time windows.
- alfiedotwtf 5mo ago… maybe we should just teach models how to get their world knowledge from a local Postgres connection! Then the model can be tiny, and it can query to its little heart desires AND run on commodity hardware TODAY!
- segmondy 5mo agoJokes on you. We are already running Deepseekv4Flash, Mimo2.5, MiniMax2.7, Qwen3-397B locally in very affordable hardware. These models are in the real of Opus4.6. For those of us a bit crazy, we are running KimiK2.6, GLM5.1 and more ...
- binyu 5mo agoThey all still fall short of Opus 4.6, definitely though. They are good but fail on extremely complex tasks, in contrast with a frontier model that will keep on trying until it succeeds or exhausts the solutions space.
- julianlam 5mo agoNot by much, and moving goalposts makes for a bad comparison. Local open weight models are already more powerful than frontier models from only a year back. If you believe what you read here, the gap is closing fast.
- segmondy 5mo agofrontier models don't keep trying until they succeed. that's a harness problem and best believe it, the best harness are private and not public.
- binyu 5mo agoIt is much more of a context window size and model capabilities problem. Local models are not even remotely close in solving complex problems, even when used with the same harness.
- root_axis 5mo agoI have two A100s and have been playing with local models for years. There's definitely moments where they are quite impressive, but small context sizes and unreliability become immediately obvious. > For those of us a bit crazy, we are running KimiK2.6, GLM5.1 Yes, those can compare to Opus, but you can't run those unquantized for less than $400k in hardware.
- stubish 5mo agoIt depends on what you mean for 'productive'. Article mainly seems to be about targeting consumer level hardware, such as the Neural Processing Unit you need for a 'Copilot PC'. Windows Recall is (was?) one such local AI application. If Microsoft get their way and my next PC has one, I look forward to using it for 'productive' purposes such as playing games, handling natural language stuff and leaving my GPU free for GPUing.
- thot_experiment 5mo agoFlat wrong. Q6 Gemma 31b feels a lot like opus 4.5 to me when run in a harness so it can retrieve information and ground itself. The gap is not that big for a lot of usecases. Qwen MoE is fast as fuck locally for things that are oneshottable. I have subscriptions to all the major providers right now and since Gemma 4 and Qwen 3.6 came out I haven't hit limits a single time. I'm actually super surprised by the number of things I try with Gemma 4 with the intent of seeing how it fails and then having Claude do it only to come away with something perfectly usable from the local model.
- alfiedotwtf 5mo agoI’m guessing Qwen3.6 for agentic coding and Gemma4 for non-coding stuff?
- thot_experiment 5mo agoNo, exactly the opposite actually. Qwen3.6 is too imprecise for long running agentic tasks. It doesn't have the same ability to check itself as Gemma does in my testing. I keep Qwen MoE in vram by default because there are tons of tasks i trust it to oneshot and it's 90tok/sec is unparalleled, anything where I don't want to have to intervene too much it can't be trusted.
- alfiedotwtf 5mo agoOh interesting. I've read that Gemma 4 is really good for creative stuff, but I'm mostly interested in agentic coding. Unfortunately, each time I use Gemma 4, I just get it stuck in loops.
- thot_experiment 5mo agoThis is probably a precision thing, I think there's a really big difference in long running tasks between q4 and q6.
- 5mo ago
- josteink 5mo ago> You are greatly underestimating the current hardware requirements for productive local LLMs. Fixed that for you. Right now most models produced are based on floating point maths and probabilities, which is "expensive" to do math on. Microsoft has researched 1-bit LLMs which can run much more efficiently, and on much cheaper hardware[1]. If this research is reproducable and reusable outside their research models, this means the cost of running self-hosted LLMs will be reduced by an order of magnitude once this hits mainstream. [1] https://github.com/microsoft/BitNet https://github.com/microsoft/BitNet
- ActorNightly 5mo agoYes and no. The best analogy is the difference between having N senior level engineers working for you, versus having N entry level engineers. With frontier cloud models, you can give a single invocation one task, and it can figure everything out. With local models, you have to manage the inputs and outputs quite a bit more, but you can achieve similar results for tasks you set up harnesses for. They are not as a good at finding the right answer internally from their own weights, but they are very capable of ingesting context and reformatting text - for example, for debugging, local models can debug issues quite well if you give them the error and documentation for a particular feature you are trying to implement.