5 ms·
“The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models hav
by chvid 2mo ago
“The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models have made this year.”
I think that should be the blinking headline - this shows what can be done with consumer hardware.
- AgentMasterRace 2mo agohis 128gb Ram laptop is quite extreme
- simonw 2mo agoIt should just about be usable in 32GB.
- npodbielski 2mo agoIt is. I am running it on R9700
- krzyk 2mo agoOn a consumer hardware it would be nicer. With no GPU/iGPU or a 6-8GB VRAM.
- bakraman 2mo agoRAM is never the issue, it's always the compute power
- spider-mario 2mo agoRAM is not “never” the issue. My iPhone and MacBook Air could both run larger and more capable models if they had more RAM.
- geek_at 2mo agoand memory bandwidth
- tuetuopay 2mo agoQuite the opposite, RAM is always the issue. More specifically, high bandwidth RAM.
- mhaberl 2mo agowhat??? not true! for inference the compute is the last thing we need more of. memory bandwidth is the numebr one blocker, after that the inefficiencies that where introduced with MoE models (and all new large models are made that way) Here is a quick read: https://news.ycombinator.com/item?id=49324600 https://news.ycombinator.com/item?id=49324600
- CamouflagedKiwi 2mo agoIt's absolutely not for these models. There are plenty of consumer GPUs out there with 8 or 12GB VRAM - they are comparatively very fast at inference but just aren't big enough to run lots of the models you want. Also context management is a massive pain.
- DanielHB 2mo agoI run qwen3.5-9B on an RTX 3080 with 10GB of vram. It runs at ~77tk/s with around 50k context size. As soon as I switch to a model that doesn't fully fit into vram it tanks to <10tk/s which makes it unusable for me for most tasks.
- 2mo ago
- aizk 2mo agoGive it 6 months, the capabilities will increase even further.
- deleted 2mo ago[deleted]
- madduci 2mo agoTried yesterday on my own laptop (a UltraCore 7 255H without dedicated GPU,with 32 GB RAM), it wasn't even starting thinking, even on a small context window (65k)
- aphroz 2mo agoI think not much can run without a dedicated GPU
- noduerme 2mo agoWhat's the story with Mac laptops? Worth a try?
- rawland 2mo agoYes. mtplx runs it at 25 tok/sec on a M4 Max with 48GB RAM.
- selcuka 2mo agoThe author tested in on an M5 laptop too: > It feels pretty slow on both the M5 Mac and the DGX Spark.
- cyberrock 2mo agoDense ones like this are more bandwidth-hungry, so you want to try MoE ones like Qwen3.6-35B-A3B (35 Billion params but only 3 Billion Active) or Gemma 4. Unfortunately it seems like we might not be getting a 3.8 MoE.
- madduci 2mo agoTill now I was using successfully Qwen 3.5 and Gemma 4 at a reasonable speed
- mobelkh 2mo agowere you running the MoE models? those perform better speed wise
- a_e_k 2mo agoLike the old proverb: "The marvel is not that the bear dances well, but that the bear dances at all."
- bitwize 2mo agoIndeed. LLMs resemble human intelligence in more or less the same way that the output of the TI-99/4A speech synthesizer resembles a human voice.
- freehorse 2mo agoI also believed that, but seeing qwen 27b overengineering solutions in a bit too familiar way in the article, I started doubting that.
- coldtea 2mo agoNot if an LLM over chat can fool most people they're talking to a human (which it can), where the TI-99 speech synthesizer voice absolutely can not.
- shafyy 2mo agoCan it? I feel like I instantly recognize if I am chatting with an LLM or a human
- tapland 2mo agoCan't even tell if you're real or a bot by reading one comment. Great times!
- IMTDb 2mo agoEmphasis on "feel"
- piva00 2mo agoI was quite surprised on how difficult it is to tell when chatting with an uncensored LLM a friend is running (it's too big to run on any of my computers but he got some B200s). You can input your own "system prompt" to make it behave like a normal internet user and the prose writes very similarly to internet comments with none of the LLMisms from ChatGPT, Claude, Grok, etc.
- ZaoLahma 2mo agoFull agree. I until very recently thought AI tools of today were limited to prohibitively expensive high end hardware hosted in data centers. I was surprised and amazed to get "decent" (with the expectations set right / low) coding performance out of Qwen3.5-9B on a decidedly medium end Radeon 9070 paired with a 5700x3d and 32GB of DDR4 RAM. We can finally reason with and "talk" to our hardware.
- eru 2mo agoYes, and we are still pretty early: AI is still advancing at breakneck speeds, and hardware is too.
- DanielHB 2mo agoIs the hardware really getting better? It feels performance per watt is not getting better at all which is the metric that will matter eventually when supply-demand stabilizes. As it is, it seems the improvements are about making the hardware cheaper (as in capex, not opex). This is just feels from me from what I hear on the news and see on the products though.
- eru 2mo agoSolar power and batteries are getting cheaper and cheaper at the moment. So Watts should become cheaper in the long run. Especially when chips are becoming cheaper (in the capex sense), then you can afford to only run them when power is cheap. Btw, from where do you take the notion that performance per Watt ain't increasing? We are also still using what's more or less general purpose GPU hardware; we could get a lot further if we were willing to specialise more. Which would be the natural avenue to explore, if progress in general purpose hardware slows down. Google is already looking.
- DanielHB 2mo agoLike I said, just feels I have from the consumer-hardware space. For several generations of GPU now most improvements come from packing more transistors into a larger die than packing more transistors closer to each other. GPUs have been getting physically bigger with huge heatsinks and fans to support those bigger dies power consumption. Just compare the TDPs: 2020 RTX 3090: 350W 2022 RTX 4090: 450W 2025 RTX 5090: 575W Bigger dies means lower capex of course, but the similar opex (maybe slightly lower as there is less physical hardware to maintain). I seen some specialized hardware like google's TPUs. Not sure how they compare on performance per watt with GPUs though. Regardless the manufacturing processes are still the same (EUV) which is the thing that hasn't been improving. A fully optimized specialized hardware can at most deliver a single-time linear improvement (that could be very significant, for example 30% is still huge of course) and then little compared to normal GPUs. I don't think renewable power generation is going to massively reduce costs for data centers, especially considering power transmission hasn't meaningfully reduced in cost. If anything the only thing that I think will have significant impact for data centers would be dedicated nuclear power plants physically located right next to the data center. In fact I expect power generation to get more expensive as demand can increase faster than supply can be established. I imagine setting up new solar farms and transmission lines to be significantly harder (as in, takes longer time due to approvals and so on) than new data centers (which requires a single large location and I assume less approvals).
- marcelo-earth 2mo agoI thought the same thing, and I generally do a lot of animation in my work, and the results in motion graphics with Qwen are impressive, I really fell in love with it
- CMay 2mo agoFor me that moment was Gemma 4 12B QAT. You're not suddenly going to start throwing your hardest programming problems at Gemma 4 12B QAT, it is still 15B parameters less. It's more that, aside from pelican art which isn't what local models are for, I didn't see anything on Simon's post that it couldn't assist with or largely succeed at. It can run 80-100t/s on a laptop, can understand images natively and do bounding boxes, read tiny text, understands audio natively as well and can transcribe or translate anything you say, can do accurate long context retrieval with pretty large context windows, tool calling, excellent reasoning and is very token efficient. It's only 7GB including the mmproj or 8GB with MTP. The Qwen 3.8 27B model Simon was using is ~18GB with MTP+mmproj, rather than 17GB alone. The point is not really that you compare these models directly, but that Gemma 4 12B QAT was really a special moment in model releases deserving of a similar reaction relative to its size, but was mutilated by Google themselves, Unsloth and Llama.cpp. The overall appreciation I think we're seeing this year in particular is that people are easily surprised when multiple things are improving simultaneously which produce seemingly exponential changes. It isn't just that models are getting smaller, or that reasoning is getting better, or that speculative decoding is becoming mainstream, or that models can understand audio and images better now, or that they can reliably call tools which expands their capabilities, or that context windows are getting larger, or that accurate retrieval is improved, or that.... and so on. It's all of them narrowing in at once that is starting to make local models incredible and truly useful for far more use cases on the existing hardware people already have.
- runtime_lens 2mo ago[flagged]
- podocarp 2mo agoOut of the loop here. What did Google and unsloth and llama do to mutilate Gemma? I can understand Google shenanigans but llama and gunsmith is kind of surprising.
- CMay 2mo agoGoogle provided incorrect settings and an imperfect template. Unsloth modified the template and then finetuned their own version of the model to optimize for some benchmarks as a means of validating quants. Google and Llama.cpp then adopt template changes by default, so anyone downloading the new model or even using the original model will now automatically be using it incorrectly. Llama.cpp also uses the same inference setting defaults regardless which version of the model you use and some settings are simply defaults it uses for all models. Then even if you account for all of these, you have to be using Gemma 4 itself correctly, which many people do not. All of these little changes and inconsistencies hurt some of the model's original capabilities. Even if you go directly to Google's repo and download the full float 16 weights with the template they have there now, you cannot simply assume you're getting the best results.
- mhaberl 2mo agoI wish we could have better hardware and I think the tech is there for a few years already. I've gone in (too many) details last night with the calcs: https://news.ycombinator.com/item?id=49324600 https://news.ycombinator.com/item?id=49324600
- wejick 2mo agoIn this kind of moment, I really wished hardware manufacturing and demand situation is in much state. Imagine this can be accessible by everyday people with only 6 months hardware market gap. The societal impact would be much bigger.
- jillesvangurp 2mo agoThe other implication here is that this is all software improvements and optimization. There might be a lot more wiggle room for improving quality over time. It seems the model and reasoning quality is improving faster than the hardware currently. The over reasoning that Simon Willison highlights here is a real issue though. I've observed it with some of the OpenAI models as well. They are prone to overthinking and overengineering things. What I would love is models that figure out their own appropriate reasoning effort given a task. I'm spending too much brain cycles worrying on what model speed, reasoning, and quality settings to pick. It's not just a cost concern it's also a time concern. Wasting a lot of time for simple UI tweaks because the model is set to high or ultra or whatever is counter productive. The last few iterations of frontier models seem to emphasize benchmarks and reasoning effort. But of course the day to day reality of many developers is that they are trying to solve relatively simple problems compared to e.g. proving some so far unproven theorems, solving some Nobel prize level problems, etc. I'd love my tools to start making sane choices based on what I ask rather than defaulting to "boil the oceans". These tools need some kind of Auto select. Mostly Ultra is overkill and a waste of time and resources. And of course with local models, keeping simple things local is a nice option. It's nice to have Sol Ultra extra fast as an option in my back pocket. But it's complete overkill 99% of the time. And it's not like most users make good choices here or are even capable of making good, informed choices. The models are more intelligent than the tool UX. Arguably, a local model of very modest size might be able to do better for this specific choice.
- stellamariesays 2mo ago[flagged]
- Bender 2mo agoDo these self hosted models avoid "protecting the user" or protecting big businesses? In other words can I just ask it any question and if it has the answer, I will get an answer rather than telling me it can't answer the question? I ask because Claude is fun for rewriting abandoned code and I am not a proper developer so it's been great for me. Claude refuses to answer questions about science and medicine that stray outside of the officially supported narratives of the AMA and I have issues that have surpassed anything a doctor can do so I am entirely on my own. Will the self hosted models answer such questions or will it also try to put walls or bumper guards around topics?
- CamperBob2 2mo agoOut of the box, open-weight models have guardrails similar to the rest. But unlike the closed models they can be 'abliterated' with varying degrees of success. If you run an aggressive Heretic abliteration of Qwen 3.6 27B you will not generally experience either refusals or an obvious degradation in quality. I personally like the https://huggingface.co/HauhauCS https://huggingface.co/HauhauCS version of Qwen 27B from a purely subjective point of view as a user, but that particular one has come under criticism for reasons that don't necessarily affect its quality or usability. Some of the larger models have also undergone similar treatment, but it's less common.
- dragonwriter 2mo ago> Do these self hosted models avoid "protecting the user" or protecting big businesses? In other words can I just ask it any question and if it has the answer, I will get an answer rather than telling me it can't answer the question? For the most part, open base models are corporate releases with somewhat similar guardrails to commercial hosted models (not quite as complete, because hosted models tend to have a combination of trained and external guardrails applied); but no one is monitoring and trying to terminate your account for using jailbreak prompts, and there are often community finetunes available that (among other things) weaken the trained-in guardrails. Of course, even if the model does answer, it may nto answer according to the particular worldview that produces hostility to the “narratives of the AMA”.
- defilan 2mo agoI totally agree. Is a local model running on your laptop going to outperform the latest frontier model? No, but that's not the point. Many of the use cases folks have can be done well with these newer smaller models. What amazes me is that these keep getting better with existing hardware you have. It's been fun to benchmark and test as these keep coming out.