10 ms·
Running local models on an M4 with 24GB memory
- sbassi 5mo agoA useful data to know about this setup is how many tokens/sec generates.
- JBorrow 5mo agoIt’s started in TFA
- NDlurker 5mo agoYou can't expect someone to read 4 paragraphs into an article before commenting
- kennywinker 5mo ago@grok is this true?
- DrBenCarson 5mo agoSorry, @grok is offline after declaring himself MechaMussolini earlier today
- NBJack 5mo agoI'm puzzled. The M4, as far as I know, doesn't have 24GB. Did the author mean a M40?
- spoonyvoid7 5mo agoM4 = M4 Macbook Pro
- teaearlgraycold 5mo agoOr Air
- sertsa 5mo agoM4 Mac Mini w/24GB sitting right here on my desk.
- NBJack 5mo agoThanks; I assumed the author was talking about an Nvidia Tesla M4 (hence my confusion and assumption that they meant the M40 series, which has 24GB of VRAM).
- tra3 5mo agoThere’s definitely an option with 24 gigs of ram: https://support.apple.com/en-ca/121552 https://support.apple.com/en-ca/121552
- NBJack 5mo agoAh, thank you. I was assuming a Nvidia Tesla M4.
- canpan 5mo agoRecent models (Qwen 3.6 and Gemma) can really do coding locally. Feels like SOTA from maybe a year ago? But you would want about 32-40GB total memory. 24GB is just a bit short of that. A gaming PC with 16GB graphics card and 32GB RAM brings you very close to a usable coding system.
- DrBenCarson 5mo agoHow are you using that RAM with the GPU?
- canpan 5mo agoLlama.cpp with automatic offload to main memory. You can also use Ollama, it is easier, but slower.
- reverius42 5mo agoFor those who want a GUI, LM Studio does this too (with llama.cpp as the backend I think). I'm getting great (albeit slow) results with Qwen3.6-35B MoE on 8GB GPU RAM, 40GB system RAM.
- ai_fry_ur_brain 5mo ago[flagged]
- deleted 5mo ago[deleted]
- solenoid0937 5mo ago> Feels like SOTA from maybe a year ago? Agree but only for small projects. SOTA from a year ago still wins on larger projects
- wktmeow 5mo agoThat’s the exact ram/vram combo of my desktop - what model would you suggest for that gaming pc setup?
- rtpg 5mo agoWhat kinda harness do people use with these local models? I am quite happy with the Claude Code permission model and interface in general for coding stuff (For chat-y interfaces I have no real opinion)
- sourc3 5mo agoI am running qwen 3.6 9b quantized model on my m4 pro 48gb and it is barely useful to do some basic pi.dev/cc driven development. I think 128gb desktops are the sweet setup to actually get meaningful work done. However, getting your hands on one of these machines is difficult at the moment. As much fun as it is to run these things locally don’t forget that your time is not free. I am slowly migrating my use cases to openrouter and run the largest qwen model for < $2-3/day with serious use for personal projects.
- hparadiz 5mo agoHow does it (the openrouter version) compare to ChatGPT 5.5 or Claude Opus 4.6?
- sourc3 5mo agoGood enough. It gets 60-70% of the work I need done for a lot less $ (keep in mind I am using these for personal projects that doesn’t generate revenue). If I was using it with the hopes of making money I think I would just use Codex at this point.
- carbocation 5mo agoWas the choice of such a small model driven by a desire for high tok/sec? I ask because an m4 pro 48gb machine can run larger models (if model intelligence is the thing that would make it more useful).
- sourc3 5mo agoYes that was my goal. Also noticed a huge performance gain going from ollama to mlx. Your mileage may vary.
- elij 5mo agoI'm using the 30b MOE model on same spec with 65k tokens as a sub agent with tooling and it absolutely writes decent code. The dense 9b I agree wasn't great.
- deleted 5mo ago
- nu11ptr 5mo agoStill trying to understand if a Macbook Pro M5 Max with 128GB is likely going to be able to run coding models well enough that I can cancel my Codex, or even go down to the $20/month plan.
- guessmyname 5mo agoA 128GiB MacBook Pro in Canada is what, north of CAD $11k after tax? That’s around USD $7k. At $20/month for a cloud AI subscription, you’re looking at almost 30 years of service for the same money. How long do people realistically expect a laptop to stay competitive with SOTA local models? Especially in a space where model sizes, context windows, and inference requirements keep moving every year. And even if the hardware lasts, the local experience usually doesn’t. A heavily quantized local model running at tolerable speeds on consumer hardware is still nowhere near frontier hosted models in reasoning, coding, multimodal capability, tool use, or reliability. The economics just don’t make sense to me unless you specifically need offline inference, privacy guarantees, or low latency for a niche workflow. Otherwise you’re tying up $10k upfront to run an approximation of what you can already access through a subscription that continuously improves over time. You could literally put the difference into index funds and probably cover the subscription indefinitely from the returns alone, even accounting for gradual price increases.
- nu11ptr 5mo agoYou are assuming I'd only get it for that. That would probably just be the straw that broke the camels back, but I'm already thinking about a purchase even if that doesn't work out.
- tom_ 5mo agoBut what if you were going to buy a laptop anyway? Obviously you can't do anything with less than 64 GBytes these days, so the question is just whether you go for the jump to 128. In the UK, it's currently an extra £800 to get a 128 GB vs the 64 GB equivalent. So that's more like 3 years of Claude - I think? - assuming current prices stay the same. Or: you might just feel like £800 isn't an unjustifiable amount of money (one way or another), and tick the box, on the basis that it might just work out. As the saying goes, in for 459,900 pennies, in for £5,399...
- nl 5mo agoI think it's useful to be realistic about what you can do with a local model, especially something as small as the 9B the author is using. A 9B model is around the level of Sonnet 3.6 - it can do autocomplete and small functions but it loses track trying to understand large problems. But the are interesting and fun to play with! I do a LOT of work on local agent harnesses etc, mostly for fun. My current project is a zero install agent: https://gemma-agent-explainer.nicklothian.com/ https://gemma-agent-explainer.nicklothian.com/ - Python, SQL and React all run completely in browser. Gemma E4B is recommended for the best experience! This is under heavy development, needs Chrome for both HTML5 Filesystem API support and LiteRT (although most Chromium based browsers can be made to work with it) It's different to most agents because it is zero install: the model runs in the browser using LiteRT/LiteLLM (which gives better performance than Transformers.js), and Filesystem API gives it optional sandbox access to a directory to read from. It is self documenting - you can ask questions like "How is the system prompt used" in the live help pane and it has access to its own source code. There's quite a lot there: press "Tour" to see it all. Will be open source next week.
- ai_fry_ur_brain 5mo ago[flagged]
- nl 5mo agoI think knowledge is power. I think that the more people who try local models (especially the larger ones) the better. I sometimes get the impression that many people claiming that local models are as good as frontier models work in "token poor" environments. If you can't build large-scale programs using at least Opus 4.5+ then it's difficult to compare. They compare something like Qwen 27B with Sonnet and see that it is nearly as good, but miss that the frontier models are a lot better. That knowledge is power, too. I personally can help making local models more accessible. I can't make Opus cheaper.
- bachmeier 5mo ago> I sometimes get the impression that many people claiming that local models are as good as frontier models work in "token poor" environments. If you can't build large-scale programs using at least Opus 4.5+ then it's difficult to compare. I sometimes get the impression that people posting comments on HN don't realize that LLMs do more than vibe coding.
- soganess 5mo agoGetting so close to good! I consider Gemma 4 31B (dense / no MoE), the new baseline for local models. It's obviously worse than the frontier models, but it feels less like a science experiment than any previous local model I’ve run, including GPT OSS 120B and Nemotron Super 120B. On my M5 Max with 128 GB of RAM and the full 256K context window, I see RAM use spike to about 70 GB, with something like 14 GB of system overhead. A 64 GB Panther Lake machine with the full Arc B390, or a 48 GB Snapdragon X2 Elite machine, could probably run it with a 128K to 256K context window. Maybe you can squeeze it into 32GB (27.5GB usable) with a 32K context window? Even last year, seeing this kinda performance on a mainstream-ish/plus configuration would have seemed like a pipe dream.
- discordance 5mo agoCould you please share your time to first token and tok/s?
- ls612 5mo agoI’m on an M2 Max and get 10 tok/s with Gemma 4 8bit MLX
- isomorphic 5mo agoM4 Pro 64GB (14 CPU / 20 GPU), Gemma 4 31B Q4_K_M GGUF, LM Studio: time to first token 0.92s, 11.56 tokens/s. Edit: For comparison with the other poster, same setup as above, but with Gemma 4 31B Instruct 8bit MLX (not sure if exactly the same model): time to first token 4.62s, 7.20 tokens/s; with a different prompt, 1.17s and 7.24 tokens/s.
- zozbot234 5mo agoCould you (or anyone with the same hardware) try antirez's ds4 and report how gracefully it degrades with only the 64GB RAM? Obviously it's going to be dog slow at best for any single inference flow, but can you meaningfully improve on that by running many sessions in parallel? (Ideally you'd need roughly on the order of model sparsity in order to get meaningful sharing of MoE weights, but whether that's genuinely achievable is anyone's guess!)
- Ngraph 5mo ago[dead]
- zoomuser 5mo ago[dead]
- reillyse 5mo agoso, interested how many people are running higher end AI models locally? Figure if I'm spending $800/month on tokens I can build a pretty beefy local machine for the cost of a few months spend - what is people's experience with say a $5k server custom built (and only for) running an AI model.
- entrope 5mo agoYou will likely have to compromise on memory bandwidth or capacity under a $10k price. The Radeon R9700 has 32 GB of VRAM and is pretty cheap (~$1500 right now), which is what I primarily use. My home desktop has 128 GB RAM and my laptop has 96 GB RAM, but bandwidth limits make most models slow on those CPUs. Models with multi-token prediction are somewhat usable on them: Nemotron 3 Super runs reasonably well on my desktop but does poorly on agentic coding that I've given it; my laptop can run Qwen3.6-27B reasonably well with a version of llama.cpp that is patched for MTP support; but usually I run Qwen3.6-27B on my R9700. vLLM might support two or three R9700s on some OS, but I've not been able to get it to run at all with Ubuntu 26.04: system ROCm version is apparently different than what's in the container images, and system OpenMPI v5.0 finally removed C++ bindings that were deprecated in 2005 but are linked from some Python wheel that vLLM (probably indirectly) imports. If you are spending $800/month on tokens you are likely to notice degradation for local models compared to near-frontier models. The models I can run locally are consistently worse than Claude Sonnet 4.6 (again for the work I give them), although Qwen3.6 does feel almost like magic for its size because it can do a lot. The really big open-weight models should be better, but they want 200+GB RAM, which will need a correspondingly expensive CPU.
- adornKey 5mo agoI'm running a server in the 5K-league. And the results are very good. I get about 150 Tokens/s from Qwen3 for coding. And about 50 Tokens/s from the newer non-MoE Qwens. I wouldn't bother with less than 32GB of VRAM. With 16GB you can already run something usable, but 32GB gives you much more power. 9B and 14B are only interesting if you want to tune models yourself. The sweet spot now seem to be around 27B-35B.
- 2ndorderthought 5mo ago
- quacker 5mo agoI could have used this article before I spent the weekend arriving to the same conclusion! Same laptop, and my contrived test was having it fix 50 or so lint errors in a small vibe-coded C++ repo. I wanted it to be able to handle a bunch of small tasks without getting stuck too often. GPT OSS 20B was usable but slow, and actually frequently made mistakes like adding or duplicating statements unnecessarily, listing things as fixed without editing the code, and so on. Qwen 3.5 9B with Opencode was much faster and actually able to work through a majority of the lint warnings without getting stuck, even through compaction and it fixed every warning with a correct edit. I tried 4bit MLX quants of Qwen 3.5 9B but it eventually would crash due to insufficient memory. I switched to GGUF, which I run with llama.cpp, and it runs without crashing. It is absolutely not comparable to frontier models. It’s way slower and gets basic info wrong and really can’t handle non trivial tasks in one go. I asked it for an architecture summary of the project and it claimed use of a library that isn’t present anywhere in the repo. So YMMV, but it’s still nice to have and hopefully the local LLM story can get much better on modest hardware over time.
- solenoid0937 5mo ago> It is absolutely not comparable to frontier models. This is not said often enough. Yes, local LLMs are great! But reading most HN posts on the subject, you'd think they're within reach of Opus 4.7. There is a very small, very vocal, very passionate crowd that dramatically overstates the capabilities of local LLMs on HN.
- HDBaseT 5mo agoAt least in my experience, local models are very far away from models like Opus 4.7 or ChatGPT 5.5 in coding and problem solving areas. I find them useful in basic research and learning and question asking tasks. Although at the same time, a Wikipedia page read or a few Google searches likely could accomplish the same and has been able to for decades.
- darkstar_16 5mo agoI think you're doing it wrong. Use the frontier moddels for the research, planning etc and once you have a plan give it to a local model for implementation.
- tjpnz 5mo agoHow about a M4 with 16GB of memory?
- spike021 5mo agoI'll have to try some more. I've been playing with gpt-oss 20b on my M4 24GB but it hasn't been the best experience.
- BubbleRings 5mo agoPeople do use SOTA LLM’s for other things besides computer programming. For instance, if you are an independent inventor trying to write a patent while keeping your patent lawyer expenses to a minimum, you want to write as much of the first draft(s) of the patent as possible yourself. (You’ll save billable hours with your patent lawyer, and you’ll end up with a better patent because you’ll communicate your innovations more clearly to your lawyer.) However, and this is the big thing, you absolutely do not want to be asking a SOTA LLM for help with the language in your patent application. This is because describing your invention to a web based LLM could be considered a public “disclosure” of your invention, which, (after a one year grace period goes by), could put your invention in the public domain, basically… and thereby prevent you (or anyone else) from being able to ever patent the invention. Plus, you know, a random unscrupulous employee at the SOTA company could be reviewing logs and notice your great idea, and file a patent on it before you do. Remember, the United States patent office went to “first inventor to file” in 2013. Oh and don’t take legal advice from random people on the internet by the way.
- dempedempe 5mo agoIt takes people years to learn how to write a good patent. If you gave your lawyer your attempt at writing your own patent, they might use the info to understand what you want (you're right about that), but a good lawyer would probably just start from scratch. Imagine you're a contractor. You have a client who knows nothing about software development that wants you to write some software for them. They give you some code they generated with an LLM to get you started. Would you use the code or start over?
- rapatel0 5mo agoI got qwen3.6:27B running on my 4090 (24GB) with ~128K context leveraging some of the recent turboquant/rotorquant memory optimizations for activations. Highly suggest going up to that. the q4_xl+rotorquant combo is pretty good. Some reference code if you want to throw your agent at it. https://github.com/rapatel0/rq-models https://github.com/rapatel0/rq-models
- dmichulke 5mo agoForgive my ignorance but aren't they already on huggingface? I assumed turboquant optimizations are already everywhere - in llama-cpp, or the quantization machinery of unsloth and the likes.
- rapatel0 5mo agoI forked it to also add rotorquant. This is a specific optimization that uses clifford rotors instead of static compile time random purmutation to store the activations. Reduces space and parameter count for the storage.
- altruios 5mo agoWhat is your exp on performance +40k tokens? I've not gone past that as I've heard reports that were problems start to arise. I'd be happy to know your experience in that regard.
- rapatel0 5mo agoI'm super happy with the performance, I generally run with 2 parallel slots so I only get about 128K context window. My experience with all llms is that they get more forgetful if you use the full window. (256-512K is the sweet spot for frontier models, 128k works for me with this current qwen)
- shouvik12 5mo ago[flagged]
- bluequbit 5mo agoThe site does not have ssl. Please can you enable it so that I can read the article?
- MinimalAction 5mo agoWell, but if I have a MacBook Air M4 with 16GB, I don't know what useful models can I run.
- jen20 5mo ago`brew install llmfit`
- prettyblocks 5mo agoIf you use lmstudio it will tell you which models will fit in what you've got available.
- ThomasBb 5mo agoBeyond the models getting better; there are still huge gains available in the inference engine side with new tricks like Dflash, MRT, turboquant - for some usecases these can multiply the speeds. There are even some model specific optimized kernels like for DeepSeek 4 flash that seem wild. Makes me feel we are nowhere near the optimum yet. Examples: https://dasroot.net/posts/2026/05/gemma-4-speed-hacks-mtp-dflash-local-inference/ https://dasroot.net/posts/2026/05/gemma-4-speed-hacks-mtp-df... https://x.com/bindureddy/status/2052982206344409242?s=46 https://x.com/bindureddy/status/2052982206344409242?s=46
- isaisabella 5mo agoI'd rather spend thousands dollars on a Mac than subscribing API. The local model allows me to do my work any time and anywhere, without worrying about privacy leak.
- claysmithr 5mo agome too. plus, I don't like the idea of needing massive datacenters, it's not good for anybody
- losvedir 5mo agoDatacenters are more efficient, though, because of batching.
- kristianpaul 5mo agoGood to keep hideThinkingBlock default, is on purpose to be able to steer de model.
- y42 5mo agoHaving an M3 with 36 GByte I was under the assumption, that I can utilize like Qwen and similar models. It's quite easy to set up, you can use pi or hermes for CLI access, or "Continue" to use it in VS Code. You can choose between omlx, Ollama and even more to run the model itself. It's no rocket science, but the results are also not satisfying. I use it occassionally for very easy tasks, fix typos or update meta data in blog posts. So yeah, it improves productivity. But coding-wise it's far away from Codex, Claude et al.
- stuaxo 5mo ago"What does work is a more interactive workflow where you’re clearly communicating with the model step by step, and giving it a lot of guidance. I’m sure that sounds pointless to many of you, why use a model where you have to babysit it as it works, but I actually found that it encouraged me to be more engaged. " This sort of thing is key to knowing what's going on and bit having your brain fully atrophy.
- PAndreew 5mo agoCritics are (rightly) pointing to the fact that these models are not on par with SOTA for complex coding tasks. But many seems to forget that a large part of white collar office work is Excel crushing, file moving, translating dry legal documents, e-mail drafting, PPT drudgery, etc. These are absolutely doable with 30-35b+ models with the added benefit of keeping company data private.
- tjoff 5mo agoArguably excel and legal are much worse than code because catching the mistakes can be much harder. Case in point, JPMorgan London Whale incident, $6 billion loss caused by an excel error...
- PAndreew 5mo agoYes... I mean organisations have to adapt to this new working scheme. First they need new processes (maybe borrowed from SW development) that enables them to triage work products on a risk/reward scale. For example my wife works on medical device tenders. It is obligatory to translate every frikkin Word document to our native language which in the end noone will read. Do we use LLMs to do the translation? Hell yeah. For a critical legal document? Eeee. Also I think enablers like speical harnesses shall be developed/improved by keeping these folks in mind. For example to build hooks into the harness that forces the LLM to test/review/sample its output. So yes it's a complex topic, but my point was rather that the inherent capabilities of medium-large-ish open LLMs are sufficient for let's say 70-80% of such office work, and it's a huge market.
- 2ndorderthought 5mo agoI think the conclusion is flawed here? Sure qwen3.5 9b is nowhere near the sota models. It's 9b and was made a year ago? Everyone taking about local models is pumped about the models released in April this year. Qwen 3.6 27b and qwen 35b a3b if you have a sad GPU. Those are comparable to sota models, seriously.
- amelius 5mo agoI'd pick a much more open system with more capabilities for a little bit more money, e.g. a Jetson Orin 64GB (unified memory). Runs Linux out of the box.
- redsocksfan45 5mo ago[dead]
- dizlexic 5mo agoThanks for sharing. I made a post earlier on bluesky describing my random setup on 32gb M2 studio. I'd love feedback. I'm a monkey and if I don't see I can't do. https://bsky.app/profile/mooresolutions.io/post/3mliilyf2i227 https://bsky.app/profile/mooresolutions.io/post/3mliilyf2i22...
- busfahrer 5mo agoI am considering a M5 Pro (18/20C) Macbook with 64GB of RAM, but I'm having a really hard time finding benchmarks of real world models: Could somebody please provide some tokens-per-second numbers for example for Qwen 3.6 35B/A3B, specifically for Q4 and Q6 quants?
- Galanwe 5mo agoMy advice: don't just look at tokens per second, but also at time to first token (TTFT). The local inference space is leaning to MoE models, and a lot of them have decent tokens / second, but horrible TTFT.
- Casteil 5mo agoYou can expect around 55-60t/s with Qwen3.5:35b-a3b or gemma4:26b-a4b Q4
- sourcecodeplz 5mo agoRunning LLMs local is fun and powerful but if you want to get work done... it is a big headache. You have to pre-plan and plan, and make specs, etc... The big OpenAI, Claude models just get you with just a few sentences..
- rvnx 5mo agoIt's actually technically easy now to run a large model at home for offline use (thanks to the Chinese who release their top-notch models). The main problem is finding the money :/
- Farmadupe 5mo agoYup, especially when for a lot of us, the price of the frontier subscription has become a cost of doing business over the last 6 months. If you're already doing big boy stuff with big boy models, then... just carry on trucking! Only place I'd differ is for vision/OCR tasks. Small/medium open weights models are as good as SoTa, and token prices for prefill are kinda very not worth it for larger batch tasks. Other thing that people forget is, if you want to have even a smallish LLM as a reliable personal service, you've got to carve out 16-24 of (V)RAM and leave it permanently running.
- ionwake 5mo ago"burn your thighs without getting anything out of it." what a phrase. love it.
- ionwake 5mo agoI have an M4 Macbook Air with 32Gb. These are my current results for my models: ┌──────────────────────┬───────────┬─────────────┐ │ Model │ Size │ Tokens/sec │ ├──────────────────────┼───────────┼─────────────┤ │ gemma-4-e4b-it-mlx │ ~4B (MLX) │ ~10.5 tok/s │ ├──────────────────────┼───────────┼─────────────┤ │ qwen3-8b-uncensor-v2 │ 8B │ ~6.3 tok/s │ ├──────────────────────┼───────────┼─────────────┤ │ qwen3-14b-uncensored │ 14B │ ~3.5 tok/s │ └──────────────────────┴───────────┴─────────────┘ I seem to be doing ok with the Gemma model for file parsing / handling.
- ActorNightly 5mo ago<=10 tok/sec is unusable. You are faster writing the code yourself.
- rs38 5mo agomy latest experiments with local LLM (mistral coder variations) fitting in older 6 GB GTX1060 were disappointing as long as you try to hook Copilot (CLI or VScode) to it and are used to provide a lot tooling. this seems to bloat initial prompt to 20k and more which seems the bottleneck if I did not completely misconfigured things. output tokens/s are more than fine, but PP is frustrating / unusable.
- compiler-devel 5mo agoI don’t understand the bipolar nature on hacker news towards LLMs. On the one hand, they’re destroying the art of software development and we shouldn’t use them. But on the other hand, there’s a lot of excitement around running them locally. I understand that multiple things can be true at the same time. Is the concern for centralized AI monopolization? Or is the concern for the art of software engineering?
- noashavit 5mo agoGemma4 is a huge improvement and it's fast. Qwen 3.5 really slows down my machine though. LMK if there is a better model to use for the code assist aspect- performance wise
- Casteil 5mo agoQwen3.5/3.6 are really prone to looping and 'overthinking'. Gemma4 doesn't seem to have the same problems.
- altruios 5mo agoGemma also doesn't have the same 'agentic' capabilities of qwen3.6. Simple test failed: sending "1","2","3" as separate messages using an openclaw harness. I tested a few other "follow these instructions" tests. Qwen3.5/6 were able to follow along, gemma was not able to.
- Xeoncross 5mo agoIs it better to have an M4-M5 Pro with 32GB of ram or an M1-M2 Max with 64GB of ram? They seem about the same price. It seems like cache layers like https://omlx.ai https://omlx.ai make more RAM better than more GPU cores or faster CPUs cores, but I'm curious if someone has tested both.
- threetonesun 5mo agoWhen I was considering a local setup the M1 Ultra Studios with 128 GB of RAM seemed to be the best price:performance at the time. I think RAM always wins out. Also minor note: the M4/5 Pros come in multiples of 12, so it's a 24/36 or 48GB set up.
- ChrisMarshallNY 5mo ago> The longer you let it drive without constraints, the worse the wreckage gets. The velocity makes you think you're winning right up until the moment everything collapses simultaneously. In my experience (so far), I can’t let the LLM write too much in one go. I need to test the hell out of what it gives me, and I can’t ask for too much, at one time. I tend to ask it to “flesh out” functions, where I have a signature, and a detailed headerdoc comment. I will provide a lot of guidance about the context, often attaching relevant files. Even then, it often doesn’t give me what I need, first time, unless it’s a small function, with extremely limited scope. That said, it’s been extremely helpful. It has accelerated my development greatly. I have found that it gives me much better PHP, than Swift. I suspect that may be because PHP is extremely mature, and there’s millions and millions of lines of high-quality code out there, in open-source repos, while Swift is probably mostly in closed repos, with open stuff not really provided by experienced developers (it’s a proprietary language used for shipping commercial software, so that may also apply to other languages). What it gives me in Swift, most closely resembles stuff that enthusiastic newer folks would do, and want to show off.
- hugmynutus 5mo ago> What it gives me in Swift, most closely resembles stuff that enthusiastic newer folks would do, and want to show off. The same is true for rust-lang. Code that will immediately clone/re-allocate anything passed by reference and collect everything to the heap that is passed by `Iterator`/`IntoIterator`. It is a massive performance anti-pattern and the hallmark of somebody "struggling" with the borrow checker. Naturally a lot of 1st & 2nd 'I just learned rust' projects lean on it. Which is totally fine for humans, you're learning. But with LLMs that pattern is now burned into their eigenvectors with the heat of a billion hours of H100 training time. It has gotten to a point that all code I generate with Opus or Codex if there as iterator or reference in the argument, I start a fresh context, with a sort of `remove unnecessary clones, collections, and copies from the following code: {{code}}`
- krferriter 5mo ago> It has gotten to a point that all code I generate with Opus or Codex if there as iterator or reference in the argument, I start a fresh context, with a sort of `remove unnecessary clones, collections, and copies from the following code: {{code}}` What does it do if you put "Avoid unnecessary clones, collections, and copies" in your CLAUDE.md/AGENTS.md?
- claysmithr 5mo agoThis is very good, local models run well on my M3 Air 24GB, to the point where I may prefer it even if it takes longer. The benefits are - private - local - no internet required - works well enough for most tasks - "free" - will pop this "AI" bubble as word spreads I got pretty good results with the model in the article on my machine. Sure, it took forever, but that doesn't matter to me as much, and it's kind of cool just watching it do its thing through LM studio. The result was also impressive enough for me that I would actually use it. Why pay $20/mo when local is good enough?
- adam_arthur 5mo agoI recently found Gemma 4 e4b surprisingly effective for small "classification" style tasks for something I'm doing at work. In this case, picking out "semantic" css classes on single dom nodes. Was able to run it on my 4(?) year old M2 mbp with 16GB of ram and it runs in only 100ms or so per query. Probably it can run much faster, but haven't experimented with batching etc With tight and targeted context control, you can use extremely small models for useful things. Ideally with problems where the harness can be mostly deterministic and you have known bounds on what you're trying to do
- yesman_x 5mo ago[flagged]
- deferredgrant 5mo ago[flagged]
- mrdependable 5mo agoI've been thinking about getting the M5 Max Macbook Pro when it comes out to run local models on, but I'm worried it would make the computer sluggish to work with. Is it better to just get some Strix Halo mini pc to be dedicated local machine?
- adamsb6 5mo agoI realized that should I end up getting laid off soon I won't have an unlimited token budget and for the workflows I've settled into it would be quite expensive. So I was exploring what it would take to run open models at home. Was quite disappointed to see that the PC side hasn't kept up. The unified architecture on Macs makes it very hard to justify spending money on a Linux machine for inference workloads.
- neverrroot 5mo agoA few interesting undertone points in the article, for those who care: - reliance on US technologies is not so good, but on Chinese is not discussed, just chosen - environmental cost is of concern - so are the energy costs In the end, there are some clear tips on how to configure the LLM, but overall the article is a bit thin and rather biased.
- abalashov 5mo agoI have an M4 Max MBP with w/128 GB of RAM, and have been very into local models. Qwen3.6-35B-A3B has been very kind to me for almost any purpose. No, it's no Opus 4.7, but it's shockingly good, and I don't have to share my code with Dario.
- lazylizard 5mo agoi thought it was a small typo of https://www.techpowerup.com/gpu-specs/tesla-m40-24-gb.c3838 https://www.techpowerup.com/gpu-specs/tesla-m40-24-gb.c3838 and wanted to ask what version of nvidia driver and cuda...
- SundarSharma 5mo ago[flagged]