7 ms·
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
- brrrrrm 2mo agothis is cool but like, are we just vibe coding NAND burners at this point? these decode times don't really tell the whole story, because prefill becomes the bottleneck. half an hour to process 10k tokens on an M5 seems... not great
- kennywinker 2mo agoNot great for coding, or realtime agent interactions. But for background processing tasks overnight? Seems like it’d work pretty well
- selcuka 2mo agoI'm pretty sure one can rent a GPU for a few minutes with the electricity cost of leaving an M5 overnight.
- hdgvhicv 2mo agoDomestic electricity is free nowadays, certainly for most of the year, as solar plus battery covers your usage for a tiny percentage of the cost of your house.
- wccrawford 2mo agoOnly if you don't count the cost of the equipment and installation.
- hdgvhicv 2mo agoOr the cost of the house. Given the cost of a building is far more than the cost of generating enough power for that building it doesn’t really matter
- kennywinker 2mo agoSure. One could. But then one wouldn’t be in control of every step of the process.
- fsuts 2mo agoThis is how progress happens, someone gets to 3t/s, the next person gets to6/s and eventually we get to 100t/s. People like this person are laying the foundations.
- IsTom 2mo agoThere is a limit how much you can squeeze out of given hardware. Betting on it being closer to 100t/s than 6t/s is only that, a bet.
- rhdunn 2mo agoOn my 4090 setup I'm getting 86t/s on a 12B Q6_K quantized model running entirely in VRAM. The current GPUs are optimized for processing huge numbers of triangles per second. There are three things at play here: 1. the organization of the data being sent to the GPU to optimize throughput; 2. the speed at which the GPU can read that data from its VRAM; 3. how many triangles it can process in parallel by using individual compute units. I suspect that given parallel improvements for neural networks, we'll see similar improvements: 1. optimizing the structure of the weights in the model for efficient access by the CPU/GPU/NPU/TPU; 2. efficient access of data strides (matrix rows) in the memory, e.g. being able to read multiple 2x2 matrix values in one clock cycle, or stepwise pairs of values (a(i,j), b(j,k)) needed for matrix multiplication; 3. parallel compute for matrix and tensor multiplication and other operations needed by neural networks.
- IsTom 2mo agoIf you've got X GB of weights in slow access memory (be it RAM vs VRAM or SSD vs RAM) and Y GB of fast memory then no matter what, if you want to use them you'll need to transfer X-Y GB and will be bound by memory throughput. You can try to reduce number of activated weights, but how much can be gained that way is speculative so far.
- leonickson 2mo agoagree, prefill is the weak spot right now. it goes through the same per-token path as decode, which is dumb for long prompts. The fix is on the list: during prefill we can batch the expert reads for the whole prompt per layer instead of per token, that amortizes the IO a lot. until that lands, honest answer is this is good for chat-length stuff, not for feeding it a 10k token document.
- piyh 2mo agoSSD NAND reads are nearly infinite. Still makes me uncomfortable, but writing is what kills. There's a reason SSDs are rated by TBW, not TBR.
- jbird99 2mo agoAt what, 10 tokens per hour? These disk swapping methods all have the same drawbacks - kill your drive early, and slow as hell.
- kennywinker 2mo agoIt says very prominently in the post: 4.5-5t/s for 80b on an M5
- wat10000 2mo agoIsn’t it only writes that kill drives?
- Alpha3031 2mo agoYes for NAND, and I suppose nobody is using mechanical hard drives for this.
- wat10000 2mo agoI'd like to see someone try it, just to see how incredibly slow and noisy it is.
- sudo_cowsay 2mo agoYeah, that's why most of these comments seem weird to me.
- petu 2mo agoThere's read disturb on SSDs, enough reads will eventually force controller to rewrite the cell and it's neighbours. Practically if you're not streaming weights 24/7 from a full SSD, then it shouldn't be a problem.
- zozbot234 2mo agoRead disturb ought to be quite rare, especially on a fresh drive that was written only once or a handful of times (WORM-like usage). Practically, it's not likely to be an issue even with very heavy read workloads.
- dghlsakjg 2mo agoI know everyone wants to crap all over these setups that are impractical, but this is how progress happens. People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc. Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.
- arjie 2mo agoHaha 1T on $50k might be a bit hopeful, mate, even at FP8. But I too am hopeful.
- hedora 2mo agoAMD already demonstrated 1T on strix halo clusters. << $10K at original MSRP.
- arjie 2mo agoProblem is prefill on these, right? Initial prompt processing takes forever? I suppose you’re right. Cost is not a thing on its own. It’s a performance-cost frontier and one can do CPU inference in the worst case.
- apimade 2mo ago8800 GTX in 2006. Cutting-edge, an insanely powered consumer card for the time. Theoretically around 0.3456 TFLOPS. 1080 GTX in 2016. Cutting-edge, an insanely powerful consumer card for the time. Theoretically around 8.87 to 8.9 TFLOPS. 5090 RTX in 2026. Cutting-edge, an insanely powerful consumer card for today. Theoretically around 104.8 TFLOPS. In the same timeframe mobile processor CPU's went from 0.001 TFLOPS, to today's Apple's A19 Pro chip which delivers 2.074 TFLOPS. That's _without_ getting into ASIC's, or purpose-built hardware like Taalas's model on silicon HC1, or generic AI dies like what they're planning with HC2 or Cerebras, which will massively compress the timeline.
- cududa 2mo ago
- AHASIC 2mo agoI read a comment on here a few months back I wanna restate. Basically, there is a good chance that Apple is betting that the LLMs in the future will be so efficient that those that consumers will use everyday will be easily computed by the iPhone or even bigger ones on Macs. Honestly makes the most sense that we are heading that way in a few years latest.
- Mistletoe 2mo agoWhat hardware advances would we need to see for that to happen? It feels like everything in that arena has kind of plateaued.
- sudo_cowsay 2mo agoIt could be on software side too. OpenAI has certainly not plateaued.
- bobbylarrybobby 2mo agoThe models themselves have far from plateaued. Maybe someone finds a way to get a really capable model down to, say, 12GB of ram. Then we'd be in business.
- swiftcoder 2mo agoAgreed. We've just seen DeepSeek post-train their ~300 billion parameter flash model to outperform their 1.6 trillion parameter pro model, in the space of a few months. There would seem to still be quite a few opportunities on the table to bring big model smarts down to the smaller models
- deleted 2mo ago[deleted]
- CircuitSeuss 2mo agoA lot of this will come from co-optimizing hardware and low level machine code for this specific use case… something apple is coincidently very good at. Apple has worked very hard to make unified memory a feasible approach, and the benefits of that are pretty clear in apple silicon- that efficiency not only results in power and therefore thermal gains, but also in a significantly faster full loop per process: or a faster time to token. This is why even their single core mobile chips in the budget line Neo out perform PC processors with several times more threads and RAM[1]. Turns out, unified memory lets you have a whole lot more control over things like RAM bussing and core use for specific workflows. Speculatively, a unified memory approach could also allow you to more easily integrate things like ReRAM to solve the current memory swapping bottleneck. Let’s say a friend of mine works hardware at apple and works on exactly this… on device processing is the future I’m betting on. [1] https://youtu.be/x26A28DoT-w?t=605 https://youtu.be/x26A28DoT-w?t=605
- CyLith 2mo agoI know relatively little about the workings of LLMs, but I keep seeing projects like this that run massive MoE models using very modest amounts of RAM, perhaps excessively so. I wonder, is there a way to make the RAM usage tunable? I have a Macbook with 32 GB of RAM, and it'd be great if I could run the same model but take advantage of the additional RAM to make it run faster.
- ianmurrays 2mo agoI guess you have to know which experts to keep “hot” in ram, which you can’t know beforehand, so there wouldn’t be much gain.
- spockz 2mo agoI do wonder if there are some experts that are more likely to be hit. So if the normal optimised setup runs in 12GiB an you have 4GiB extra to spare, you could say “promote the most used X experts to this stable (old gen in GC parlance) region and don’t swap it out. Maybe you could even do something like profiling and remember over multiple sessions (per project/workspace) what the most used agents are and load those up before hand.
- zamadatix 2mo agoThat's about the turning point for just using typical quants for me. Larger still and you can just do the full model. Smaller to this degree and you need all sorts of extra tricks to get anything.
- fodkodrasz 2mo ago> I wonder, is there a way to make the RAM usage tunable? In LM Studio I can tune it by selecting different quantation of the model, by selecting how many layers of the neural net to be loaded to GPU (rest stays in main mem, evaluated by the CPU), and by adjusting context window.
- leonickson 2mo agoIt's tunable, --cache-gb N on the CLI. In my sweep the speed barely moved between a 1GB and 6GB cache (43% vs 70% hit rate, same tok/s) because right now the bottleneck is GPU dispatch, not the SSD. so more RAM doesnt buy much yet. once the kernel work lands it should start to matter, so on 32GB I would just set 8 and let it age well. Also the hit rates themselves answer the "can you even know which experts stay hot" question, reuse across tokens is very real.
- adrianco 2mo agoThis looks useful, you can increase the RAM cache so if you have a Mac with 24-32GB it should speed up a lot and still run models that wouldn’t normally fit. I’m going to run some tests…
- crossroadsguy 2mo agoI see this at the end of the README > Swiftlet was built in collaboration with Claude Code. Did this really happen (some sort of working with Anthropic or Claude Code team) or is it some kind of requirement when you develop some software with Claude Code (I see the other author is: https://github.com/claude https://github.com/claude), or sort of reuse some of its parts? Is it like someone saying "built in collaboration with VS Code" or ".. in collaboration with <xyz> autocomplete plugin"? Or merely a disclaimer about vibe-coding or AI written tool?
- iamflimflam1 2mo agoIt means they used Claude code to write the software.
- vermarish 2mo agoWhen you have Claude Code indepedently author commits and PRs and merge them in, it'll always credit itself as an author. I assume it showing up in the README is a byproduct of the same logic.
- leonickson 2mo agono Anthropic involvement, I just used Claude Code heavily while building this and putting that in the README felt more honest than not mentioning it. Now that I think about it may be it shuld be "built with claude code" instead of "built in collaboration...". Changed it.
- crossroadsguy 2mo agoThank you for explaining. > Swiftlet was built with Claude Code. This is great and correct. It's a tool. Again, thanks. PS. I just don't why people downvoted me. I am just someone, after a longish sabbatical/gap, exploring and getting used to the agnatic world, though very slowly :)
- gitpusher42 2mo agoThank you for using TurboFieldfare as a starting point for this project and thank you for mentioning it at the README. I am glad it inspired more people to explore area of on-device AI further!
- leonickson 2mo ago[dead]
- sallymander 2mo ago"As far as we know, that is the first time a model of this class has run natively on a phone.". I feel like I've seen a similar statement on a lot of these streaming weight projects. 400b model on an iPhone: https://x.com/anemll/status/2035901335984611412 https://x.com/anemll/status/2035901335984611412
- leonickson 2mo ago[dead]
- throwawayffffas 2mo ago> One expectation to set honestly Hello Claude!
- sbsbdbdbfndj 2mo agoLet me just make sure first instead of guessing
- tredre3 2mo agoThat line was authored by Claude. There you go: https://github.com/leonickson1/Swiftlet/commit/3ac64020eadcb9629abcbc162bab2f71f44af3f7 https://github.com/leonickson1/Swiftlet/commit/3ac64020eadcb...
- vancekai 2mo ago[dead]
- gizmodo59 2mo agoThe web and connecting to other services is very important for almost all of my use cases. While I believe we are going to get better and faster models, the web index is certainly not downloadable and maintainable for 99.99% of the folks who are able to use local models. Any good solutions exist?
- pbronez 2mo agoThere are many search APIs available, I like Kagi's. Microsoft and Amazon both provide web snapshot services that purport to give you a kind of agent-first internet archive. You can approximate something like that using common crawl, but it's a huge amount of data. Downloading the internet is impossible or a bad idea for almost everyone.
- morgoo 2mo agoIf you're really into self-hosting I've been experimenting with SearXNG and early signs are promising
- myshapeprotocol 2mo ago[flagged]
- hn974izqdv 2mo ago[dead]
- nc55g3g 2mo agoRunning a 35B on iPhone at 1 tok/s with 2.5 GB RAM… this is the future of on-device inference. Insane work.
- lern_too_spel 2mo agoSee also: https://github.com/Helldez/BigMoeOnEdge https://github.com/Helldez/BigMoeOnEdge
- lenerdenator 2mo agoHow can I do this with, say, Gemma?
- gitpusher42 2mo agoYeah, original TurboFieldfare supports Gemma https://github.com/drumih/turbo-fieldfare https://github.com/drumih/turbo-fieldfare
- erelong 2mo agoIs there something that already runs like this on android / linux ( / windows)? (ollama or something?) Or could this be ported to work on other such platforms? edit: AI mentions a "BigMoeonEdge" project