11 ms·
Qwen3 30B A3B Hits 13 token/s on 4xRaspberry Pi 5
- geerlingguy 1y agodistributed-llama is great, I just wish it would work with more models. I've been happy with ease of setup and its ongoing maintenance compared to Exo, and performance vs llama.cpp RPC mode.
- alchemist1e9 1y agoAny pointers to what is SOTA for cluster of hosts with CUDA GPUs but not enough vram for full weights, yet 10Gbit low latency interconnects? If that problem gets solved, even if for only a batch approach that enables parallel batch inference resulting in high total token/s but low per session, and for bigger models, then it would he a serious game changer for large scale low cost AI automation without billions capex. My intuition says it should be possible, so perhaps someone has done it or started on it already.
- echelon 1y agoThis is really impressive. If we can get this down to a single Raspberry Pi, then we have crazy embedded toys and tools. Locally, at the edge, with no internet connection. Kids will be growing up with toys that talk to them and remember their stories. We're living in the sci-fi future. This was unthinkable ten years ago.
- taminka 1y ago[flagged]
- tonyhart7 1y agothis is very pessimistic take there are lot of bad people on internet too, does that make internet is a mistake ??? Noo, the people are not the tool
- Twirrim 1y agoIt's not unrealistically pessimistic. We're already seeing research showing the negative effects, as well as seeing routine psychosis stories. Think about the ways that LLMs interact. The constant barrage of positive responses "brilliant observation" etc. That's not a healthy input to your mental feedback loop. We all need responses that are grounded in reality, just like you'd get from other human beings. Think about how we've seen famous people, businesses leaders, politicians etc go off the rails when surrounded by "yes men" constantly enabling and supporting them. That's happening with people with fully mature brains, and that's literally the way LLMs behave. Now think about what that's going to do to developing brains that have even less ability to discern when they're being led astray, and are much more likely to take things at face value. LLMs are fundamentally dangerous in their current form.
- quesera 1y agoObsequiousness seems like the easiest of problems to solve. Although it's quite unclear to me what the ideal assistant-personality is, for the psychological health of children -- or for adults. Remember A Young Lady's Illustrated Primer from The Diamond Age. That's the dream (but it was fiction, and had a human behind it anyway). The reality seems assured to be disappointing, at best.
- tonyhart7 1y agoYes this is flaw on we train them, we must rethink on how rewards reinforced learning works but that doesn't mean its not fixable, that doesn't mean progress must stop if the earliest inventor of plane think like you, human would never conquer skies we are in explosive growth that many brightest mind in planet get recruited to solve this problem, in fact I would be baffled if we didn't solve this by the end of year if humankind cant fix this problem, just say goodbye at those sci-fi interplanetary tech
- Twirrim 1y agoWow. That's... one hell of a leap you're making.
- 1y ago
- striking 1y agoI think it's worth remembering that there's room for thoughtful design in the way kids play. Are LLMs a useful tool for encouraging children to develop their imaginations or their visual or spatial reasoning skills? Or would these tools shape their thinking patterns to exactly mirror those encoded into the LLM? I think there's something beautiful and important about the fact that parents shape their kids, leaving with them some of the best (and worst) aspects of themselves. Likewise with their interactions with other people. The tech is cool. But I think we should aim to be thoughtful about how we use it.
- manmal 1y agoAn LLM in my kids‘ toys only over my cold, dead body. This can and will go very very wrong.
- supportengineer 1y ago[flagged]
- ugh123 1y agoWhat about a kid who lives in an urban area without parks?
- hkt 1y agoCampaign for parks
- supportengineer 1y agoMy snarky answer would be that a lack of parks presents a clear and present danger to the kid.
- bongodongobob 1y agoYou can do both bro.
- Aurornis 1y agoParent here. Kids have a lot of time and do a lot of different things. Some times it rains or snows or we’re home sick. Kids can (and will) do a lot of different things and it’s good to have options.
- supportengineer 1y agoKids can play in the rain and in the snow.
- bigyabai 1y ago> Kids will be growing up with toys that talk to them and remember their stories. What a radical departure from the social norms of childhood. Next you'll tell me that they've got an AI toy that can change their diaper and cook Chef Boyardee.
- fragmede 1y agoIf a raspberry pi can do all that, imagine the toys Bill Gates' grandkids have access to! We're at the precipice of having a real "A Young Lady's Illustrated Primer" from The Diamond Age.
- 9991 1y agoBill Gates' grandkids will be playing with wooden blocks.
- 1gn15 1y agoThis is indeed incredibly sci fi. I still remember my ChatGPT moment, when I realized I could actually talk to a computer. And now it can run fully on an RPi, just as if the RPi itself has become intelligent and articulate! Very cool.
- dingdingdang 1y agoVery impressive numbers.. wonder how this would scale on 4 relatively modern desktop PCs, like say something akin to a i5 8th Gen Lenovo ThinkCentre, these can be had for very cheap. But like @geerlingguy indicates - we need model compatibility to go up up up! As an example it would amazing to see something like fastsdcpu run distributed to democratize accessibility-to/practicality-of image gen models for people with limited budgets but large PC fleets ;)
- rthnbgrredf 1y agoI think it is all well and good, but the most affordable option is probably still to buy a used MacBook with 16/32 or 64 GB (depending on the budget) unified memory and install Asahi Linux for tinkering. Graphics cards with decent amount of memory are still massively overpriced (even used), big, noisy and draw a lot of energy.
- ivape 1y agoIt just came to my attention that the 2021 M1 Max 64gb is less than $1500 used. That’s 64gb of unified memory at regular laptop prices, so I think people will be well equipped with AI laptops rather soon. Apple really is #2 and probably could be #1 in AI consumer hardware.
- jeroenhd 1y agoApple is leagues ahead of Microsoft with the whole AI PC thing and so far it has yet to mean anything. I don't think consumers care at all about running AI, let alone running AI locally. I'd try the whole AI thing on my work Macbook but Apple's built-in AI stuff isn't available in my language, so perhaps that's also why I haven't heard anybody mention it.
- ivape 1y agoPeople don’t know what they want yet, you have to show it to them. Getting the hardware out is part of it, but you are right, we’re missing the killer apps at the moment. The very need for privacy with AI will make personal hardware important no matter what.
- mehdibl 1y ago[flagged]
- hidelooktropic 1y ago13/s is not slow. Q4 is not bad. The models that run on phones are never 30B or anywhere close to that.
- lostmsu 1y agoIt is very slow and totally unimpressive. 5060Ti ($430 new) would do over 60, even more in batched mode. 4x RPi 5 are $550 new.
- magicalhippo 1y agoSo clearly we need to get this guy hooked up with Jeff Geerling so we can have 4x RPi5s with a 5060 Ti each... Yes, I'm joking.
- deleted 1y ago[deleted]
- kosolam 1y agoHow is this technically done? How does it split the query and aggregates the results?
- magicalhippo 1y agoFrom the readme: More devices mean faster performance, leveraging tensor parallelism and high-speed synchronization over Ethernet. The maximum number of nodes is equal to the number of KV heads in the model #70. I found this[1] article nice for an overview of the parallelism modes. [1]: https://medium.com/@chenhao511132/parallelism-in-llm-inference-c0b6bdc5f693 https://medium.com/@chenhao511132/parallelism-in-llm-inferen...
- varispeed 1y agoSo would 40x RPi 5 get 130 token/s?
- SillyUsername 1y agoI imagine it might be limited by number of layers and you'll get diminishing returns as well at some point caused by network latency.
- VHRanger 1y agoMost likely not because of NUMA bottlenecks
- reilly3000 1y agoIt has to be 2^n nodes and limited to one per attention head that the model has.
- behnamoh 1y agoEverything runs on a π if you quantize it enough! I'm curious about the applications though. Do people randomly buy 4xRPi5s that they can now dedicate to running LLMs?
- blululu 1y agoFor $500 you may as well spend an extra $100 and get a Mac mini with an m4 chip and 256gb of ram and avoid the headaches of coordinating 4 machines.
- MangoToupe 1y agoI don't think you can get 256 gigs of ram in a mac mini for $600. I do endorse the mac as an AI workbench tho
- ryukoposting 1y agoI'd love to hook my development tools into a fully-local LLM. The question is context window and cost. If the context window isn't big enough, it won't be helpful for me. I'm not gonna drop $500 on RPis unless I know it'll be worth the money. I could try getting my employer to pay for it, but I'll probably have a much easier time convincing them to pay for Claude or whatever.
- mmastrac 1y agoIs the network the bottleneck here at all? That's impressive for a gigabit switch.
- tarruda 1y agoI suspect you'd get similar numbers with a modern x86 mini PC that has 32GB of RAM.
- misternintendo 1y agoAt this speed this is only suitable for time insensitive applications..
- shaaca 1y ago[dead]
- YJfcboaDaJRDw 1y ago[dead]
- poly2it 1y agoNeat, but at this price scaling it's probably better to buy GPUs.
- rao-v 1y agoNice! Cheap RK3588 boards come with 15GB of LPDDR5 RAM these days and have significantly better performance than the Pi 5 (and often are cheaper). I get 8.2 tokens per second on a random orange pi board with Qwen3-Coder-30B-A3B at Q3_K_XL (~12.9GB). I need to try two of them in parallel ... should be significantly faster than this even at Q6.
- jerrysievert 1y ago> a random orange pi board with Qwen3-Coder-30B-A3B at Q3_K_XL (~12.9GB) fantastic! what are you using to run it, llama.cpp? I have a few extra opi5's sitting around that would love some extra usage
- rao-v 1y agoYup! Build and ignore KleinAI and Vulkan etc. I’ve found that a clean CPU only build is optimal
- ThatPlayer 1y agoIs that using the NPU on that board? I know it's possible to use those too.
- rao-v 1y agoIt is possibly (superb subreddit) but painful to convert a modern model and takes ages for them to be supported. The NPU is energy efficient but no faster than CPU for generation (and has lousy software support). I’m mostly interested in the NPu to run a vision head in parallel with an LLM to speed up time to first token with VLLMs (kinda want to turn them into privacy safe vision devices for consumer use cases)
- ThatPlayer 1y agoSince my comment, I remembered I had a RK3588 board, a Rock 5B, and tried llama.cpp CPU over that, and performance was not amazing. But also I realized this is LPDDR4X, so don't get the cheapest RK3558 boards. My Orange Pi 5 is actually worse. This one has LPDDR4. Looking at the rest of Orange Pi's line-up, they don't actually have a board with both LPDDR5 and 32GB, only 16GB or LPDDR4(X). Using llama-bench, and Llama 2 7B Q4_0 like https://github.com/ggml-org/llama.cpp/discussions/10879 https://github.com/ggml-org/llama.cpp/discussions/10879 how does yours compare? Cuz I'm also comparing it with a few a few Ryzen 5 3000 Series mini-pcs for less than 150$, and that gets 8 t/s on this list and I've gotten myself With my Rock 5B and this bench, I get 3.65 t/s. On my Orange Pi 5 (not B) 8GB LPDDR4 (not X), I get 2.44 t/s.
- ineedasername 1y agoThis is highly usable in an enterprise setting when the task benefits from near-human level decision making and when $acceptable_latency < 1s meets decisions that can be expressed in natural language <= 13tk. Meaning that if you can structure a range of situations and tasks clearly in natural language with a pseudo-code type of structure and fit it in model context then you can have an LLM perform a huge amount of work with Human-in-the-loop oversight & quality control for edge cases. Think of office jobs, white colar work, where, business process documentation and employee guides and job aids already fully describe 40% to 80% of the work. These are the tasks most easily structured with scaffolding prompts and more specialized RLHF enriched data, and then perform those tasks more consistently. This is what I decribe when I'm asked "But how will they do $X when they can't answer $Y without hallucinating?" I explain the above capability, then I ask the person to do a brief thought experiment: How often have you heard, or yourself thought something like, "That is mindnumbingly tedious" and/or "a trained monkey could do it"? In the end, I don't know anyone whose is aware of the core capabilities in the structured natural-language sense above, that doesn't see at a glance just how many jobs can easily go away. I'm not smart enough to see where all the new jobs will be or certain there will be as many of them, if I did I'd start or invest in such businesses. But maybe not many new jobs get created, but then so what? If the net productivity and output-- essentially the wealth-- of the global workforce remains the same or better with AI assistance and therefore fewer work hours, that means... What? Less work on average, per capita. More wealth, per work hour worked per Capita than before. Work hours used to be longer, they can shorten again. The problem is getting there. To overcoming not just the "sure but it will only be the CEOs get wealthy" side of things to also the "full time means 40 hours a week minimum." attitude by more than just managers and CEOs. It will also mean that our concept of the "proper wage" for unskilled labor that can't be automated will have to change too. Wait staff at restaurants, retail workers, countless low end service-workers in food and hospitality? They'll now be providing-- and giving up-- something much more valuable than white colar skills that are outdated. They'll be giving their time to what I've heard, and the term is jarring to my ears but it is what it is, I've heard it described as "embodied work". And I guess the term fits. And anyway I've long considered my time to be something I'll trade with a great deal more reluctance than my money, and so demand a lot money for it when it's required so I can use that money to buy more time (by not having to work) somewhere in the near future, even if it's just by covering my costs for getting groceries delivered instead of the time to go shopping myself. Wow, this comment got away from me. But seeing Qwen3 30B level quality with 13tk/s on dirt cheap HW struck a deep chord of "heck, the global workforce could be rocked to the core for cheap+quality 13tk/s." And that alone isn't the sort of comment you can leave as a standalone drive-by on HN and have it be worth the seconds to write it. And I'm probably wrong on a little or a lot of this and seeing some ideas on how I'm wrong will be fun and interesting.
- drbscl 1y agoDistributed compute is cool, but $320 for 13 tokens/s on a tiny input prompt, 4 bit quantization, and 3B active parameter model is very underwhelming
- ab_testing 1y agoWould it work better on a used GPU?
- bjt12345 1y agoDoes Distributed Llama use RDMA over Converged Ethernet or is this roadmapped? I've always wondered if RoCE and Ultra-Ethernet will trickle down into the consumer market.
- rldjbpin 1y agohow would llm-d [1] work compared to distributed-llama? is the overhead or configuration too much to work with for simple setups? [1] https://github.com/llm-d/llm-d/ https://github.com/llm-d/llm-d/