9 ms·
Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5
- 8thcross 2y agowhats the point of this? serious question - can someone provide usecases for this?
- cwoolfe 2y agoIt is shown running on 2 or 4 raspberry pis; the point is that you can add more (ordinary, non GPU) hardware for faster inference. It's a distributed system. The sky is the limit.
- 8thcross 2y agoah, Thanks! but what can a distributed system like this do? is this a fun to do, for the sake of doing it project or does it have practical applications? just curious about applicability thats all.
- __MatrixMan__ 2y agoI'm going to get downvoted for saying the B-word, but I imagine this growing up into some kind of blockchain thing where the AI has some goal and once there's consensus that some bit of data would further that goal it goes in a block on the chain (which is then referenced by humans who also have that goal and also is used to fine tune the AI for the next round of inference). Events in the real world are slow enough that gradually converging on the next move over the course of a few days would probably be fine. The advantage over centralizing the compute is that you can just connect your node and start contributing to the cause (both by providing compute and by being its eyes and hands out there in the real world), there's no confusion over things like who is paying the cloud compute bill and nobody has invested overmuch in hardware.
- 8thcross 2y agogot it. plan B for skynet. one baby transformer at a time.
- deleted 2y ago[deleted]
- __MatrixMan__ 2y agoEh, something that's trapped in a blockchain and can only move forward when people vote to approve its next thought/block is a lot less scary to me than some centralized AI running behind closed doors and taking direction from some rich asshole who answers only to himself. I think of it more like government 2.0.
- semi-extrinsic 2y agoIt doesn't even scale linearly to 4 nodes. It's slower than a five year old gaming computer. There is definitely a hard limit on performance to be had from this approach.
- gatienboquet 2y agoIn a distributed system, the overall performance and scalability are often constrained by the slowest component. This Distributed Llama is over Ethernet..
- walrus01 2y agoThe raspberry pis aren't really the point, since the raspberry pi os is basically debian, this means you could do the same thing on four much more powerful but still very cheap ($250-300 a piece) x86-64 systems running debian (with 32, 64 or 128GB RAM each if you needed). Also opening up the possibility of relatively cheap pci-express 3.0 based 10 Gbps NICs and switch between them, which isn't possible with raspberry pi.
- memhole 2y agoThis is the modern Beowulf cluster.
- mjhagen 2y agoBut can it run Crysis?
- samstave 2y agoIt can run doom in a .ini
- semi-extrinsic 2y agoI honestly don't understand the meme with RPi clusters. For a little more money than 4 RPi 5's, you can find on eBay a 1U Dell server with a 32 core Epyc CPU and 64 GB memory. This gives you at least an order of magnitude more performance. If people want to talk about Beowulf clusters in their homelab, they should at least be running compute nodes with a shoestring budget FDR Infiniband network, running Slurm+Lustre or k8s+OpenStack+Ceph or some other goodness. Spare me this doesnt-even-scale-linearly-to-four-slowass-nodes BS.
- madduci 2y agoThe TDP of 4 PIs combined is still smaller than a larger server, which is probably the whole point of such an experiment?
- znpy 2y agoThe combined TDP of 4 raspberry PIs is likely less than what the fans of that kind of server pull from the power outlet.
- znpy 2y agoYou can buy it but you can't run it, unless you're fairly wealthy. In my country (italy) a basic colocation service is like 80 euros/month + vat, and that only includes 100Wh of power and a 100mbps connection. +100wh/month upgrades are like +100 euros. I looked up the kind of servers and cpus you're talking about and the cpu alone can pull something like 180W/h, without accounting for fans, disks and other stuff (stuff like GPUs, which are power hungry). Yeah you could run it at home in theory, but you'll end up paying power at consumer price rather than datacenter pricing (and if you live in a flat, that's going to be a problem). Unless you're really wealthy, you have your own home with sufficient power[1] delivery and cooling. [1] not sure where you live, but here most residential power connections are below 3 KWh. If otherwise you can point me at some datacenter that will let me run a normal server like the ones you're pointing at for like 100-150 euros/month, please DO let me know and i'll rush there first thing next business day and I will be throwing money at them.
- NitpickLawyer 2y agoAs always, take those t/s stats with a huge boulder of salt. The demo shows a question "solved" in < 500 tokens. Still amazing that it's possible, but you'll get nowhere near those speeds when dealing with real-world problems at real-world useful context lengths for "thinking" models (8-16k tokens). Even epyc's with lots of channels go down to 2-4 t/s after ~4096 context length.
- numba888 2y agoSmaller robots tend to have smaller problems. Even little help from the model will make them a lot more capable than they are today.
- b4rtazz 2y agoI checked how it performs in long run (prediction) on 4 x Raspberry Pi 5: * pos=0 => P 138 ms S 864 kB R 1191 kB Connect * pos=2000 => P 215 ms S 864 kB R 1191 kB . * pos=4000 => P 256 ms S 864 kB R 1191 kB manager * pos=6000 => P 335 ms S 864 kB R 1191 kB the
- tofof 2y agoThis continues the pattern of all other announcements of running 'Deepseek R1' on raspberry pi - that they are running llama (or qwen), modified by deepseek's distillation technique.
- corysama 2y agoYeah. People looking for “Smaller DeepSeek” are looking for the quantized models, which are still quite large. https://unsloth.ai/blog/deepseekr1-dynamic https://unsloth.ai/blog/deepseekr1-dynamic
- zozbot234 2y agoYes this is just a fine-tuned LLaMa with DeepSeek-like "chain of thought" generation. A properly 'distilled' model is supposed to be trained from scratch to completely mimick the larger model it's being derived from - which is not what's going on here.
- kgeist 2y agoI tried the smaller 'Deepseek' models, and to be honest, in my tests, the quality wasn't much different from simply adding a CoT prompt to a vanilla model.
- rcarmo 2y agoYet for some things they work exactly the same way, and with the same issues :)
- whereismyacc 2y agoI really don't like that these models can be branded as Deepseek R1.
- sgt101 2y agoWell, Deepseek trained them?
- 2y ago
- behnamoh 2y agoOkay but does any one actually _want_ a reasoning model at such low tok/sec speeds?!
- ripped_britches 2y agoLots of use cases don’t require low latency. Background work for agents. CI jobs. Other stuff I haven’t thought of.
- behnamoh 2y agoIf my "automated" CI job takes more than 5 minutes, I'll do it myself..
- bee_rider 2y agoI bet the raspberry pi takes a smaller salary though.
- chickenzzzzu 2y agoThere are tasks that I don't want to do whose delivery sensitivity is 24 hours, aka they can be run while I'm sleeping.
- baq 2y agoWhere I’ve been doing CI 5 minutes was barely enough to warm caches on a cold runner
- Xeoncross 2y agoNo, but the alternative in some places is no reasoning model. Just like people don't want old cell phones / new phones with old chips - but often that's all that is affordable in some places. If we can get something working, then improving it will come.
- rvnx 2y agoYou can have questions that are not urgent. It's like Cursor, I'm fine with the slow version until a certain point, I launch the request then I alt-tab to something else. Yes it's slower, but well, for free (or cheap) it is acceptable.
- c6o 2y agoThat’s the real future
- rahimnathwani 2y agoThe interesting thing here is being able to run llama inference in a distributed fashion across multiple computers.
- samstave 2y agoWhich begs the question; Where is the equiv of distributed GPU? (Seti@HOME) and just pipe to a tool which is globally ditributed R1 full, but slow... and let it reason in the open on deep and complex tasks
- zdw 2y agoDoes adding memory help? There's a Rpi 5 with 16GB RAM recently available.
- zamadatix 2y agoMemory capacity in itself doesn't help so long as the model+context fits in memory (and and 8B parameter Q4 model should fit in a single 8 GB Pi).
- cratermoon 2y agoIs there a back-of-the-napkin way to calculate how much memory a given model will take? Or what parameter/quantization model will fit in a given memory size?
- monocasa 2y agoq4=4bits per weight So Q4 8B would be ~4GB.
- zamadatix 2y agoTo find the absolute minimum you just multiply the number of parameters by the bits per parameter, divide by 8 if you want bytes. In case 8 billion parameters of 4 bits each means "at least 4 billion bytes". For back of the napkin add ~20% overhead to that (it really depends on your context setup and a few other things but that's a good swag to start with) and then add whatever memory the base operating system is going to be using in the background. Extra tidbits to keep in mind: - A bits-per-parameter higher than the model was trained adds nothing (other than compatibility on certain accelerators) but a bits-per-parameter lower than the model was trained degrades the quality. - Different models may be trained at different bits-per-parameter. E.g. 671 billion parameter Deepseek R1 (full) was trained at fp8 while llama 3.1 405 billion parameter was trained and released at a higher parameter width so "full quality" benchmark results for Deepseek R1 require less memory than Llama 3.1 even though R1 has more total parameters. - Lower quantinizations will tend to run proportionally faster if you were memory bandwidth bound and that can be a reason to lower the quality even if you can fit the larger version of a model into memory (such as in this demonstration).
- blackeyeblitzar 2y agoCan’t you run larger models easily on MacBook Pro laptops with the bigger memory options? I think I read that people are getting 100 tokens a second on 70B models.
- efficax 2y agohaven’t measured but the 70b runs fine on an m1 macbook pro with 64gb of ram, although you can’t do much until it has finished
- politelemon 2y agoYou can get even faster results with GPUs, but that isn't the purpose of this demo. It's showcasing the ability to run such models on commodity hardware, and hopefully with better performance in the future.
- sgt101 2y agoI find 100tps unlikely, I see 13tps on an 80GB A100 for a 70B 4bit quantized model. Can you link - I am answering an email on monday where this info would be very useful!
- JKCalhoun 2y agoPeople on Apple Silicon are running special models for MLX. Maybe start with the MLX Community page: https://huggingface.co/mlx-community https://huggingface.co/mlx-community
- amelius 2y agoWhen can I "apt-get install" all this fancy new AI stuff?
- deleted 2y ago[deleted]
- Towaway69 2y agoOn my mac: brew install ollama might be a good start ;)
- dzikimarian 2y ago"ollama pull" is pretty close
- dheera 2y agoWhy tf isn't ollama in apt-get yet? F these curl|sh installs.
- diggan 2y ago> Why tf isn't ollama in apt-get yet? There are three broad groups of people in packaging: A) People who package stuff you and others need. B) People who don’t package stuff but use what’s available. C) People who don’t package stuff and complain about what’s available without taking further action. If you find yourself in group C while having no interest in contributing to group A or working within the limits of group B, then you have two realistic options: either open your wallet and pay someone to package it for you or accept that your complaints won’t change anything. Most packaging work is done by volunteers. If something isn’t available, it’s not necessarily because no one sees value in it—it could also be due to policy restrictions, dependency complexity, or simply a lack of awareness. If you want it packaged, the best approach is to contribute, fund the work, or advocate for it constructively.
- amelius 2y agoThis is not entirely fair. We can't all be involved in package management, and GP might contribute in other ways.
- replete 2y agoThat's not a bad result, although for £320 for 4x Pi5s you could probably find a used 12GB 3080 and probably more than 10x token speed
- varispeed 2y ago> Deepseek R1 Distill 8B Q40 on 1x 3080, 60.43 tok/s (eval 110.68 tok/s) That wouldn't get on Hacker News ;-)
- jckahn 2y agoHNDD: Hacker News Driven Development
- geerlingguy 2y agoOr attach a 12 or 16 GB GPU to a single Pi 5 directly, and get 20+ tokens/s on an even larger model :D https://github.com/geerlingguy/ollama-benchmark?tab=readme-ov-file#deepseek https://github.com/geerlingguy/ollama-benchmark?tab=readme-o...
- littlestymaar 2y agoReading the beginning of your comment I was like “ah yes I saw Jeff Geerling do that on a video”. Then I saw you github link and your HN handle and I was like “Wait, it is Jeff Geerling!”. :D
- ziml77 2y agoHaha I had nearly the same thing happen. First I was like "that sounds like something Jeff Geerling would do". Then I saw the github link and was like "ah yeah Jeff Geerling did do it" and then I saw the username and was like "oh it's Jeff Geerling!"
- replete 2y agoThanks for sharing. Pi5 + cheap AMD GPU = convenient modest LLM api server? ...if you find the right magic rocm incantations I guess Double thanks for the 3rd party mac mini SSD tip - eagerly awaiting delivery!
- ineedasername 2y agoAll these have quickly become this generation’s “can it run Doom?”
- JKCalhoun 2y agoI did not see (understand) how multiple Raspberry Pis are being used in parallel. Maybe someone can point me in the right direction to understand this.
- jonatron 2y agoBlog post from same author explaining https://b4rtaz.medium.com/how-to-run-llama-3-405b-on-home-devices-build-ai-cluster-ad0d5ad3473b https://b4rtaz.medium.com/how-to-run-llama-3-405b-on-home-de...
- walrus01 2y agoNoteworthy nothing there really seems to be raspberry pi specific, as the raspberry pi os is based on debian, the same could be implemented on N number of ordinary x86-64 small desktop PCs for a cheap test environment. You can find older dell 'precision' series workstation systems on ebay with 32GB of RAM for pretty cheap these days, four of which together would be a lot more capable than a raspberry pi 5.
- dankle 2y agoOn cpu or the NPU “AI” hat?
- ninetyninenine 2y agoReally there needs to be a product based off of LLMs similar to Alexa or Google home where instead of connecting to the cloud it’s a locally run LLM. I don’t know why one doesn’t exist yet or why no one is working on this
- fabiensanglard 2y ago> locally run LLM You mean like Ollama + llamacpp ?
- ninetyninenine 2y agoYeah but packaged into a singular smart speaker product. We know there’s a market out there for Alexa and Google home. So this would be the next generation of that. It’s the nobrainer next step.
- unshavedyak 2y agoWouldn't it be due to price? Quality LLMs are expensive, so the real question is can you make a product cheap enough to still have margins and a useful enough LLM that people would buy?
- johntash 2y agoYou can kind of get there with home assistant. I'm not sure if it can use tools yet, but you can expose stuff you'd ask about like the weather/etc.
- ninetyninenine 2y agoI have both and the one offered is really weak. If you pay you can use Gemini but it doesn’t have agentic access to smart home controls meaning you can no longer ask it to turn on or off the lights
- jaybro867 2y ago[dead]
- czk 2y agois it just me or does calling these distilled models 'DeepSeek R1' seem to be a gross misrepresentation of what they actually are people think they can run these tiny models distilled from deepseek r1 and are actually running deepseek r1 itself its kinda like if you drove a civic with a tesla bodykit and said it was a tesla
- simonw 2y agoIf anyone wants to try this model on a Mac (it looks like they used something like DeepSeek-R1-Distill-Llama-8B) my new llm-mlx plugin can run it like this: brew install llm # or pipx install llm or uv tool install llm llm install llm-mlx llm mlx download-model mlx-community/DeepSeek-R1-Distill-Llama-8B llm -m mlx-community/DeepSeek-R1-Distill-Llama-8B 'poem about an otter' It's pretty performant - I got 22 tokens/second running that just now: https://gist.github.com/simonw/dada46d027602d6e46ba9e4f484775b5 https://gist.github.com/simonw/dada46d027602d6e46ba9e4f48477...
- lostmsu 2y agoI wonder what are the bandwidth requirements at higher rates, and is this still feasible to utilize GPUs?