3 ms·
That's not a bad result, although for £320 for 4x Pi5s you could probably find a used 12GB 3080 and probably more than 10x token speed
by replete 2y ago
That's not a bad result, although for £320 for 4x Pi5s you could probably find a used 12GB 3080 and probably more than 10x token speed
- varispeed 2y ago> Deepseek R1 Distill 8B Q40 on 1x 3080, 60.43 tok/s (eval 110.68 tok/s) That wouldn't get on Hacker News ;-)
- jckahn 2y agoHNDD: Hacker News Driven Development
- geerlingguy 2y agoOr attach a 12 or 16 GB GPU to a single Pi 5 directly, and get 20+ tokens/s on an even larger model :D https://github.com/geerlingguy/ollama-benchmark?tab=readme-ov-file#deepseek https://github.com/geerlingguy/ollama-benchmark?tab=readme-o...
- littlestymaar 2y agoReading the beginning of your comment I was like “ah yes I saw Jeff Geerling do that on a video”. Then I saw you github link and your HN handle and I was like “Wait, it is Jeff Geerling!”. :D
- ziml77 2y agoHaha I had nearly the same thing happen. First I was like "that sounds like something Jeff Geerling would do". Then I saw the github link and was like "ah yeah Jeff Geerling did do it" and then I saw the username and was like "oh it's Jeff Geerling!"
- replete 2y agoThanks for sharing. Pi5 + cheap AMD GPU = convenient modest LLM api server? ...if you find the right magic rocm incantations I guess Double thanks for the 3rd party mac mini SSD tip - eagerly awaiting delivery!
- geerlingguy 2y agollama.cpp runs great with Vulkan, so no ROCm magic required!
- HPsquared 2y agoOr a couple of 12GB 3060s.
- agilob 2y agoSome of us also worry about energy consumption
- talldayo 2y agoThey idle at pretty low wattages, and since the bulk of the TDP is rated for raster workloads you usually won't see them running at full-power on compute workloads. My 300w 3070ti doesn't really exceed 100w during inference workloads. Boot up a 1440p video game and it's a different story altogether, but for inference and transcoding those 3060s are some of the most power efficient options on the consumer market.
- aljarry 2y agoInteresting, my 3060 uses 150-170W with 14B model on Ollama, according to nvidia-smi.
- moffkalast 2y ago48W vs 320W as well though.
- replete 2y ago10x pi5s would be 480w vs 360w
- moffkalast 2y agoA Pi 5 uses about 12W full tilt, so 120W in that case. But the price comparison is with only 4 of them.