4 ms·
I get excited when hackers like Justine (in the most positive sense of the word) start working with LLMs. But every time, I am let down. I still dream of some
by acatton 3y ago
I get excited when hackers like Justine (in the most positive sense of the word) start working with LLMs.
But every time, I am let down. I still dream of some hacker making LLMs run on low-end computers like a 4GB rasbperry pi. My main issues with LLMs is that you almost need a PS5 to run the them.
- simonw 3y agoLLMs work on a 4GB Raspberry Pi today, just INCREDIBLY slowly. There's a limit to how much progress even the most ingenious hacker can make there - LLMs are incredibly computationally intensive. Those billion item matrices aren't going to multiply themselves!
- deleted 3y ago[deleted]
- plagiarist 3y agoPeople are working on it but it's just down to how many floating point operations can you do? I wonder if something like the Coral would help? I'd love having an LLM on a Pi, but I'll have to settle for a larger machine I can turn on to get more compute. At least for the time being.
- filterfiber 3y agoThe current bottleneck for most current hardware is RAM capacity than memory bandwidth and last is FLOPS/TOPS. The coral has 8 MB of SRAM which uh, won't fit the 2GB+ that nearly any decent LLM require even after being quantized. LLMs are mostly memory and memory bandwidth limited right now.
- jart 3y agoThanks for saying that. The last part of my blog post talks about how you can run Rocket 3b on a $50 Raspberry Pi 4, in which case llamafile goes 2.28 tokens per second.