5 ms·
Let’s not pretend they’re free. Even if you’re heavily optimizing, one of the best local models you can run today is CodeLlama 34B. To run that effectively, you
by easygenes 3y ago
Let’s not pretend they’re free. Even if you’re heavily optimizing, one of the best local models you can run today is CodeLlama 34B. To run that effectively, you need at least a used RTX 3090: $800 (assuming your computer already has everything else you need to effectively use one). They take 350W by default, but you can power limit them to 200W and preserve most of the LLM performance. If you’re in an average place where residential power is on the order of $0.2/kWh then you’ll pay $1 for every ~25 hours of use in power.
The output will be significantly slower than what you’d get from GPT-3.5, so if your time is worth anything then that will be a huge factor that’s hard to calculate. Just assume everything will take about 3 times longer though.
If you amortize the cost of the 3090 for as long as you’d expect it to be relevant, you can estimate a rough monthly cost of the hardware. I’d give it just over a year at the current pace before you’d probably want a replacement. That’s about $50/mo.
Operation at commodity scale with the still significant lead OAI or Anthropic has on efficient operations has its advantages beyond just SOTA performance.
- grishka 3y agoOf course it's not free either way. But when you're running it on your own infrastructure, you are free from OpenAI's whims, and that's extremely important.
- kkielhofner 3y agoOverall I think this is the wrong view. Let’s go back to the 90s - OpenAI is a word processor you rent at the library, an RTX 3090 is a computer you own. Just like the 90s, early on you were buying or upgrading a computer every Christmas if not sooner. Hell, when I got a Pentium 100 (OMG) my heart sunk when the 120 came out like a month later. 20% increase in clock speed in a month!?! Obviously we all know now that it was well worth it, this is how we got our start, and likely why we’re on HN today. That RTX 3090 can also run thousands upon thousands of other models for any number of use cases at the same cap and op ex. It can enable and build anything from a sense of wonder to an extremely valuable, fulfilling, and long lasting career. Worst case it’s fun, a relatively cheap hobby. Speaking of fun it also plays games. The RTX 3090 is an investment, not a cost. That said, some fair points generally but: 1) 25 hours of use is A LOT. GPUs scale power consumption too. The GPU will use slightly higher than idle power to keep the model in memory, then spike to whatever your power limit is while the model is executing. Let’s say 30 seconds to generate as worst case. Human reads that, comprehends it, does something. Meanwhile the power scales back down until you hit it again. Point is, 25 hours is 3,000 30 second generation sessions - 100/day, every day. For $1 (plus base/idle load of machine). You can also suspend the machine or turn it off to cut down on idle-idle waste. 2) Are you doing any optimization? I mostly work on high scale and inference serving frameworks (not on 3090s) but if I had to guess with something like lmdeploy, AWQ int4, int8 kv cache, etc you should be far exceeding ChatGPT in tokens/s. For example, I do concurrency testing with 13b on the RTX 4090 in my workstation. It hits 115 tokens/s - beating the 15 tokens/s or so from ChatGPT should be pretty straightforward with 34b on an RTX 3090.