4 ms·
> Even at non VC subsidized $/token prices, its still much cheaper to run cloud based models. On a price-per-wattage level, this is not true, people have done
by dvt 4mo ago
> Even at non VC subsidized $/token prices, its still much cheaper to run cloud based models.
On a price-per-wattage level, this is not true, people have done the math on /r/LocalLLaMA many times over[1]. Local models, while not as good as premier models (GPT 5.5, etc.), are like ~80%+ of the way there, and often converge to a similar solution after a few dead ends.
[1] https://www.reddit.com/r/LocalLLM/comments/1kshq4f/electricity_cost_of_running_local_llm_for_coding/ https://www.reddit.com/r/LocalLLM/comments/1kshq4f/electrici...
- fwip 4mo agoMaybe not per watt, but unless you already happen to own a 3900 cited by that post, you'd have to buy that as well, which is currently selling for around $1400 used.
- dvt 4mo agoI do have a 3090 Ti on my gaming PC, but even my old M1 MBP (with a mere 32gb of RAM) is quite competent and can run a quantized `Gemma4-26B-A4B` in the background while I do other stuff.
- ActorNightly 4mo agoThe MBP running Gemma4 is absolutely is useless for any real work.
- nozzlegear 4mo agoWhat is "real work"?
- ActorNightly 4mo agoWhere you are developing software. Its significantly faster to use google gemini and copy paste code back and forth compared to having gemini edit files for you.
- strictnein 4mo ago3090s are running $1400 now? Wowsers. I thought I was overspending when I bought 6x of them for around $800 a pop. Might be time to sell, to be honest. It's fun to have that at home, but I can't justify having $10k (with memory, mobo, cpu, etc) sitting in my basement without being fully utilized.
- karim79 4mo agoI'll take two of them. A thousand a piece.
- ClikeX 4mo agoTo be fair, I can also use that 3900 for other things locally. Not just AI.