3 ms·
Unless I misread it, are they saying local GPUs use less energy? That’s surprising, almost unbelievable, due to batching. Local is usually not batched.
by api 17d ago
Unless I misread it, are they saying local GPUs use less energy?
That’s surprising, almost unbelievable, due to batching. Local is usually not batched.
- stymaar 17d agoSmall models are much smaller than frontier models though, which is how they end up consuming less energy despite low batch count. (Though with local models growing strong agentic capabilities, batching becomes a reality with local models as well).
- frumiousirc 17d ago> Unless I misread it, are they saying local GPUs use less energy? You misread it. From the abstract: > local accelerators achieve at least 1.4× lower IPW than cloud accelerators running identical models That's "intelligence per watt". They also have IPJ, per Joule. So, they find local is 40% "dumber" than cloud for the same power or 40% more power for the same "intelligence". Tables 13 and 14 summarize their IPW and IPJ metrics. But, to your actual point, I think the "local is 40% dumber per watt than cloud" message is still an understatement. And maybe this is something I failed to find in the paper but they seem to ignore the "idle baseline" costs and talks about explicitly focusing on the power consumption of just the accelerator under load. There is a large baseline power consumption just to support the accelerator. CPUs, memory, PS losses, network, fans, general environment cooling. This "cost floor" is different for data centers and a "random local computer" and I think must be in favor of data centers which are designed and built with efficiency in mind. Idleness should also be considered. My local GPUs at $WORK and home are idle more than they are used. Idle time energy in real world scenarios should be somehow attributed to those brief, punctuated times when LLM functions are actually active on the accelerator. Actual, local LLM usage of a GPU is brief (assuming one user per PC). Even with my heavy usage developing s/w I'd guess I heat up a GPU about one hour per day total, sometimes much less. If that is local then one must pay 23 hours of idleness for that 1 hour of "intelligence". Of course a local PC is used for other things and the idleness penalty must somehow account for that. OTOH, data centers try to maximize utilization so their idle time penalty would be much less, perhaps close to zero, by construction.
- fc417fc802 16d ago> Of course a local PC is used for other things and the idleness penalty must somehow account for that. I'd argue the idleness penalty only applies if you wouldn't have otherwise had equivalent hardware. If you already have the exact same dGPU for gaming and it doubles up for inference then the only inefficiency is power consumption (both by the GPU and potentially by AC for your living space). Conversely I think we should consider that privacy, distributed compute that you can access without an intermediary, and a more distributed power grid all provide net benefits to society at large.
- zozbot234 17d agoYou can most definitely batch local models and do unattended inference on a 24/7 basis to maximize utilization on local hardware too. The limits are usually set by some combination of memory utilization for KV cache (particularly on small dGPUs) and overall thermals/power limits (particularly on iGPUs with unified RAM/VRAM). (If you're not near thermal limits, the main alternative to batching is to use MTP or speculative decoding in order to raise arithmetic intensity and speed with the same memory utilization. But batching requests is generally viewed as preferable.) Newer models, especially from DeepSeek, do a nice job of reducing KV cache memory impact for any given context length and/or amount of parallel sessions, so batching on local hw really ought to be quite feasible.