5 ms·
Like I said, just feels I have from the consumer-hardware space. For several generations of GPU now most improvements come from packing more transistors into a
by DanielHB 2mo ago
Like I said, just feels I have from the consumer-hardware space. For several generations of GPU now most improvements come from packing more transistors into a larger die than packing more transistors closer to each other.
GPUs have been getting physically bigger with huge heatsinks and fans to support those bigger dies power consumption. Just compare the TDPs:
2020 RTX 3090: 350W
2022 RTX 4090: 450W
2025 RTX 5090: 575W
Bigger dies means lower capex of course, but the similar opex (maybe slightly lower as there is less physical hardware to maintain).
I seen some specialized hardware like google's TPUs. Not sure how they compare on performance per watt with GPUs though. Regardless the manufacturing processes are still the same (EUV) which is the thing that hasn't been improving. A fully optimized specialized hardware can at most deliver a single-time linear improvement (that could be very significant, for example 30% is still huge of course) and then little compared to normal GPUs.
I don't think renewable power generation is going to massively reduce costs for data centers, especially considering power transmission hasn't meaningfully reduced in cost. If anything the only thing that I think will have significant impact for data centers would be dedicated nuclear power plants physically located right next to the data center.
In fact I expect power generation to get more expensive as demand can increase faster than supply can be established. I imagine setting up new solar farms and transmission lines to be significantly harder (as in, takes longer time due to approvals and so on) than new data centers (which requires a single large location and I assume less approvals).
- ndriscoll 2mo agoIf you have somewhere to put them, you can get 2 440W panels for ~$350 if you're okay with intermittent or mostly daytime use. Or for ~$130/kWh you can extend that with batteries. Then you need an inverter or some kind of regulator, but all in your fully capitalized power is still less than an AMD or Intel GPU and a lot less than an nVidia GPU for home use (the context of this thread is how intelligence is not limited to datacenter deployments and can be done at home). Most home users probably aren't going to leave it running overnight all the time, so you really only need storage for morning/evening, maybe. If you had a 27B model on a cerebras-like chip, it'd probably be much faster than any human could interact with it, so it'd race to idle just like CPUs.
- eru 2mo ago> If you had a 27B model on a cerebras-like chip, it'd probably be much faster than any human could interact with it, so it'd race to idle just like CPUs. I don't think the main use case is for a human to directly interact with the raw token stream. You probably want reasoning and you want the thing to be able to program on its own. That uses way more tokens than you can read.
- ksec 2mo agoOn the RTX wattage over the years, perhaps it is important to note 5090 is nearly 20% larger die size. In terms of Pref per Watt is also a higher.
- eru 2mo ago> GPUs have been getting physically bigger with huge heatsinks and fans to support those bigger dies power consumption. Just compare the TDPs: You can also look at what's been happening in mobile and especially with Apple's integrated processors. They are more power constrained, so people worried more about power there.
- DanielHB 1mo agoI believe the power gains from apple stuff mostly come from CPU and memory interconnect, unrelated to consumer-grade PC GPUs. Server-grade GPUs have different memory interconnect architecture which I assume already has similar power efficiency gains. I think raw flops per watt come mostly from fab process, not architecture. This was my original point, fab process is not getting better at a linear (much less exponential) scale anymore.