3 ms·
This is somewhat orthogonal to the article, but the whole bubble on AI data centers seems to presume that the need for compute is so massive that it far exceeds
by harshaw 1mo ago
This is somewhat orthogonal to the article, but the whole bubble on AI data centers seems to presume that the need for compute is so massive that it far exceeds the expected optimizations we would expect with at scale inference (PIM, ASICs, etc). I would expect that there is a set of optimizations like this one (or variations) that would someone negate the buildout. But it's not really discussed.
- Tenoke 1mo agoThere's been a ton of optimizations already, it hasn't remotely reduced demand even temporarily. More efficiency just makes the compute have even higher ROI per $ and watt spent.
- roryirvine 1mo agoWith sufficient optimisation, there ought to be a tipping point beyond which local inference is good enough. And, sure, datacentre compute will still be needed for training but one of the biggest current uses will begin to taper off. The question really is how soon we reach that tipping point, and whether it's before or after the current bubble runs out of steam for some other reason.
- Tenoke 1mo ago>there ought to be a tipping point beyond which local inference is good enough There's no such ought really. Even at current levels you'd need like a 100x gain from here to approach current top proprietary models (probably a lot more for say Mythos or Mythos 2), and it's not like they are stoppng to improve. This is before we even account that you'd just be running 1 agent then, and not a swarm like you'd be able to in the cloud or that you can do only so much compression before you are losing out
- eureka7 1mo ago> (probably a lot more for say Mythos or Mythos 2) Not everyone needs that large of a model, though.
- UltraSane 1mo agoJevons Paradox shows that increasing efficiency can increase demand for a product by making it cost effective for more uses.
- deleted 1mo ago[deleted]
- petra 1mo agoIt's not just inference, some things done in data centers like simulations, testing, are complementary to inference. And in these types of hardware, the time between a successful prototype and a fully deployed product is pretty long. Maybe they count on that to know when to stop?