4 ms·
I think the profits depend on how well they manage their fleet purchases (or possible sub-leasing?) to get high utilization without overloading or idle racks.
by ilaksh 3mo ago
I think the profits depend on how well they manage their fleet purchases (or possible sub-leasing?) to get high utilization without overloading or idle racks.
Because accelerators like H200, B300 etc. are highly parallel and designed to run like 200 or maybe 300 sequences at once (depends on the model, just guessing). I assume they finance the hardware and that cost per device or rack is the same whether each unit is handling 10 requests or 150 requests (aside from electricity).
And probably international customers factor into it to get good utilization over more of the night time. And it likely is something that they look at quarterly more seriously than monthly. The biggest risk to profits might be a downturn in business that causes some portion of the financed AI accelerators to go idle or get low utilization for some weeks (that they can't sublease).
- jaggederest 3mo agoSomeone on HN made a comment in one of these threads that we could bake the weights into something like Cerebras's wafer scale chips and serve essentially the entire world off a single wafer, which is a pretty wild thing to think about. You'd have to make new hardware any time you trained a model but that seems really worth it.
- ilaksh 3mo agoWell, Taalas has that kind of technology, but the chip they demoed is probably 20-100 times smaller than necessary since it's only an 8b model. But let's say they could someday scale that up to a much larger model, 72 large chips per wafer and each chip can do 1000 LLM requests at once (Vera Rubin?). So it's roughly the equivalent of an NVL72 rack. You might be able to serve something like 50000-60000 requests at once. So I think it's more like handling a small city's worth of customers per wafer than the world if you had that. I believe in less than 5 years we will get to that, but the model size and/or number of agents is going to keep going up also.
- Atotalnoob 3mo agoYou’d never be able to update it’s knowledge. LLMs need retraining to incorporate new knowledge. Baking them into wafers means they will be out of date by the time they finish the first wafers.
- jaggederest 3mo agoYes, of course, but all the LLMs are already out of date, so that doesn't seem to me to be a hard limiting factor. Even if they had a knowledge basis ~3 months out of date additionally, being able to serve 100x the requests per watt seems totally reasonable to me.
- Shorel 3mo agoSo, what? I don't see the C++ compiler standards or Newton's laws changing every day.
- lelanthran 3mo agoI think the future will have to include specialised host boards for memory chips. What I actually want is an FPGA board with a very large number of DDR3/DDR4 RAM slots arranged in banks (2, 4, 8 or even more banks). I want an FPGA board that can hold 1TB of DDR3/DDR4 RAM. The throttling point right now is not RAM, it's bus speed. Having different busses for banks of RAM alleviates that.