2 ms·
I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on th
by hyperhello 11d ago
I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
- gruez 11d ago"If they are selling it for less than it cost to make, buy as much as you can." -- Warren Buffett
- taraindara 11d agoOnly caveat is you’re buying time. Not a physical good. It’s only worth what you’re able to get out of it in that time.
- rlindsey123 11d agoIs that a real quote? Golden if true
- tyre 11d agoFor their current models, served directly from their infrastructure, they are profitable after training (which all present models are.) I don't know when we'll have an open equivalent to Fable, let alone whatever (insane) hardware you'd need to run it locally.
- deleted 11d ago[deleted]
- epistasis 11d agoMore than that, running hundreds of conversation streams at once is essentially the same cost as running a single conversation. And then you add on the secondary benefit of having the GPUs running nearly all the time rather than mostly idle... Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.
- ASalazarMX 11d agoThe comparison is about how many tokens you buy vs how much hardware you could buy with the same money. It's as saying "if you have rib eyes at Applebee's every day, how long until cooking your own rib eyes pays for itself". If you don't consume many of tokens, it will likely never pay for itself. If you do, though, it will have trade-offs, but you'll probably save money in the end.