5 ms·
> It’s just that none of the stages are great. Really? Nobody uses TF but plenty of people use XLA, which is what you allude to saying that TPU in Pytorch is p
by kajecounterhack 2y ago
> It’s just that none of the stages are great.
Really? Nobody uses TF but plenty of people use XLA, which is what you allude to saying that TPU in Pytorch is possible. That's arguably the most important piece; the sledgehammer that makes good performance and good devex compatible.
> I wish Google would compete with Nvidia directly and let me buy a TPU
I don't see the utility of buying _any_ GPUs. It's mainly a cost + availability optimization to own them yourself if you're doing a lot of continual training. Outside of the foundation model companies, most of us just use cloud services -- and I want those to be cheap and always use the latest thing.
Google has a highly optimized infrastructure to make sure unused resources (CPU, TPU, RAM, disk) are properly allocated. This means preemptible instances, highly co-located compute + data, etc. In theory everyone should win when you use ML hw through hyperscalers, because they collect on their structural advantages and your TCO is lower.
> Google could have Nvidia’s business and OpenAI’s
They arguably have a better _business_ on their hands but a worse _speculative outlook_ in the eyes of Mr. Market. Those are different things. There's an argument to be made that Google's valuation could be as high if they ran their public relations strategy as well as OpenAI / Nvidia / Tesla, which are all riding monumental hype.
- janalsncm 2y ago> most of us just use cloud services Yes, but because cloud services buy GPUs almost exclusively Nvidia, it would likely drive cloud prices down as well.
- kajecounterhack 2y agoCloud prices are already driven down by the fact that you can deploy with Google's cloud offerings. Also, if you weren't going to use Google's offerings you probably wouldn't be down to buy their HW either :) FWIW your PTSD with Tensorflow is shared by everyone at Google, it was just unavoidable because Google needed ML to be performant and they started by sacrificing devex for performance whereas Pytorch went ergo-first, performance later. In retrospect a prescient move by Meta, but now making the high performance stuff ergonomic is proving out for Google via XLA and JAX.