4 ms·
Power concerns aside, individual chips in a TPU pod don't actually have a ton of vRAM; they rely on fast interconnects between a lot of chips to aggregate vRAM
by kajecounterhack 2y ago
Power concerns aside, individual chips in a TPU pod don't actually have a ton of vRAM; they rely on fast interconnects between a lot of chips to aggregate vRAM and then rely on pipeline / tensor parallelism. It doesn't make sense to try to sell the hardware -- it's operationally expensive. By keeping it in house Google only has to support the OS/hardware in their datacenter and they can and do commercialize through hosted services.
Why do you want the hardware vs just using it in the cloud? If you're training huge models you probably don't also keep all your data on prem, but on GCS or S3 right? It'd be more efficient to use training resources close to your data. I guess inference on huge models? Still isn't just using a hosted API simpler / what everyone is doing now?