10 ms·
It's true that we were initially using GPUs mostly for historical reasons, but over the last several years modern GPUs have been optimized for ML as much as any
by why_only_15 4y ago
It's true that we were initially using GPUs mostly for historical reasons, but over the last several years modern GPUs have been optimized for ML as much as anything else. If you read NVidia's marketing documents, they talk constantly about ML. The A100 is about as good, if not better, than the TPUv4 in terms of raw performance on ML workloads. The A100 can do 312 bf16 TFLOPs and costs $0.88/hr on Google Cloud [0] whereas the TPUv4 can do 275 bf16 TFLOPs and costs $0.97/hr on Google Cloud [1] [2]. The A100 is also generally speaking easier to program: it's supported by more frameworks and can perform more operations. The TPUv4 is in my understanding still worth it if you like JAX and/or you're doing lots of networking though.
WRT putting a TPU on a separate die -- this has been done for several years in the mobile space: Apple Neural Engine for iPhones, TPU (not same as server TPU) on Pixel, SNPE on Qualcomm, etc.
[0] https://cloud.google.com/compute/gpus-pricing https://cloud.google.com/compute/gpus-pricing
[1] https://cloud.google.com/tpu/pricing#v4-pricing https://cloud.google.com/tpu/pricing#v4-pricing
[2] this is somewhat unfair, because the GPU pricing number is for just the GPU and not the host it runs on, whereas the TPU pricing number (for TPU VMs) includes the host it runs on. If you include the price GCP charges for the host, preemptible A100s are about $1.20/hr. Why does Google make GPUs look cheaper than TPUs when they're not? Your guess is as good as mine.
- synergy20 4y agoMaybe Google is favoring TPUv4 over whatever GPU runs on its platform? With Hopper 100 on the way, I wonder when TPUv5 will come out. I also wonder how Intel's Gaudi2 vs Ponte Vecchio will work together, looks like duplicate efforts for me. AMD has its MI300 on the way, but it seems still far behind Nvidia|TPU|Intel at this point.