4 ms·
Here is a larger-scale comparison of Cloud TPU and Google Cloud GPU performance and cost (focused on Cloud TPU Pods): https://cloud.google.com/blog/products/ai-
by zak 8y ago
Here is a larger-scale comparison of Cloud TPU and Google Cloud GPU performance and cost (focused on Cloud TPU Pods):
https://cloud.google.com/blog/products/ai-machine-learning/now-you-can-train-ml-models-faster-and-lower-cost-cloud-tpu-pods https://cloud.google.com/blog/products/ai-machine-learning/n...
All the code used in that comparison is open source, and there is a detailed methodology page with instructions that you can follow if you want to reproduce the results:
https://github.com/tensorflow/tpu/blob/master/benchmarks/ResNet-50_v1.5_Performance_Comparison_TensorFlow_1.12_GCP.md https://github.com/tensorflow/tpu/blob/master/benchmarks/Res...
Also, Cloud TPUs are available to everyone for free via Colab. Here is a sample Colab that shows how to train a Keras model on the Fashion MNIST dataset using the Adam optimizer:
https://colab.research.google.com/github/tensorflow/tpu/blob/master/tools/colab/fashion_mnist.ipynb https://colab.research.google.com/github/tensorflow/tpu/blob...
(I work on Cloud TPUs)
- twtw 8y agoAre there plans for preemptible TPU pods? As is, it looks like p3 16xlarge spot instance (or probably preemptible gcp 8xV100) are still by far the most cost effective option. You have to do a bit extra to tolerate preemption, but it's worth the ~80% savings. Also, the "TPU pod is 200x faster than single v100" comparison is a little goofy. Might as well say Summit or Titan is faster than my desktop.
- zak 8y agoAt present, preemptible Cloud TPU v2 and v3 devices are widely available, and they are likely to be the most cost-effective option for training any of the models listed here, often by a wide margin: https://cloud.google.com/tpu/docs/tutorials https://cloud.google.com/tpu/docs/tutorials Definitely appreciate your point about comparing Summit to your desktop - however, the difference is that you can't rent Summit via any public cloud, whereas you _can_ rent Cloud TPU Pods. You might find the comparison between 8 x V100 GPUs on GCP and a full Cloud TPU Pod more relevant - in that case, as of the time the Google Cloud blog post linked above was published, a full Cloud TPU Pod delivered a 27X speedup at 38% lower cost for a large-scale ResNet-50 training run, all without requiring any code changes to scale beyond a single device.
- twtw 8y agoThanks for your response. I'm interested particularly in preemptible pods - I see here https://cloud.google.com/tpu/docs/pricing https://cloud.google.com/tpu/docs/pricing pricing for preemptible single units, but there is no indication that preemptible pods are available.
- deleted 8y ago[deleted]