3 ms·
Are there plans for preemptible TPU pods? As is, it looks like p3 16xlarge spot instance (or probably preemptible gcp 8xV100) are still by far the most cost ef
by twtw 8y ago
Are there plans for preemptible TPU pods?
As is, it looks like p3 16xlarge spot instance (or probably preemptible gcp 8xV100) are still by far the most cost effective option. You have to do a bit extra to tolerate preemption, but it's worth the ~80% savings.
Also, the "TPU pod is 200x faster than single v100" comparison is a little goofy. Might as well say Summit or Titan is faster than my desktop.
- zak 8y agoAt present, preemptible Cloud TPU v2 and v3 devices are widely available, and they are likely to be the most cost-effective option for training any of the models listed here, often by a wide margin: https://cloud.google.com/tpu/docs/tutorials https://cloud.google.com/tpu/docs/tutorials Definitely appreciate your point about comparing Summit to your desktop - however, the difference is that you can't rent Summit via any public cloud, whereas you _can_ rent Cloud TPU Pods. You might find the comparison between 8 x V100 GPUs on GCP and a full Cloud TPU Pod more relevant - in that case, as of the time the Google Cloud blog post linked above was published, a full Cloud TPU Pod delivered a 27X speedup at 38% lower cost for a large-scale ResNet-50 training run, all without requiring any code changes to scale beyond a single device.
- twtw 8y agoThanks for your response. I'm interested particularly in preemptible pods - I see here https://cloud.google.com/tpu/docs/pricing https://cloud.google.com/tpu/docs/pricing pricing for preemptible single units, but there is no indication that preemptible pods are available.
- deleted 8y ago[deleted]