6 ms·
Cloud TPU pods are seriously amazing. I'm a researcher at Google working on speech synthesis, and they allow me to flexibly trade off resource usage vs. time to
by rryan 7y ago
Cloud TPU pods are seriously amazing. I'm a researcher at Google working on speech synthesis, and they allow me to flexibly trade off resource usage vs. time to results with nearly linear scaling due to the insanely fast interconnect. TPUs are already fast (non-pods, i.e. 8 TPU cores are 10x faster for my task than 8 V100s) but having pods open up new possibilities I couldn't build easily with GPUs. As a silly example, I can easily train on a batch size of 16k (typical batch size on one GPU is 32) if I want to by using one of the larger pod sizes, and it's about as fast as my usual batch size as long as the batch size per TPU core stays constant. Getting TPU pod quota was easily the single biggest productivity speedup my team has ever had.
- totoglazer 7y agoHow, if at all, do you have to tweak model architecture or hyperparameters for pods vs tpu vs gpu?
- rytill 7y ago16k batch size going that fast... Damn. Use it for good.
- p1esk 7y agoI thought TPUs are comparable in speed to V100. What makes them 10x faster for your task?
- sdenton4 7y agoIIUC, the difference is that the TPU allows insane data parallelism - exactly the giant batch sizes that rryan mentions. Google carried out extensive research on batch size, demonstrating that you can greatly increase training efficiency with giant batches. (And with tweaks to optimizers or some regularization, can push the "efficient" batch size domain even further.) https://ai.googleblog.com/2019/03/measuring-limits-of-data-parallel.html?m=1 https://ai.googleblog.com/2019/03/measuring-limits-of-data-p...
- p1esk 7y agoBatch size is only limited by memory. For comparison Nvidia DGX-2 has 512GB.
- sdenton4 7y agoHere's a writeup on tpu vs got throughput: https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transformers-bert/ https://timdettmers.com/2018/10/17/tpus-vs-gpus-for-transfor... The underlying principle is that even if the silicon isn't strictly faster, you can shove a dumb amount of data through with the right architecture. And efficient training with data parallelism means that it's a worthwhile strategy.
- p1esk 7y agoI prefer actual benchmark results, which indicate TPU is similar in performance to v100. Unless you have some data that proves otherwise?
- sdenton4 7y agoThis manifests in the training time + cost for imagenet in the stanford DAWN benchmark. The (partial) TPU pods win on training time, with a clear exponential speedup as the pod fraction increases. (How do you use more cores? Increase the batch size.) https://dawn.cs.stanford.edu//2018/04/30/dawnbench-v1-results/ https://dawn.cs.stanford.edu//2018/04/30/dawnbench-v1-result...
- p1esk 7y agoI’m looking at Imagenet training time for Resnet-50 [1] and it appears that 32 V100 chips are faster than half of a TPUv2 pod (64 chips). Are we looking at different benchmarks? [1] https://dawn.cs.stanford.edu//benchmark/ImageNet/train.html https://dawn.cs.stanford.edu//benchmark/ImageNet/train.html
- 7y ago
- sytelus 7y agoAre TPUs drop-in replacement for CUDA if you were using TF?Can I simply change device from CUDA to TPU and run any TF code? Last I heard, TPUs still had long way to go towards making this happen...
- nl 7y agoIf you are using TF, you can look in Tensorboard at your model graph and it will show you any incompatible operations. These days it's pretty good. Not perfect, but you can do RNNs on it now. The FAQ has decent docs: https://cloud.google.com/tpu/docs/faq https://cloud.google.com/tpu/docs/faq