17 ms·
Author of the blog post here. We submitted multiple results using various Cloud TPU v3 Pod slice sizes to show the current achievable Transformer training effi
by zak 7y ago
Author of the blog post here.
We submitted multiple results using various Cloud TPU v3 Pod slice sizes to show the current achievable Transformer training efficiency at several scales:
https://mlperf.org/training-results-0-6 https://mlperf.org/training-results-0-6
We're actively improving the whole TPU software stack, so training efficiency is likely to continue to increase over time.
- gok 7y agoWhat I'm really asking is: how much effective compute throughput are you able to get during Transformer training relative to the amount of theoretical raw compute available?