4 ms·
Author of the blog post here. Cloud TPUs are designed to maximize performance-per-dollar, so you are right that pure performance comparisons at maximum scale d
by zak 7y ago
Author of the blog post here.
Cloud TPUs are designed to maximize performance-per-dollar, so you are right that pure performance comparisons at maximum scale don't tell the whole story.
The most straightforward performance-per-dollar comparisons would be among several different hardware configurations across the major public clouds. However, we haven't yet seen any other public cloud MLPerf submissions at scales comparable to Cloud TPU Pods, so there isn't currently a strong baseline available for comparison. It's also not clear whether public cloud networking will ultimately be able to match the performance of the network hardware that was used to produce the largest-scale on-premise MLPerf submissions.
- rrss 7y agoWhy performance per dollar over performance per watt? Performance per watt doesn't care where the systems are (on-prem vs cloud) and should track perf/$ (unless perf/$ is mostly determined by subsidies).
- frankchn 7y agoI think hardware acquisition costs can dominate GPU prices, so perf/watt might not be a realistic measure for the cost of running such large scale/high performance experiments. For instance, the DGX-2 has an MSRP of $399,000 and consumes 10 kW of power [1]. The average commercial electricity cost across the US is about $0.11/kWh [2], so a DGX-2 running at full tilt costs $1.10 an hour in electricity. Thus, a DGX-2 running at 100% utilization for 3 years costs $9,636 in electricity, which is ~2.5% of the cost of the box itself. Of course, you probably could get DGX-2s for a lot cheaper if you are buying 100 of them, but the acquisition costs are still going to be significant vis-a-vis power costs. [1]: https://www.anandtech.com/show/12587/nvidias-dgx2-sixteen-v100-gpus-30-tb-of-nvme-only-400k https://www.anandtech.com/show/12587/nvidias-dgx2-sixteen-v1... [2]: https://www.pacificpower.net/about/rr/cpc.html https://www.pacificpower.net/about/rr/cpc.html
- zak 7y agoWhen comparing different hardware configurations within or between public clouds, power measurements generally aren't available, whereas prices or realistic price estimates generally are.
- yaroslavvb 7y agoCan you elaborate on the performance-per-dollar part? IE, I'm seeing GCP providing V100 at $2.48/hour and a TPUv3 at $8.00/hour
- rrss 7y agoProbably Google is still doing that 1 "Cloud TPUv3" = 4 TPUv3 chips.
- zak 7y agoSure. The $2.48/hour per V100 GPU on GCP does not include the price of the CPU host; that is purely the price to rent a single accelerator. By contrast, a network-attached Cloud TPU v3 device includes both a CPU host and four connected TPU v3 chips that collectively deliver up to 420 teraflops. Furthermore, each individual V100 GPU on GCP has 16 GB of memory, whereas the Cloud TPU v3 device has 128 GB of HBM. The best apples-to-apples performance-per-dollar comparison we have publicly available was published last fall, and it compared the performance and cost of using various Cloud TPU v2 Pod slice sizes with the performance and cost of using various numbers of V100 GPUs attached to a single GCP host: https://cloud.google.com/blog/products/ai-machine-learning/now-you-can-train-ml-models-faster-and-lower-cost-cloud-tpu-pods https://cloud.google.com/blog/products/ai-machine-learning/n... We went to great lengths to ensure that we trained exactly the same version of ResNet-50 to the same accuracy in the same way across all hardware configurations. The methodology predated MLPerf and is documented in full here: https://github.com/tensorflow/tpu/blob/master/benchmarks/ResNet-50_v1.5_Performance_Comparison_TensorFlow_1.12_GCP.md https://github.com/tensorflow/tpu/blob/master/benchmarks/Res... If you were going to do a similar performance-per-dollar comparison today, the simplest approach might be to try to get the code from NVIDIA's MLPerf 0.6 submissions running at scale on one or more major public clouds using the fastest-available networking technology that each cloud provides: https://github.com/mlperf/training_results_v0.6/tree/master/NVIDIA/benchmarks https://github.com/mlperf/training_results_v0.6/tree/master/... It would be very interesting to see how distributed training performance using large-scale GPU clusters in public clouds compares with the published on-premise MLPerf performance numbers using exactly the same MLPerf code and methodology. With these measurements in hand, it would then be straightforward to make performance-per-dollar comparisons with Cloud TPU v3 Pod slices of various sizes.
- justicezyx 7y agoI cannot stand Google people's tendency to explain or sometimes rebuttal comments. You are talking to your customers who is paying or considering paying, or in search of products. Take the feedback, if it can be done, and it's beneficial, do it and report so. Or stop explaining... That's simply not professional for a cloud provider...