3 ms·
The conclusions section of the paper is a good summary: "In the process, we learned ten lessons about DSAs and DNNs in general and about DNN DSAs specifically
by cokernel_hacker 5y ago
The conclusions section of the paper is a good summary:
"In the process, we learned ten lessons about DSAs and
DNNs in general and about DNN DSAs specifically that
shaped the design of TPUv4i:
1. Logic improves more quickly than wires and SRAM
⇒ TPUv4i has 4 MXUs per core vs 2 for TPUv3 and 1 for
TPUv1/v2.
2. Leverage existing compiler optimizations
⇒ TPUv4i evolved from TPUv3 instead of being a brand
new ISA.
3. Design for perf/TCO instead of perf/CapEx
⇒ TDP is low, CMEM/HBM are fast, and the die is not big.
4. Backwards ML compatibility enables rapid deployment
of trained DNNs
⇒TPUv4i supports bf16 and avoids arithmetic problems by
looking like TPUv3 from the XLA compiler’s perspective.
5. Inference DSAs need air cooling for global scale
⇒ Its design and 1.0 GHz clock lowers its TDP to 175W.
6. Some inference apps need floating point arithmetic
⇒ It supports bf16 and int8, so quantization is optional.
7. Production inference normally needs multi-tenancy
⇒ TPUv4i’s HBM capacity can support multiple tenants.
8. DNNs grow ~1.5x annually in memory and compute
⇒ To support DNN growth, TPUv4i has 4 MXUs, fast onand off-chip memory, and ICI to link 4 adjacent TPUs.
9. DNN workloads evolve with DNN breakthroughs
⇒ Its programmability and software stack help pace DNNs.
10. The inference SLO is P99 latency, not batch size
⇒ Backwards ML compatible training tailors DNNs to
TPUv4i, yielding batch sizes of 8–128 that raise throughput
and meet SLOs. Applications do not restrict batch size."
- ArtWomb 5y ago>>> 8. DNNs grow ~1.5x annually in memory and compute Wow! That's a massive growth rate for ML. TPUv3 was already faster than A100 in MLPerf. But this suggests a real breakthrough is needed to keep pace with future requirements. Each MXU already handles 16k ops per tick. And with the additional constraint of optimizing per watt rather than dollar, its quite the challange ;)
- omegalulw 5y ago> But this suggests a real breakthrough is needed to keep pace with future requirements. Not necessarily. Look at papers like the lottery ticket hypothesis - big ML models may be doing better simply because gradient descent just isn't doing a good enough job. Better optimizers would go a long way than just throwing compute at the the problem. Even if you can, it's impractical to use something like GPT-3 all the time.