40 ms·
This is very impressive technology and engineering. However, I remain a bit skeptical of the business case for TPUs for 3 core reasons: 1) 100000x lower unit
by obblekk 4y ago
This is very impressive technology and engineering.
However, I remain a bit skeptical of the business case for TPUs for 3 core reasons:
1) 100000x lower unit production volume than GPUs means higher unit costs
2) Slow iteration cycle - these TPUv4 were launched in 2020. Maybe Google publishes one gen behind, but that would still be a 2-3 year iteration cycle from v3 to v4.
3) Constant multiple advantage over GPUs - maybe 5-10x compute advantage over off the shelf GPU, and that number isn't increasing with each generation.
It's cool to get that 5-10x performance over GPUs, but that's 4.5yrs of Moore's Law, and might already be offset today due to unit cost advantages.
If the TPU architecture did something to allow fundamentally faster transistor density scaling, it's advantage over GPUs would increase each year and become unbeatable. But based on the TPUv3 to TPUv4 perf improvement over 3 years, it doesn't seem so.
Apple's competing approach seems a bit more promising from a business perspective. The M1 unifies memory reducing the time commitment required to move data and switch between CPU and GPU processing. This allows advances in GPUs to continue scaling independently, while decreasing the user experience cost of using GPUs.
Apple's version also seems to scale from 8GB RAM to 128GB meaning the same fundamental process can be used at high volume, achieving a low unit cost.
Are there other interesting hardware for ML approaches out there?
- kccqzy 4y ago> Are there other interesting hardware for ML approaches out there? Google also has Coral, which is a non-cloud mobile-focused TPU that you can buy and plug in (USB or PCIe). https://coral.ai/products/ https://coral.ai/products/ Naturally, this is orders of magnitude less powerful than the kind of TPU being discussed here but I just like the idea of local compute.
- 1MachineElf 4y agoJust received my mSATA Coral TPU in the mail. I'd been waiting 11 months for it after backordering on Digikey. Perhaps this speaks to the parent comment's concerns over unit volume and iteration cycle? Hopefully that will improve in the future and modules like these will become widespread.
- cubefox 4y ago> If the TPU architecture did something to allow fundamentally faster transistor density scaling, it's advantage over GPUs would increase each year and become unbeatable. It is completely unreasonable to expect something like that.
- sebzim4500 4y ago>100000x lower unit production volume than GPUs This is obviously an exaggeration, I wonder what the actual ratio is between TPU proudction and e.g. A100 production.
- summerlight 4y agoProbably more close to 1000x? I see a fairly large number of TPU pod these days and I don't think A100 is as prevalent as high end consumer GPU, which is typically measured in millions, not billions.
- sliken 4y ago> 100000x lower unit production volume than GPUs means higher unit costs Two points. Nvidia's RTX 3000 series (3060 Ti, 3080, and many other flavors) ships 6 or more flavors per generation. The related silicon has names like the GA102, GA103, GA104, GA106, and GA107. So only 1/6th of the consumer market for Nvidia silicon can be amortized over any single design. I wouldn't be at all surprised to see Google making the TPUs by the million. I found a vague reference to 9 exaflops and single facilities (one of many) costing $4 billion to $8 billion. So I wouldn't assume that the consumer GPU market/number of silicon designs is 100,000 times larger than the TPUv4 market. > Slow iteration cycle True. Then again generations make much less difference than they used to. Gone are the days where even after a multiple generations that average performance increases by 2x. Sure nvidia's 4000 series claims 2x ... on raytracing. But normal game performance seems to be more like 15%. Sure various trickery like DLSS helps, but similar tricks are increasing the performance of older cards as well. Similarly apple's a14 -> a15 -> a16 (or m1 -> m2 if you prefer) chips have had modest performance increases and mostly have increases in perf/watt. > 4.5yrs of Moore's Law It's dead Jim.
- pclmulqdq 4y agoI believe that Nvidia uses the same chip from A16 up to the A100, and maybe for some of the Quadro chips. That easily puts the unit count into the several millions. The picture in this article shows 8 racks with (according to the paper on arxiv) has 16 TPU sleds each, 4 TPUs per sled. that's only 512 chips. According to the paper, it is one of eight in a 4096-chip supercomputer. If you give them 10-100 of those around the world, you get 40,000-400,000 chips. That's enough for reasonable scale. Nvidia should still have 100x (or more) their scale.
- the-rc 4y agoThey are not all the same chips; some use different processes, for one. I assume that it would eat into any large volume efficiency claims. I am sure there are more A100s than TPUv4 chips out there, but I would hesitate to say the scale is over 100x.
- PragmaticPulp 4y ago> 3) Constant multiple advantage over GPUs - maybe 5-10x compute advantage over off the shelf GPU, and that number isn't increasing with each generation. The advantage of this hardware isn’t just the raw compute capacity. It’s the massive IO bandwidth. GPUs are fast, but they’re not designed for massive all-to-all communication across large networks. That’s where modules like this shine.
- lostmsu 4y agoWhy do you need all-to-all? With deep networks you could just put different layers on different GPUs, then run the forward and backward passes asynchronously. They would only need to communicate in a chain-like manner.
- atty 4y agoJust as a quick example (there are lots of other examples), I currently am working on a model that does significant tensor-level parallelism. The input to the model is huge, so we replicate the weights for the encoder part of the model over 12 A100s, and then run an all-reduce over the outputs before then distributing that reduced tensor to another set of GPUs for the next set of operations. In this setup we need significant bandwidth not only between every device in a node but also across nodes (infiniband).
- mirker 4y agoIsn’t the simpler example data parallel training? That’s a reduce and broadcast.
- atty 4y agoYou’re correct, I was going to add on to my answer that this is combined with DDP for a “3D” style parallelism that could specifically benefit from the TPU’s network topology, but by the time I got done fixing all the typos (and still missing a few) from writing it on my phone I completely forgot :)
- 4y ago
- fnbr 4y agoBiggest advantage is the high bandwidth interconnect. It’s very painful training large models on multiple GPUs. TPUs help a lot with that, although the software side of things is pretty brutal.
- jsnell 4y ago> 2) Slow iteration cycle - these TPUv4 were launched in 2020. Maybe Google publishes one gen behind, but that would still be a 2-3 year iteration cycle from v3 to v4. Is that substantially different from the Nvidia release cadence? According to the Wikipedia announcement dates, their last three generations were H100 in March 2022, A100 in May 2020, V100 in March 2017.
- ctchocula 4y agoAs of now, Apple's version is uncompetitive for training. It is decent for inference, which is good but doesn't cover the training usecase which is the main usecase for TPUv4 and Nvidia GPUs like A100. A100 provides 312 TFlops, but M1 GPU only 5.1 TFlops, which is 2 orders of magnitude slower. This blog [1] does make the argument that M1 Max has about 8x lower performance than a consumer grade Nvidia GPU 2080, but also uses 8x less watts, so it's possible Apple can make a product with similar performance/watt. However, I would say that they are farther away than TPUv4 for now, because the product doesn't exist.