6 ms·
It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8
by fancyfredbot 2y ago
It looks amazing but I wish we could stop playing silly games with benchmarks. Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Why leave out TPUv6 in the comparison?
Why compare fp64 flops in the El Capitan supercomputer to fp8 flops in the TPU pod when you know full well these are not comparable?
[Edit: it turns out that El Capitan is actually faster when compared like for like and the statement below underestimated how much slower fp64 is, my original comment in italics below is not accurate] (The TPU would still be faster even allowing for the fact fp64 is ~8x harder than fp8. Is it worthwhile to misleadingly claim it's 24x faster instead of honestly saying it's 3x faster? Really?)
It comes across as a bit cheap. Using misleading statements is a tactic for snake oil salesmen. This isn't snake oil so why lower yourself?
- shihab 2y agoI went through the article and it seems you're right about the comparison with El Capitan. These performance figures are so bafflingly misleading. And so unnecessary too- nobody shopping for AI inference server cares at all about its relative performance vs a fp64 machine. This language seems designed solely to wow tech-illiterate C-Suites.
- cheptsov 2y agoI think it’s not misleading, but rather very clear that there are problems. v7 is compared to v5e. Also, notice that it’s not compared to competitors, and the price isn’t mentioned. Finally, I think the much bigger issue with TPU is the software and developer experience. Without improvements there, there’s close to zero chance that anyone besides a few companies will use TPU. It’s barely viable if the trend continues.
- latchkey 2y agoThe reference to El Capitan, is a competitor.
- cheptsov 2y agoAre you suggesting NVIDIA is not a competitor?
- latchkey 2y agoYou said: "notice that it’s not compared to competitors" The article says: "When scaled to 9,216 chips per pod for a total of 42.5 Exaflops, Ironwood supports more than 24x the compute power of the world’s largest supercomputer – El Capitan – which offers just 1.7 Exaflops per pod." It is literally compared to a competitor.
- cheptsov 2y agoI believe my original sentence was accurate. I was expecting the article to provide an objective comparison between TPUs and their main competitors. If you’re suggesting that El Capitan is the primary competitor, I’m not sure I agree, but I appreciate the perspective. Perhaps I was looking for other competitors, which is why I didn’t really pay attention to El Capitan.
- sebzim4500 2y ago>Without improvements there, there’s close to zero chance that anyone besides a few companies will use TPU. It’s barely viable if the trend continues. I wonder whether Google sees this as a problem. In a way it just means more AI compute capacity for Google.
- mupuff1234 2y ago> besides a few companies will use TPU. It’s barely viable if the trend continues That doesn't matter much of those few companies are the biggest companies. Even with Nvidia majority of the revenue is being generated by a handful of hyperscalers.
- imtringued 2y agoAlso, there is no such thing as a "El Capitan pod". The quoted number is for the entire supercomputer. My impression from this is that they are too scared to say that their TPU pod is equivalent to 60 GB200 NVL72 racks in terms of fp8 flops. I can only assume that they need way more than 60 racks and they want to hide this fact.
- charcircuit 2y ago>Why compare fp8 performance in ironwood to architectures which don't support fp8 in hardware? Because end users want to use fp8. Why should architectural differences matter when the speed is what matters at the end of the day?
- bobim 2y agoGP bikes are faster than dirt bikes, but not on dirt. The context has some influence here.
- zipy124 2y agoBecause it is a public company that aims to maximise shareholder value and thus the value of it's stock. Since value is largely evaluated by perception, if you can convince people your product is better than it is, your stock valuation, at least in the short term will be higher. Hence Tesla saying FSD and robo-taxis are 1 year away, the fusion companies saying fusion is closer than it is etc.... Nvidia, AMD, apple and intel have all been publishing misleading graphs for decades and even under constant criticism they continue to.
- fancyfredbot 2y agoI understand the value of perception. A big part of my issue here is that they've really messed up the misleading benchmarks. They've failed to compare to the most obvious alternative, which is Nvidia GPUs. They look like they've got something to hide, not like they're ahead. They've needlessly made their own current products look bad in comparison to this one understating the long-standing advantage TPUs have given Google. Then they've gone and produced a misleading comparison to the wrong product (who cares about El Capitan? I can't rent that!). This is a waste of credibility. If you are going to go with misleading benchmarks then at least compare to something people care about.
- zipy124 1y agoThat's fair enough, I agree with all your points :)
- segmondy 2y agoWhy not? If we line up to race. You can't say why compare v8 to v6 turbo or electric engine. It's a race, the drive train doesn't matter. Who gets to the finish line first? No one is shopping for GPU by fp8, fp16, fp32, fp64. It's all about cost/performance factor. 8 bits is as good as 32bits, great performance is even been pulled out of 4 bits...
- fancyfredbot 2y agoThis is like saying I'm faster because I ran (a mile) in 8 minutes whereas it took you 15 minutes (to run two miles).
- scottlamb 1y agoI think it's more like saying I ran a mile in 8 minutes whereas it took you 15 minutes to run the same distance, but you weigh twice what I do and also can squat 600 lbs. Like, that's impressive, but it's sure not helping your running time. Dropping the analogy: f64 multiplication is a lot harder than f8 multiplication, but for ML tasks it's just not needed. f8 multiplication hardware is the right tool for the job.
- fancyfredbot 2y agoIt's even worse than I thought. El Capitan has 43,808 MI300A APUs. According to AMD each MI300A can do 3922TF of sparse FP8 for a total of 171EF sparse FP8 performance, or 85TF non-sparse. In other words El Capitan is between 2 and 4 times as fast as one of these pods, yet they claim the pod is 24x faster than El Capitan.
- adrian_b 2y agoFP64 is more like 64 times harder than FP8. Actually the cost is even much higher, because the cost ratio is not much less than the square of the ratio between the sizes of the significands, which in this case is 52 bits / 4 bits = 13, and the square of 13 is 169.
- christkv 2y agoMemory size and bandwidth goes up a lot right?
- dekhn 2y agoGoogle shouldn't do that comparison. When I worked there I strongly emphasized to the TPU leadership to not compare their systems to supercomputers- not only were the comparisons misleading, Google absolutely does not want supercomputer users to switch to TPUs. SC users are demanding and require huge support.
- meta_ai_x 2y agoGoogle needs to sell to Enterprise Customers. It's a Google Cloud Event. Of course they have incentives to hype because once long-term contracts are signed you lose that customer forever. So, hype is a necessity
- deleted 2y ago[deleted]