5 ms·
Floating point operations (TFlop) != Tensor Operations (TOp)
by aub3bhat 9y ago
Floating point operations (TFlop) != Tensor Operations (TOp)
- nhaehnle 9y agoSure, but read the slide. They have 64k multiply-accumulate units, running at 700 MHz. That means ~46T/s multiply-accumulates, which means ~92T/s individual arithmetic ops. It's a standard way to measure this. I think it's fair to say that 92T/s 8-bit arithmetic ops is much less than 25x the 27T/s half-float operations of a GPU.
- deleted 9y ago[deleted]
- Symmetry 9y agoIf an 8-bit integer is sufficient for a problem there isn't anything to be gained by moving to 32 bit floats. You can't just use a 32-bit floating point operation to emulate 4 8-bit integer operations for free (or vice versa) so you just can't compare the two the way you're trying to. Especially since moving to larger precision values would balloon memory and bandwidth requirements. For an honest comparison find out what the 8-bit integer performance of the GPU is.
- Dylan16807 9y ago> you just can't compare the two the way you're trying to I don't think that accusation is justified. The part about float operations being better is only a side note. The core of the comment is that they are not inferior. If you needed to, you could snip wires to turn that half-float unit into an 8 bit unit. So treat the numbers as if they were the same thing. 27 vs. 92. That's not a 25x increase. Not even close. Something about this comparison seems either unfair or misleading. For example if a GPU doesn't engage most of its ALUs for certain sizes of input (cough GP10x cough), that's not a point in favor of the google design, that's just the GPU being broken.
- Symmetry 9y agoThey aren't inferior but you can't just multiply by 4 here either. You could turn a int32 adder into 4 int8 addres if the larger adder works on a ripple-carry principle but really everybody uses ripple-carry or carry-bypass. And a float32 is more complex and you could get 3 int8s out of it but you'd have lots of transistors left over in the execution logic. But simple quantity of execution logic is almost never an interesting constraint in a design. But the actually important part here is that the register and bypass networks to pass 4 bits of int8 data around are way more complicated than those required to pass a single float32 around and that's where Google's decision to restrict the flexibility of its TPU pays big dividends. NVidia's GPUs do not have broken designs. They're just making compromises based on the need to handle a wider variety of use cases.
- Dylan16807 9y ago> They aren't inferior but you can't just multiply by 4 here either. Yeah but there wasn't a suggestion to do so. Just by raw count there are issues with 25x. > NVidia's GPUs do not have broken designs. The part where the current generation sticks in a single FP16x2 unit per 128 FP32 units, so that if your code triggers them it runs 64x slower on FP16 while leaving all the FP32 units idle? That's broken as far as I can see, there to upsell you the pro cards. Anything that would make 8 bit math slower than 32 bit math is just a fundamental lack of forethought. It's not preferred by GPU design, and shouldn't be used as a point against GPUs in general.