4 ms·
Weeks later and we still don't know nothing. How do M1's TPU cores (the dedicated, so called "AI engine", not M1' GPU!) compare to current, consumer Nvidia TPU
by desmap 6y ago
Weeks later and we still don't know nothing. How do M1's TPU cores (the dedicated, so called "AI engine", not M1' GPU!) compare to current, consumer Nvidia TPU cores, eg those from the 3080 or 3090 (not Nvidia's GPU cores and please, no CPU and not some random Intel GPU!) re model training (not execution!)?
If there's someone who has an educated guess of the performance difference between...
M1's TPU vs 3090's TPU re training
...please let us know.
- EvgeniyZh 6y agoApple claims 11 trillion OPS, NVIDIA claims 285 TOPS, my educated guess is that real-life performance is off by same factor, leading to ~25x advantage of 3090 compared to M1
- trhway 6y agonvidia is ~30 TFLOPS, i.e 3x.
- EvgeniyZh 6y ago30 TFLOPS are general-purpose ones, for tensor cores, depending on what exactly you count can be 71 TFLOPS (fp32 accumulator), 142 TFLOPS (fp16 accumulator), or 282 TFLOPS (sparse tensors)
- andy_threos_io 6y agoOn all of this type of computation the real bottleneck is memory, size and bandwidth. You have to read and write back the data, you just simply can not perform faster than your memory bandwidth. ex If you have 2GB model ( 1G float16 ), which does not fit in any cache, and have to perform calculation on all of the data, than for 1 operation on all of the data 4 GB bandwidth is required. M1 17 times/s (if you don't make anything, so its a impossible max) 3090 233 times/s (that also not so practical) so 3090 10x M1 is a lower estimate. 3090 memory 24GB bandwidth 936.2 GB/s M1 memory 16GB max bandwidth 68GB/s (shared with other parts of the system)