4 ms·
Weeks :? Looking at every other project from Apple that try to integrate with some opensource lib it'll be months/years if they didn't drop support for it befo
by machinekob 4y ago
Weeks :?
Looking at every other project from Apple that try to integrate with some opensource lib it'll be months/years if they didn't drop support for it before that :P
Did you get only 2.55x on BERT/other transformer vs CPU version?
- microtonal 4y agoI think in this case the PyTorch team also involved. They have already fixed some annoying bugs in the last two days (like matrix multiplication often failing because of buffer size mismatches). As for BERT inference performance, I am not sure what kind of speedups you are expecting. The M1 Pro gives me 2.6 TFLOPs in single precision matrix multiplication of 768x768 matrices. M1 Pro GPU performance is supposed to be 5.3 TFLOPS (not sure, I haven’t benchmarked it).
- machinekob 4y agoAhh nvm I was thinking about m1 max (my brain is damaged by the m1 naming and i didnt saw information that you are using m1 pro gpu)
- microtonal 4y agoRight, the Max should make a much bigger difference, since it has the same number of AMX units as the Pro, but double the GPU cores.