3 ms·
MPS (Metal backend for PyTorch) is pretty poor integrated right now most of my custom models are crashing and pure Swift + MPSGraph versions is working 3-10x fa
by machinekob 4y ago
MPS (Metal backend for PyTorch) is pretty poor integrated right now most of my custom models are crashing and pure Swift + MPSGraph versions is working 3-10x faster then PyTorch.
So I'm pretty sure there is A LOT of optimizing and bug fixing before we can even consider PyTorch on apple devices (and this is ofc. 0.1 or smth like that so maybe we'll get something working fine in next few years).
Also it'll be cool to compare Tensorflow vs PyTorch on metal device :)
- microtonal 4y agoMPS (Metal backend for PyTorch) is pretty poor integrated right now most of my custom models are crashing and pure Swift + MPSGraph versions is working 3-10x faster then PyTorch. So I'm pretty sure there is A LOT of optimizing and bug fixing before we can even consider PyTorch on apple devices Also, even for ops that do not crash, it often returns garbage. I built PyTorch from git this morning: >>> torch.arange(10, device="mps") tensor([0, 0, 0, 0, 0, 0, 0, 0, 0, 0], device='mps:0') >>> torch.ones(10, device="mps").type(torch.int32) tensor([1065353216, 1065353216, 1065353216, 1065353216, 1065353216, 1065353216, 1065353216, 1065353216, 1065353216, 1065353216], device='mps:0', dtype=torch.int32) Many bugs will probably be squashed over the coming weeks, so I am still very exited about MPS support. I have done some preliminary benchmarks with a spaCy transformer model and the speedup was 2.55x on an M1 Pro. Which is quite nice because transformers were already very fast on M1 Macs, thanks to the AMX units (PyTorch links against Accelerate, so the AMX units are used for matrix multiplication). If they can squeeze more performance out of M1 GPUs in the future, it will be a very nice speedup.
- machinekob 4y agoWeeks :? Looking at every other project from Apple that try to integrate with some opensource lib it'll be months/years if they didn't drop support for it before that :P Did you get only 2.55x on BERT/other transformer vs CPU version?
- microtonal 4y agoI think in this case the PyTorch team also involved. They have already fixed some annoying bugs in the last two days (like matrix multiplication often failing because of buffer size mismatches). As for BERT inference performance, I am not sure what kind of speedups you are expecting. The M1 Pro gives me 2.6 TFLOPs in single precision matrix multiplication of 768x768 matrices. M1 Pro GPU performance is supposed to be 5.3 TFLOPS (not sure, I haven’t benchmarked it).
- machinekob 4y agoAhh nvm I was thinking about m1 max (my brain is damaged by the m1 naming and i didnt saw information that you are using m1 pro gpu)
- microtonal 4y agoRight, the Max should make a much bigger difference, since it has the same number of AMX units as the Pro, but double the GPU cores.