13 ms·
Accelerated PyTorch Training on M1 Mac
- singularity2001 4y agoAnyone else getting "illegal hardware instruction"? (pytorch_env) ~/dev/ai/ python -c "import torch"
- zimpenfish 4y agoIIRC, when I had that problem, it was because it was loading the wrong arch for Python.
- dilielloneluca 4y agoI started collecting benchmarks of the M1 Max on PyTorch here: https://github.com/lucadiliello/pytorch-apple-silicon-benchmarks https://github.com/lucadiliello/pytorch-apple-silicon-benchm...
- MasterScrat 4y agoSmall code example in the PyTorch doc: https://pytorch.org/docs/master/notes/mps.html https://pytorch.org/docs/master/notes/mps.html
- nxpnsv 4y agoTried https://pytorch.org/tutorials/beginner/basics/quickstart_tutorial.html https://pytorch.org/tutorials/beginner/basics/quickstart_tut... with mps vs cpu. mps worked, but cpu actually was faster (16 vs 21s). Perhaps I am doing it wrong...
- nafizh 4y agoExciting!! But don't see comparison with any laptop Nvidia GPUs in terms of performance. That would be insightful.
- sudosysgen 4y agoIt compares unfavourably, but then again NVidia GPUs on laptop are massive powerhogs.
- smlacy 4y agoDo apple users really require the ability to train large ML models while mobile and without access to A/C power? Is this a real-world use case for the target market?
- sudosysgen 4y agoIndeed, I doubt anyone really needs that. And anyways while training a model you'd be lucky to get an hour of battery life even on an M1 Max.
- sbeckeriv 4y agoWhat is the * in the chart referencing?
- mrchucklepants 4y agoProbably supposed to be referencing the text under the plot stating the specific configuration of the hardware and software.
- sbeckeriv 4y agolooks like the website was updated after I posted. I used page search to look for the *.
- arecurrence 4y agoThis is much nicer ergonomics than what I had to do for tensorflow. It’s ostensibly out of the box support as a different torch device.
- mark_l_watson 4y agoI agree. I appreciated the M1/Metal TensorFlow support, but that was not as easy to setup.
- alfalfasprout 4y agoI mean, building tensorflow is generally an awful experience.
- dangrie158 4y agoYou must have installed it a while ago then :). I just recently did and only needed to I stall 2 packages via pip (i think tensoflow-macos and tenderfoot-metal ) which I found much better than wrangling with cuda and cudnn versions for Nvidia cards
- Scene_Cast2 4y agoI'm curious about the performance compared to something like, say, the RTX 3070.
- my123 4y agoLow. Apple doesn't have matrix math accelerators in their current GPUs. The neural engine is small and inference only. It's also only exposed by a far higher level interface, CoreML. Where it could still make sense is if you have a small VRAM pool on the dGPU and a big one on the M1, but with the price of a Mac, not sure that makes a lot of sense either in most scenarios compared to paying for a big dGPU.
- Kon-Peki 4y ago> Apple doesn't have matrix math accelerators in their current GPUs. That's because the M1 has a dedicated matrix math accelerator called AMX [1]. I've used it with both Swift and pure C. https://medium.com/swlh/apples-m1-secret-coprocessor-6599492fc1e1 https://medium.com/swlh/apples-m1-secret-coprocessor-6599492...
- my123 4y agoAMX is indeed very nice for FP64 where customer GPUs aren't an alternative at all. However, for lower precisions (which is what deep learning uses), you're much better off with a GPU.
- brrrrrm 4y agohave you actually benchmarked that? I think (someone please correct me if I'm way off here) the AMX instructions can hit ~2.8tflops (fp16) per co-processor and there are 2 on the 7-core M1. That's 5.6tflops vs the 4.6tflops the GPU can hit.
- johndough 4y agoOften the limiting factor is memory bandwidth instead of raw FLOPS, so dealing with 4 times larger data types (FP64 vs FP16) is a disadvantage.
- ekelsen 4y agoNice results! But why are people still reporting benchmark results on VGG? Does anybody actually use this network anymore? Better would be mobilenets or efficientNets or NFNets or vision transformers or almost anything that's come out in the 8 years since VGG was published (great work it was at the time!).
- 6gvONxR4sf7o 4y ago> But why are people still reporting benchmark results on VGG? It makes me feel like i’m missing something! Is is still used as a backbone in the same way as legacy code is everywhere, or is it something else entirely??
- sanxiyn 4y agoVGG works better for style transfer than ResNet (this is a surprising result, but empirically true), but that's the only case I am aware of. https://arxiv.org/abs/2104.05623 https://arxiv.org/abs/2104.05623
- plonk 4y ago> Does anybody actually use this network anymore? Why not? It's still good for simple classification tasks. We use it as an encoder for a segmentation model in some cases. Most ResNet variants are much heavier.
- jph00 4y agoI don't think that's true - have a look at this analysis here: https://www.kaggle.com/code/jhoward/which-image-models-are-best https://www.kaggle.com/code/jhoward/which-image-models-are-b... Those slow and inaccurate models at the bottom of the graph are the VGG models. A resnet34 is faster and more accurate than any VGG model. And there are better options now -- for example resnet34d is as fast as resnet34, and more accurate. And then convnext is dramatically better still.
- YetAnotherNick 4y ago> ResNet > VGG: ResNet-50 is faster than VGG-16 and more accurate than VGG-19 (7.02 vs 9.0); ResNet-101 is about the same speed as VGG-19 but much more accurate than VGG-16 (6.21 vs 9.0). https://github.com/jcjohnson/cnn-benchmarks#:~:text=ResNet%20%3E%20VGG%3A%20ResNet%2D50,16%20(6.21%20vs%209.0) https://github.com/jcjohnson/cnn-benchmarks#:~:text=ResNet%2....
- toppy 4y agoDoes speed up refer to absolute value or percentage?
- deleted 4y ago[deleted]
- cj8989 4y agoreally hope to see some comparisons with nvidia gpus!
- amelius 4y ago> Accelerated GPU training is enabled using Apple’s Metal Performance Shaders (MPS) as a backend for PyTorch. What do shaders have to do with it? Deep learning is a mature field now, it shouldn't need to borrow compute architecture from the gaming/entertainment field. Anyone else find this disconcerting?
- my123 4y agoApple doesn't have a separate API tailored towards compute only, but a single unified API that makes concessions to both. Concessions towards compute: a C++ programming language for device code (totally unlike what's done for most graphics APIs!) Concessions towards graphics: no single-source programming model at all for example...
- sudosysgen 4y agoMany GPUs allow you to write device code in C++ via SYCL. It works well enough.
- dagmx 4y agoShaders are just the way compute is defined on the GPU. Why is that concerning to you?
- my123 4y agoThat terminology isn't used at all in GPGPU compute APIs specifically tailored for that purpose, which use quite different programming models where you can mix host and device code in the same program. And there are "GPUs" today that can't do graphics at all (AMD MI100/MI200 generations) or in a restricted way (Hopper GH100) which has the fixed function pipeline only on two TPCs, for compatibility, but running very slowly due to that.
- alfalfasprout 4y agoThere's absolutely a lot of "graphics" terminology that spills into GPGPU. For example, texture memory in CUDA :) The reality is that GPU's, even the ones that can't output video, are ultimately still using hardware that largely is rooted in gaming. Obviously the underlying architectures for these ML cards are moving away from that (increasingly using more die space for ML related operations) but many of the core components like memory are still shared. It boils down to the fact that at the end of the day they're linear algebra processors.
- alexfromapex 4y agoSince it's tangentially relevant, if you have an M1 Mac I've created some boilerplate for working with the latest Tensorflow with GPU acceleration as well: https://github.com/alexfromapex/tensorexperiments https://github.com/alexfromapex/tensorexperiments . I'm thinking of adding a branch for PyTorch now.
- masklinn 4y agoDid you compare that to Apple's tf plugin to see what was what?
- galoisscobi 4y agoThis is great! Appreciate the note on H5Py troubleshooting as well.
- lekevicius 4y agoCuriously neither PyTorch nor Tensorflow currently use M1's Neural Engine. Is too limited? Too hard to interact with? Not worth the effort?
- deleted 4y ago[deleted]
- RicoElectrico 4y agoMost probably Neural Engine is optimized for inference, not training.
- sillyinseattle 4y agoQuestion about terminology (no background in AI). In econometrics, estimation is model fitting (training, I guess), and inference refers to hypothesis testing (e.g. t or F tests). What does inference mean here?
- upwardbound 4y agoInference here means "running" the model. So maybe it has a similar meaning as in econometrics? Training is learning the weights (millions or billions of parameters) that control the model's behavior, vs inference is "running" the trained model on user data.
- iamaaditya 4y agoIn machine learning (especially deep learning or neural networks), the 'training' is done by using Stochastic Gradient Descent. These gradients are computed using Backpropagation. Backpropagation requires you to do a backward pass of your model (typically many layers of neural weights) and thus requires you to keep in memory a lot of intermediate values (called activations). However, if you are doing "inference" that is if the goal is only to get the result but not improve the model, then you don't have to do the backpropagation and thus you don't need to store/save the intermediate values. As the layers and number of parameters in Deep Learning grows, this difference in computation in training vs inference becomes signifiant. In most modern applications of ML, you train once but infer many times, and thus it makes sense to have specialized hardware that is optimized for "inference" at the cost of its inability to do "training".
- Kalanos 4y agoAnyone care to comment on how this is better than Metal's TensorFlow support?
- buildbot 4y agoThis is very interesting since the M1 studio supports 128GB of unified memory - training a large memory heavy model slowly on a single device could be interesting, or inferencing a very large model.
- zdw 4y agoEverything old is new again - the M1 studio's unified memory echos the SGI O2 which had similar unified CPU/GPU memory back in the 90's. In both cases the unified memory machines outperformed much larger machines in specific use cases.
- smoldesu 4y ago...specific use cases being the key operand here. Unified memory is cool, but there are reasons we don't use it at-scale: - It needs extremely high-bandwidth controllers, which severely limits the amount of memory you can use (Intel Macs could be configured with an order of magnitude more ram in it's server chips) - ECC is still off-the-table on M1 apparently - Most workloads aren't really constrained by memory access in modern programs/kernels/compilers. Problems only show up when you want to run a GPU off the same memory, which is what these new Macs account for. - Most of the so-called "specific workloads" that you're outlining aren't very general applications. So far I've only seen ARM outrun x86 in some low-precision physics demos, which is... fine, I guess? I still don't foresee meteorologists dropping their Intel rigs to buy a Mac Studio anytime soon.
- my123 4y ago> - It needs extremely high-bandwidth controllers, which severely limits the amount of memory you can use (Intel Macs could be configured with an order of magnitude more ram in it's server chips) In the first half of 2023, NVIDIA Grace Superchip will ship with an 1TB memory config (930GB usable because ECC bits) on a 1024-bit wide LPDDR5X-8533 config (same width as M1 Ultra, with LPDDR5-6400). So it's going to become much less of an issue really soon.
- 4y ago
- mkaic 4y agoThis is really cool for a number of reasons: 1.) Apple Silicon currently can't compete with Nvidia GPUs in terms of raw compute power, but they're already way ahead on energy efficiency. Training a small deep learning model on battery power on a laptop could actually be a thing now. Edit: I've been informed that for matrix math, Apple Silicon isn't actually ahead in efficiency 2.) Apple Silicon probably will compete directly with Nvidia GPUs in the near future in terms of raw compute power in future generations of products like the Mac Studio and Mac Pro, which is very exciting. Competition in this space is incredibly good for consumers. 3.) At $4800, an M1 Ultra Mac Studio appears to be far and away the cheapest machine you can buy with 128GB of GPU memory. With proper PyTorch support, we'll actually be able to use this memory for training big models or using big batch sizes. For the kind of DL work I do where dataloading is much more of a bottleneck than actual raw compute power, Mac Studio is now looking very enticing.
- smoldesu 4y agoThere's definitely competition, and it's going to be really interesting to watch Nvidia and Apple duke it out over the next few years: - Apple undoubtedly owns the densest nodes, and will fight TSMC tooth-and-nail over first dibs on whatever silicon they have coming next. - Apple's current GPU design philosophy relies on horizontally scaling the tech they already use, whereas Nvidia has been scaling vertically, albeit slowly. - Nvidia has insane engineers. Despite the fact they're using silicon that's more than twice as large by-area when compared to Apple, they're still doubling their numbers across the board. And that's their last-gen tech too, the comparison once they're on 5nm later this summer is going to be insane. I expect things to be very heated by the end of this year, with new Nvidia, Intel and potentially new Apple GPUs.
- my123 4y ago> but they're already way ahead on energy efficiency 1) Nope. For neural network training not the case: https://tlkh.dev/benchmarking-the-apple-m1-max https://tlkh.dev/benchmarking-the-apple-m1-max And that's with the 3090 set at a very high 400W power limit, can get far more efficient when clocked lower. (which is normal, because no dedicated matrix math accelerators on the GPU notably) 2) We'll see, hopefully Apple thinks that the market is worth bothering with... (which would be great) 3) Indeed, if you need a giant pool of VRAM above everything else at a relatively low price tag, Apple is indeed a quite enticing option. If you can stand Metal for your use case of course.
- ivstitia 4y agoThere was a report comparing M1 Pro with several other Nvidia GPUs from a few months ago: https://wandb.ai/tcapelle/apple_m1_pro/reports/Deep-Learning-on-the-M1-Pro-with-Apple-Silicon---VmlldzoxMjQ0NjY3 https://wandb.ai/tcapelle/apple_m1_pro/reports/Deep-Learning... I'm curious on how the benchmarks change with this recent new release!
- munro 4y agoyess! This is important for me, because I don't have any $$$ to rent GPUs for personal projects. Now we just need M1 support for JAX. Since there are no hard benchmarks against other GPUs, here's a Geekbench against an RTX 3080 Mobile laptop I have [1]. Looks like it's about 2x slower--the RTX laptop absolutely rips for gaming, I love it. [1] https://browser.geekbench.com/v5/compute/compare/4140651?baseline=4529092 https://browser.geekbench.com/v5/compute/compare/4140651?bas...
- jph00 4y agoYou can use GPUs for free on Paperspace Gradient, Google Colab, and Kaggle.
- in3d 4y agoIt’s surprising to see PyTorch developers working on things like that when common operations like group convolutions are still completely unoptimized on Nvidia GPUs, despite many requests.
- jacobn 4y agoGrouped convolutions can't really run faster than groups * conv(ch/group) and I believe that's close to where they're at? Note that for ch<O(512) (varies by GPU & hw) you tend to be memory-transfer-speed limited, not compute limited. So unfortunately depthwise convolutions end up having terrible performance.
- in3d 4y agoWhy wouldn’t you be able to run them in parallel using CUDA? You shouldn’t be memory-transfer speed limited when group convolution layers are a part of a bigger net. Note that pointwise 1x1 convolutions are a special case of group convolutions and actually I think they might be specially optimized in PyTorch (I’d have to run some benchmarks to test it though).
- brrrrrm 4y agopointwise isn’t a case of grouped conv, they’re orthogonal ideas. You can fuse grouped convs (depthwise is a special case of grouped convs) into preceding or following layers. Maybe JAX can do this already? No clue if any library offers such an optimization out of the box
- in3d 4y agoSorry, yes, I was replying to the post about depthwise convolution and that’s what I meant (though the naming of it is poor) - i.e. the special case of group convolutions where the number of groups is equal to the number of channels.
- macshome 4y agoDoes this work on any Metal hardware or just the M1 GPU?
- atty 4y agoThis is targeting AMD GPUs and M1 GPUs currently not targeting the integrated Intel GPUs present in Intel machines. However if you have a 16 inch Intel MBP, or a Mac Pro, etc, this should work with your AMD GPUs. That support isn’t in the nightly packages yet (only Apple Silicon support so far) but the PyTorch team is saying that it will be available by the end of the week hopefully. If you just can’t wait, you should be able to build from source to test it out right now.
- almostdigital 4y agoAnyone actually got this to run on an M1 Mac? $ conda install pytorch torchvision torchaudio -c pytorch-nightly Collecting package metadata (current_repodata.json): done Solving environment: failed with initial frozen solve. Retrying with flexible solve. Collecting package metadata (repodata.json): done Solving environment: failed with initial frozen solve. Retrying with flexible solve. PackagesNotFoundError: The following packages are not available from current channels: - torchaudio And the pip install variant installs an old version of torchaudio that is broken OSError: dlopen(/opt/homebrew/Caskroom/miniforge/base/envs/test123/lib/python3.10/site-packages/torchaudio/lib/libtorchaudio.so, 0x0006): Symbol not found: __ZN2at14RecordFunctionC1ENS_11RecordScopeEb
- boopmaster 4y agopip install --pre torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/nightly/cpu https://download.pytorch.org/whl/nightly/cpu
- fragmede 4y agopip3 install pytorch worked for me. I think it's something with your brew installation. fragmede@samairmac:~$ python Python 3.9.7 | packaged by conda-forge | (default, Sep 29 2021, 19:24:02) [Clang 11.1.0 ] on darwin Type "help", "copyright", "credits" or "license" for more information. >>> import torch >>> torch.__file__ '/Users/fragmede/projects/miniforge3/lib/python3.9/sitepackages/torch/__init__.py' >>>
- almostdigital 4y agoDoes torchaudio work for you? I can get torch and torchvision to work but not torchaudio
- deleted 4y ago[deleted]
- kristianp 4y agoA tangential thought: will we see animation studios buy mac studios for their rendering farms? What do they use these days, aws ec2?
- singularity2001 4y agoThe installation command generated on https://pytorch.org/get-started/locally/ https://pytorch.org/get-started/locally/ didn't install the latest version for me. What did it was: pip3 install --pre torch==1.12.0.dev20220518 --extra-index-url https://download.pytorch.org/whl/nightly/cpu https://download.pytorch.org/whl/nightly/cpu
- singularity2001 4y agoIf you came late make sure to update the date to 20220521 …
- tzekid 4y agoAhh just saw this after compiling pytorch from source. Thanks!