13 ms·
Curiously neither PyTorch nor Tensorflow currently use M1's Neural Engine. Is too limited? Too hard to interact with? Not worth the effort?
by lekevicius 4y ago
Curiously neither PyTorch nor Tensorflow currently use M1's Neural Engine. Is too limited? Too hard to interact with? Not worth the effort?
- deleted 4y ago[deleted]
- RicoElectrico 4y agoMost probably Neural Engine is optimized for inference, not training.
- sillyinseattle 4y agoQuestion about terminology (no background in AI). In econometrics, estimation is model fitting (training, I guess), and inference refers to hypothesis testing (e.g. t or F tests). What does inference mean here?
- upwardbound 4y agoInference here means "running" the model. So maybe it has a similar meaning as in econometrics? Training is learning the weights (millions or billions of parameters) that control the model's behavior, vs inference is "running" the trained model on user data.
- iamaaditya 4y agoIn machine learning (especially deep learning or neural networks), the 'training' is done by using Stochastic Gradient Descent. These gradients are computed using Backpropagation. Backpropagation requires you to do a backward pass of your model (typically many layers of neural weights) and thus requires you to keep in memory a lot of intermediate values (called activations). However, if you are doing "inference" that is if the goal is only to get the result but not improve the model, then you don't have to do the backpropagation and thus you don't need to store/save the intermediate values. As the layers and number of parameters in Deep Learning grows, this difference in computation in training vs inference becomes signifiant. In most modern applications of ML, you train once but infer many times, and thus it makes sense to have specialized hardware that is optimized for "inference" at the cost of its inability to do "training".
- eklitzke 4y agoJust to add to this, the reason these inference accelerators have become big recently (see also the "neural core" in Pixel phones) is because they help doing inference tasks in real time (lower model latency) with better power usage than a GPU. As a concrete example, on a camera you might want to run a facial detector so the camera can automatically adjust its focus when it sees a human face. Or you might want a person detector that can detect the outline of the person in the shot, so that you can blur/change their background in something like a Zoom call. All of these applications are going to work better if you can run your model at, say, 60 HZ instead of 20 HZ. Optimizing hardware to do inference tasks like this as fast as possible with the least possible power usage it pretty different from optimizing for all the things a GPU needs to do, so you might end up with hardware that has both and uses them for different tasks.
- sillyinseattle 4y agoThank you @iamaaditya and @eklitzke . Very informative
- dataexporter 4y agoThis sounds really fascinating. Are there any resources that you'd recommend for someone who's starting out in learning all this? I'm a complete beginner when it comes to Machine Learning.
- dr_zoidberg 4y agoDeep Learning with Python (2nd ed), by Francois Chollet. If you don't mind about learning the part where you program, it's got a lot of beginner/intermediate concepts clearly explained. If you do dive into the programming examples, you get to play around with a few architectures and ideas and you're left on the step to dive into the more advanced material knowing what you're doing.
- dekhn 4y agoit took me 20 years to learn this body of knowledge and now it can just sort of be summed up in a paragraph. When I learned and used gradient descent, you had to analytically determine your own gradients (https://web.archive.org/web/20161028022707/https://genomics.soe.ucsc.edu/sites/default/files/stormo94.pdf https://web.archive.org/web/20161028022707/https://genomics....). I went to grad school to learn how to determine my own gradients. Unfortunately, in my realm, loss landscapes have multiple minima, and gradient descent just gets trapped in local minima.
- Q6T46nT668w6i3m 4y agoI’m surprised nobody has provided the basic explanation: inference, here, means matrix, matrix or matrix, scalar multiplication.
- malshe 4y agoI have background in both and it's very confusing to me. Inference in DL is running a trained model to predict/classify. Inference in stats and econometrics is totally different as you noted.
- abm53 4y agoIt is confusing that the ML community have come to use "inference" to mean prediction, whereas statisticians have long used it to refer to training/fitting, or hypothesis testing. I'm not sure when or why this started.
- mattkrause 4y agoPrediction. The model is literally "inferring" something about its inputs: e.g., these pixels denote a hot dog, those don't.
- munro 4y agoThat /sounds/ right, but training still has a forward part, so OP does raise a really great question. And looking at the silicon, the neural engine is almost the size of the GPU. Really need someone educated in this area to chime in :)
- my123 4y agoThe neural engine is only exposed through a CoreML inference API. You can't even poke the ANE hardware directly from a regular process. The interface for accessing the neural engine is not hardened (you can easily crash the machine from it). So the matter is essentially moot in practice as you'd need your users to run with SIP off...
- munro 4y agoSounds like you've you done a bit of digging around, you're efforts are appreciated. I found and a github of people sharing what they know, here's a guy live streaming hacking it and building a tinygrad https://youtu.be/mwmke957ki4 https://youtu.be/mwmke957ki4
- viraptor 4y agoThat doesn't seem to be a huge issue. If someone actually does this for income, would they avoid disabling sip for 2x performance gain for example?
- dgacmu 4y agoYou have to stash more information from the forward pass in order to calculate the gradients during backprop. You can't just naively use an inference accelerator as part of training - inference-only gets to discard intermediate activations immediately. (Also, many inference accelerators use lower precision than you do when training) There are tricks you can do to use inference to accelerate training, such as one we developed to focus on likely-poorly-performing examples: https://arxiv.org/abs/1910.00762 https://arxiv.org/abs/1910.00762
- why_only_15 4y agoThe ANE only has support for calculations with fp16, int16 and int8 all of which are too small to train with (too much instability). A common thing to do is train in fp32 to be able to get the small differences and gradients and then once the model is frozen do inference on fp16 or bf16.
- jph00 4y agoUsing mixed precision training you can do most operations in fp16 and just a few in fp32 where it's needed. This is the norm for NVIDIA GPU training nowadays. For instance using fastai add `.to_fp16()` after your learner call, and that happens automatically.
- omegalulw 4y agoHow is the choice between fp16 and fp32 made? Is it like if any gradients in the tensor need the extra range you use fp32?
- h-jones 4y agoThe PyTorch docs give a pretty good overview of AMP here https://pytorch.org/tutorials/recipes/recipes/amp_recipe.html https://pytorch.org/tutorials/recipes/recipes/amp_recipe.htm... and an overview of which operations cast to which dtype can be found here https://pytorch.org/docs/stable/amp.html#autocast-op-reference https://pytorch.org/docs/stable/amp.html#autocast-op-referen.... Edit: Fixed second link.
- andoma 4y agoThis article [0] from Nvidia gives a good overview of how mixed precision training works. Super high level (from section 3): 1. Converting the model to use the float16 data type where possible. 2. Keeping float32 master weights to accumulate per-iteration weight updates. 3. Using loss scaling to preserve small gradient values. [0] https://docs.nvidia.com/deeplearning/performance/mixed-precision-training/index.html https://docs.nvidia.com/deeplearning/performance/mixed-preci...