5 ms·
Skimmed about 2/3s of this. It's unclear to me if code needs to be specially written to run on the NPU. Normally I think of something like CUDA to get code run
by patrickthebold 2y ago
Skimmed about 2/3s of this. It's unclear to me if code needs to be specially written to run on the NPU.
Normally I think of something like CUDA to get code running on a GPU. Can code targeting a CPU or GPU automatically be sped up running on a NPU? Or does code explicitly need to target the NPU?
- lawlessone 2y agoI agree I don't understand the difference? Are the calculations an NPU is capable of doing different to a GPU? Are they not basically identical hardware?
- klelatti 2y agoNPUs are basically specialized for matrix multiplication, GPUs for more general parallel operations on multiple data (Single Instruction Multiple Thread) although modern GPUs also contain matrix multiplication units. May be a degree of software compatibility at the highest level - eg PyTorch - but the underlying software will be very different.
- adrian_b 2y agoA NPU can do only a very small subset of the operations supported by a GPU. A NPU does strictly only the operations required for ML inference, which use data types with low precision, i.e. 16-bit or 8-bit types.
- foobiekr 2y agoDifferent optimization choices.
- xcv123 2y agoBy the same reasoning, a CPU is no different to a GPU. They can both do matrix calculations. A GPU is optimised for 3D rendering (and is useful for parallel computations in general). An NPU is optimised for neural network inferencing. These algorithms both involve matrix mathematics but they are not the same. The NPU hardware design matches the deep neural network inferencing algorithm. For example it has an "Activation Function" block dedicated to computing the activation function between neural network layers. It is optimised and specialised for one very specific algorithm: inferencing. A GPU would beat an NPU for training, and any other parallel computations besides inferencing.
- TheDudeMan 2y agoHigh-end Nvidia GPUs are not optimized for 3D rendering; they are optimized for machine learning and inference.
- taneq 2y agoThey're not really GPUs then, are they? Even if they're capable of generating a video output.
- Dalewyn 2y agoThey are GPUs, but standing for General Processing Unit. "GPU" the acronym has persisted from sheer force of tradition and habit, but what it means has changed drastically over the past several years between cryptocurrency and now "AI". I wonder if NPU will supercede GPU as in General Processing Unit now that it has finally entered the wider lexicon, relegating GPU back to Graphics Processing Unit or video cards. And no, GPGPU (General Purpose Graphics Processing Unit) is a bloody stupid term to be bluntly honest.
- taneq 2y agoThat might be how you use the term but GPU is still almost universally used to mean graphics processing unit. If it want to use a ubiquitous acronym to mean something different, you need to define it first.
- Dalewyn 2y ago>GPU is still almost universally used to mean graphics processing unit. No, GPU is almost universally used to mean GPU. There is nothing graphical about cryptocurrency, "AI" (sans image generation), protein crunching, and whatever else they are being used for that aren't graphical. I question how many people are even aware the G is supposed to stand for Graphics anymore. The nomenclature is outdated and doesn't reflect reality anymore.
- Const-me 2y ago> Are they not basically identical hardware? For example Apple’s NPU can’t do FP32 precision, it can only do FP16 and less.
- eterps 2y agoI'm also wondering if such an NPU can be targeted (in a way that can be understood) from the assembly/machine language level. Or that it needs an opaque kitchen sink of libraries, blobs and other abstractions.
- klelatti 2y agoOpaque blobs I believe in almost all cases!
- qludes 2y agoI believe it's the latter, each NPU vendor has their own software stack. Take a look at Tomeu Vizoso's work: https://blog.tomeuvizoso.net/search/label/npu https://blog.tomeuvizoso.net/search/label/npu
- wmf 2y agoYou can write assembly for NPUs although the instruction set may be quite janky compared to CPUs. Once you've written NPU code you need some libraries to load your code but that's not particularly different from the fact that CPUs now need massive firmware to boot. Back in reality, that's not how any vendor intends their NPUs to be used. They provide high-level libraries that implement ONNX, CoreML, DirectML, or whatever and they expect you to just use those.
- deleted 2y ago[deleted]
- cherioo 2y agoMy understanding is that the low level APIs are different, CUDA vs. whatever Android provides vs. whatever Apple provides. However, higher abstraction like PyTorch may be able to target different platforms with less code changes.
- fulafel 2y agoCuda is of course NVidia-proprietary. Will be interesting to see if other high level accelerator supporting languages like Chapel or Futhark or JAX end up getting NPU backends, it might give them a nice boost over the proprietary C++ inspired language. Edit: JAX has TPU support.
- fisf 2y agoSupport in for high level backends almost always comes down to c++ either way. That goes for NPUs (which are really fragmented anyway, i.e. there is no uniform API), and and also for JAX' TPU backend (which iirc is using XLA).
- deleted 2y ago[deleted]
- fulafel 2y agoC++ stings less as an API used in the low level machinery under the hood as long as you as an application author don't have to write code in it. I haven't done an in depth look but most matrix math accelerators (eg AMD, Intel and Apple) seem to provide C/C++/Python APIs for describing the computations but the code executing on the NPU is not compiled from user C++ code. Apparently eg in Intel's stuff there's a custom run-time compiler consuming this kind of IR (intermediate representation) in the accelerator sw stack: https://docs.openvino.ai/2022.3/openvino_ir.html https://docs.openvino.ai/2022.3/openvino_ir.html & https://github.com/intel/linux-npu-driver https://github.com/intel/linux-npu-driver And on AMD from user POV it doesn't seem too different: https://ryzenai.docs.amd.com/en/latest/devflow.html https://ryzenai.docs.amd.com/en/latest/devflow.html
- fisf 2y agoYes, but this is not where I am getting at. "Will be interesting to see if other high level accelerator supporting languages like Chapel or Futhark or JAX end up getting NPU backends, it might give them a nice boost over the proprietary C++ inspired language." As you say, the GPU (or NPU, TPU,..) don't run C++ or anything derived from it. The "runtime" (~backend) will usually emit some kind of hardware dependent format (or again, an IR) like SPIR-V, PTX, etc. But the backend itself is usually written in C++ (due to performance reasons), and there is really no way to get around that. Interacting with that from Python (or Jax) is a usability win, but there is zero difference in functionality. I.e. there is no proprietary C++ inspired language in play here. Hence no way to get a boost.
- ribit 2y agoMost NPUs are not directly end-user programmable. The vendor usually provides a custom SDK that allows you to run models created with popular frameworks on their NPUs. Apple is a good example since they have been doing it for a while. They provide a framework called CoreML and tools for converting ML models from frameworks such as PyTorch into a proprietary format that CoreML can work with. The main reason for this lack of direct programmability is that NPUs are fast-evolving, optimized technology. Hiding the low-level interface allows the designer to change the hardware implementation without affecting end-user software. For example, some NPUs can only work with specific data formats or layer types. Early NPUs were very simple convolution engines based on DSPs; newer designs also have built-in support for common activation functions, normalization, and quantization. Maybe one day, these things will mature enough to have a standard programming interface. I am skeptical about this becoming a reality any time soon. Some companies (like Tenstorrent) are specifically working on open architectures that will be directly programmable, I'm not sure whether their approach translates to the embedded NPUs, though. What would be nice is an open graph-based API and a model format for specifying and encoding ML models.