5 ms·
I’m thinking this is related to the machine-learning coprocessor. They likely expose this to developers through their ML API and don’t want to have to support t
by mlazos 6y ago
I’m thinking this is related to the machine-learning coprocessor. They likely expose this to developers through their ML API and don’t want to have to support this for external customers if it changes. Still this is a great find!
- londons_explore 6y agoOther vendors ML hardware is very much "A far away device over the PCI express bus which you give bunch of work to and later come back and see if its done". Looks like apple has decided to make it very tightly integrated into the CPU. It means while doing ML operations, the CPU core can't go off and do other useful work. It also forces all cores to require matching ML hardware (or have a lot of OS complexity as certain threads can only run on some cores). The benefit is latency to get ML stuff done is lowered from microseconds to nanoseconds. I suspect Apple probably messed up with that tradeoff... Very few applications can't wait a few microseconds for results of some ML computation...
- tgtweak 6y agoIt is unlikely that this is related to the much more purpose-built and conventionally-designed "neural engine" in the M1 (as they are dubbing it) which operates in the "set it up, feed it, and check the results" type of pipeline you allude to above. These AMX instructions are doing similar matrix operations (tile-based) but not at the same scale as dedicated hardware. Think of it like L1 cache for neural-nets - if you can fit your model into that on-die tile, it can happen there and be much quicker (and also leverage data already in cache or memory) than dispatching it to a separate device which is orders of magnitude higher latency. The tradeoff is not likely perceptible to developers since the compiler or runtime will decide if the net is small enough to run locally or whether it should be dispatched to the dedicated device (again, like L1 access for the majority of developers not writing inline assembly) - thus why it's not exposed. They (apple) could also do some very interesting things with it internally such as neural branch prediction and cache eviction based on contextual operations - again outside of the scope of what a developer would have access to. Intel is doing the same in their upcoming x86 silicon with Intel Advanced Matrix Extension - I would expect this ISA to be heavily inspired by that.
- Q6T46nT668w6i3m 6y agoThe ML device is pretty much just a MATMUL processor.
- texse 6y agoThese instructions are not for the Neural Engine. That is a separate hardware block outside of the CPUs. Apple refers to this feature as "AMX" in their marketing documentation. While AMX could be used for deep learning in a pinch despite the lack of support for common formats like fp16 and int8, I suspect Apple had some other use cases in mind as well. For example, 64-bit float support is expensive and generally useless for ML. However, they are useful (though not necessarily required) in problems such as bundle adjustment that may appear in the context of frameworks like ARKit.