16 ms·
Apple AMX instruction set (M1/M2 matrix coprocessor)
- 16890c 4y agoSo what software actually uses this / what compiler supports this if it's neither supported nor documented by Apple?
- kennywinker 4y agoPresumably first party software that qualifies. I’m thinking siri, or their AR stuff. It’s possible it’s also used by something like CoreML where there is a public facing framework that utilizes this under the hood. I have no special info, these are just guesses.
- garblegarble 4y agoApplications use it via a higher level interface: Apple's Accelerate framework[1] 1: https://developer.apple.com/accelerate/ https://developer.apple.com/accelerate/
- peaslock 4y agoSpeaking of which, is anyone aware of example code using the LSTM? I've been trying to get this to work, but there seems to be information missing e.g how to setup the input/output descriptors and how to manage input data: https://developer.apple.com/documentation/accelerate/bnns/using_long_short-term_memory_layers_lstm https://developer.apple.com/documentation/accelerate/bnns/us...
- Someone 4y agoIt’s just like with private APIs. It’s functionality that they don’t know (¿yet?) whether they want to support it in the future, but they do provide calls that use them indirectly. There also could be bugs in the hardware that they carefully programmed around.
- scarface74 4y agoWhy do people act as if “private APIs” (which is somewhat of a contradiction in terms) is nefarious? An API is something that is publicly documented that the vendor promises to support where the behavior won’t change. Raymond Chen has been blogging for close two decades about all of the hacks that MS had to put into Windows because vendors used undocumented APIs.
- matthewmacleod 4y agoI agree that people tend to get a bit too worked up about it, but I don’t think it’s a contradiction in terms as such - an API is really just any interface where two distinct pieces of software interact in some way. It doesn’t need to be formally described or published or anything like that, and the idea of a private API is pretty common generally. Where it starts to piss people off though is where those private APIs are used to allow first-party software access to platform features that third-party software doesn’t get. For some applications that’s not so bad, but in the case of a general-purpose operating system platform or similar it’s kind of an anticompetitive move and we should complain when companies do it.
- scarface74 4y agoOnce you make an API public, no matter how badly designed it is, you have to support it forever. I would much rather an API be private, let the company dog food it and let their internal employees use it, and then make it public. It also gives them the freedom of completely changing the internal workings. The extensions API and the Siri integration for third parties are great examples. The Siri intent based API is very usable and reminds of the Amazon Lex based API - the AWS version of the consumer Alexa skills SDK.
- thfuran 4y ago>Once you make an API public, no matter how badly designed it is, you have to support it forever That's simply not true.
- 4y ago
- michelb 4y agoHere are the docs: https://developer.apple.com/documentation/accelerate https://developer.apple.com/documentation/accelerate
- zeristor 4y agoIs this a repost the other day, I thought that was too good have been missed out. Also I’m keen to see if this 60Gb.a-1 near field wireless data link for Apple Watches for diagnosis will be able to be used in some sort of MagSafe/usb for iPhones.
- altairprime 4y ago(This isn’t a post about near-field, Watches, MagSafe, or USB; perhaps this comment was meant for another post?)
- deleted 4y ago[deleted]
- saboot 4y agoIs there a comparison with other fast cpu methods of matrix multiplication?
- moonchild 4y agoIIRC I measured something like 10-50% performance difference (don't remember exactly, but it was somewhere in there), vs a reasonably well-regarded blas implementation. This was for dgemm specifically; I don't know if the story changes for smaller floats.
- saboot 4y agoNot too bad, I do scientific computing and choose Intel/Nvidia as their APIs for accelerated math operations are documented and supported for developers. I've been paying attention to what Apple has been pushing with their M1/M2 chips, and I'm pretty tempted to try it out, but unless these features are documented and supported I can't feel comfortable writing programs relying on them.
- mhh__ 4y agoThe API apple wants you to use is documented and presumably is here to stay. That doesn't help if there's some edge case you'd need access to the raw ISA but still.
- Someone 4y agoAlso, BLAS is part of that interface (https://developer.apple.com/documentation/accelerate/blas https://developer.apple.com/documentation/accelerate/blas) Of course it is a black box in that you can’t (realistically) try and speed it up. You still run the risk of Apple’s priorities being different from yours.
- kergonath 4y agoOTOH the interfaces are standard. Using Accelerate instead of the standard BLAS is just a compiler switch.
- sposeray 4y ago[dead]
- stephc_int13 4y agoKnowing the history of Apple an open standard and given their success with their implementation of the ARM64 ISA, it is unfortunately highly probable that they will follow the proprietary route once again. Indeed, they are already doing it, we're lucky they weren't in a dominant position when TCP/IP or HTML were invented.
- Reason077 4y ago> "we're lucky they weren't in a dominant position when TCP/IP or HTML were invented" TCP/IP and HTML became dominant because they were open standards. Should they have been proprietary, they would have floundered and something else would have emerged instead.
- rjzzleep 4y agoI guess people have forgotten, that Webkit came out of KHTML from the KDE team and that Apple was a nightmare when it came to contributing code back. They just released a huge dump of the whole thing.
- klodolph 4y agoI remember this... it was just a fork. Projects get forked. It's unfortunate from some perspectives, but from other perspectives you can understand why forks happen. When you have a long-running fork, especially one that is so active, merging it naturally becomes a nightmare. This is expected and ordinary. The Linux kernel gets forked by Android vendors and others all the time. A lot of the changes never make it upstream, for various reasons. At least the story ends a bit better for KHTML / WebKit.
- rjzzleep 4y agoEvery single Apple patch to GitHub projects is done by the same single indistinguishable user account. This isn't just "some long-running fork". It is Apple culture to actively prohibit contributions to open source projects unless 5 managers sign off on it.
- scarface74 4y agoWhat is the “non proprietary” alternative and why should Apple or it’s users be forced to wait on consensus? This is the same reason that Apple wasn’t saddled with the horrible PC “standards” before USB became ubiquitous. Not to mention even today, Bluetooth is a shit show outside of the Apple ecosystem as far as handoff an ease of pairing.
- viraptor 4y agoOne of the non-extreme solutions is publishing the doc for the extension + explicit usage grant on any related patents. They don't need to go full standardisation route before the first release. It would still be a proprietary extension under their control, but not a haha-screw-you proprietary.
- samwillis 4y agoApple are religious about interoperability within their platform, making it easy and reliable. If they were to do as you suggested with some of their proprietary tech, there will be products that implement it badly. To the user they would have no idea who’s at fault, and would probably blame the tech in general, damaging Apple. Standardisation, in combination with certification to use the “label”, ensures that people developing on top of their innovation do so well enough that it doesn’t damage the brand. (Somewhat less relevant to an instruction set, and not something I particularly agree with)
- viraptor 4y agoThis doesn't really make sense. There are already systems implementing connections to Apple stuff badly due to lack of documentation. The situation would only improve with the publication. For example we already have most of m1 hardware reverse engineered in Asahi - that's not going away. We've had things like air drop and earbuds charge state RE'd too. We'll get amx libraries as well soon.
- scarface74 4y ago
- coder543 4y agoDoesn't M2 add support for SVE2 as well?
- aseipp 4y agoNo, there are no publicly available chips with SVE2 support at all; the closest you can get is SVE1 on Graviton 3 from AWS.
- deleted 4y ago[deleted]
- sufiyan 4y agoI don’t get why other chip manufacturers don’t go this same route. For example AVX is done on the same core that also supports integer math. Many companies have a separate GPU but AVX seems to always come prepackaged.
- guipsp 4y agoThere is a cost to moving data to another chip(let). I think this is the main reason
- viktorcode 4y agoQualcomm most likely will go this route (proprietary instructions) as this now officially blessed by Arm. I'd bet we'll see that in their 2023 chips.
- ndesaulniers 4y agoIsn't ARM suing Qualcomm?
- ladyanita22 4y agoWhich compiler would be required for this (https://github.com/corsix/amx/blob/main/aarch64.h https://github.com/corsix/amx/blob/main/aarch64.h)? I understand the limitation is not at the OS side, as nothing can be done there, but at the compiler-side (I mean that the Apple-supplied compiler doesn't compile against the AMX instruction set, so you'd need a compatible one that, I understand, doesn't exist). Or is it just undocumented and you can actually get it to work with a Standard xcode and macOS installation given the headers provided?
- dougall 4y agoThis header works with standard Xcode/macOS, by taking advantage of inline assembly in a slightly-cursed way (turning register names into numbers and encoding the instruction itself).
- corsix 4y agoSlightly cursed is how I roll. I would like to know where the trick first originated from though (I found it at https://github.com/yvt/amx-rs/blob/main/src/nativeops.rs#L22 https://github.com/yvt/amx-rs/blob/main/src/nativeops.rs#L22 rather than inventing it de novo)
- ndesaulniers 4y agoThis happens often in the Linux kernel to continue to support older assemblers for newer instruction set extensions. The x86 retbleed mitigation uses .inst to trick the hardware instruction decoder...different instructions are encoded/run than what is speculatively decoded.
- dagmx 4y agoCompilers wouldn’t target it. You’d get to it via https://developer.apple.com/documentation/accelerate https://developer.apple.com/documentation/accelerate
- dan-robertson 4y agoMaybe I’m just slow but it wasn’t immediately obvious to me how to use this for matrix multiplication. Let me now try to explain. Suppose we have some matrices we would like to multiply, a_ok and b_ij (and let’s say their sizes line up with the hardware because I think those details aren’t so relevant). Their product is c_ik = a_ij b_jk = sum(a_ij * b_jk for all j). The hardware lets us cheaply compute and accumulate an outer product (see picture in OP): r_ij = r’_ij + p_i * q_j Now start with r = 0 and accumulate: r_ik = a_i1 * b1k + a_i2 * b2k + ... + a_in * b_nk = c_ik Each row corresponds to one AMX op on all the cells of the matrix. Writing it out like this it seems quite straightforward. I think I was caught up on thinking about the per-cell computation too much. When computing based on cells in the output, you take a row from the left hand side and dot it with a column from the right hand side (nn dot products). Here, we take a column* from the left hand side and a row from the left hand side and outer product them (n outer products) and add up the result. Perhaps this is partly a victory for this kind of symbolic index notation. I think this would all be much less obvious if I wrote it all out as a sum of outer products with eg the tensor product symbol.
- magwa101 4y ago