7 ms·
They have not adopted ARMv9. This is still ARMv8, but with SME.
by ribit 2y ago
They have not adopted ARMv9. This is still ARMv8, but with SME.
- axoltl 2y agoYep, the binaries are all arm64e.
- saagarjha 2y agoThis doesn’t really say much
- hajile 2y agoARMv9.0 is very similar to ARMv8.5 (9.0 supersets 8.5 with SVE2, TME, TLA, and CCA), so it's not a massive deal. SME implies v8.7 which is basically identical to v9.2 except for those couple extensions previously mentioned. I wonder if there is licensing at play though. Apple may have gotten a really great licensing deal on ARMv8 that they wouldn't be offered for ARMv9.
- ribit 2y agoMy guess is that Apple is simply not interested in some of the ARMv9 features. They are not eager to implement SVE and the se Ure virtualization features are probably not that relevant to them.
- skavi 2y agoDoes anyone have insight into why arm CPU vendors seem so hesitant about implementing SVE2? ~They seem~ *Apple seems to have no issue with SSVE2 or SME. Edit: Only Apple has implemented SSVE and SME I think.
- hajile 2y agoSVE2 is an extension on top of SVE which some stuff already implements. The issue is more likely to be the politics of moving to ARMv9 than anything else. As to SVE though, I'd guess variable execution time makes the implementation require a bit of work. Normally, multi-cycle tasks have a fixed number. Your scheduler knows that MUL takes N cycles and plans accordingly. SVE seems like it should require N-M cycles depending on what is passed. That must be determined and scheduled around. This would affect the OoO parts of the core all the way from ordering through to the end of the pipeline. That's definitely bordering on new uarch territory and if that is the case, it would take 4-5 years from start to finish to implement. This would explain why all the ARMv8 guys never got around to it. ARMv9 makes it mandatory, but that was released in 2021 or so which means non-ARM implementors probably have a ways to go.
- skavi 2y agoThis isn’t a convincing explanation to me. There are plenty of variable latency instructions on existing high performance arm64 cores.
- dzaima 2y agoSVE doesn't need variable-execution-time instructions, outside of perhaps masked load/store, but those are already non-constant. Everything else is just traditional instructions (given that, from the perspective of the hardware, it has a fixed vector size), with a blend.
- anticensor 2y agoVariable execution time instructions can always be divided into smaller fixed execution time microinstructions.
- ribit 2y agoI am curious, which SVE instructions imply variable execution time? I’d guess that first fault load could be tricky to implement…
- ribit 2y agoWhat do you mean? Apple is the only one who has an SME/SSVE implementation.
- skavi 2y agoI misremembered. Looks like it is only Apple. I appreciate the correction.
- brigade 2y agoWhat is the measurable benefit to implementing 128b SVE2? Like, ARM has CPUs that implement that, and it's not even disabled on some chips. So there must be benchmarks somewhere showing how worthwhile it is. And implementing 256b SVE has different issues depending on how you do it. 4x256b vector ALUs are more power hungry than generally useful. 2x256b is only beneficial over 4x128b if you're limited by decode width, which isn't an issue now that A32/T32 support has been dropped. 3x256b would probably imply 3x128b which would regress existing NEON code. And little cores don't really want to double the transistors spent on vector code, but you can't have a different vector length than the big cores...
- skavi 2y agoMasked instructions primarily. But apart from that it’s just a more complete ISA vs NEON. More comparable to AVX512/AVX10. > 2x256b is only beneficial over 4x128b if you're limited by decode width This is only true if we ignore more complex instructions and focus on things like adding two vectors.
- brigade 2y agoWhat is the percentage gain of using masked instructions on any benchmark/task of your choice? It can be negative on weird kernels that do lots of vector cmp since even ARM decided the cost of more than one write port in the predicate register file wasn't worth it, or if the masking adds lots of unnecessary and possibly false dependencies on the destination registers. > This is only true if we ignore more complex instructions and focus on things like adding two vectors. ARM implemented a CPU that had 2x256b SVE and 4x128b NEON. Literally the only benchmarks that benefitted from SVE were because they were limited by the 5-wide decode in NEON. Do you have an actual real-world counterexample?
- skavi 2y agoI think it's somewhat unfair to ask for real world examples when there really aren't many people writing optimized SVE code right now. Probably because there are hardly any devices with the extension. I think the transition from AVX2 to AVX512 is comparable in that it provided not only larger vectors, but also a much nicer ISA. There were certainly a few projects that benefited significantly from that move. simdjson is probably the most famous example [0]. [0]: https://lemire.me/blog/2022/05/25/parsing-json-faster-with-intel-avx-512/ https://lemire.me/blog/2022/05/25/parsing-json-faster-with-i...
- rmccue 2y agoFrom what I’ve read previously, Apple has a special licensing deal already as they were part of founding Arm, although I don’t know if there’s any details on exactly how that works.
- deleted 2y ago[deleted]
- astrange 2y agoThat seems like something people just made up, seeing as Apple didn't use ARM for something like a decade or two after that. However, Apple basically commissioned ARMv8 in the first place to develop the A/M chips, so that presumably helps.
- zimpenfish 2y ago> Apple didn't use ARM for something like a decade or two after that. They used the ARM610 in the Newton in 1993 (ARM was founded in late 1990) and then an 8 year gap to the iPod in 2001 (ARM7TDMI which are ARM designs.) Their first in-house ARM design (I believe) is the iPhone 4 in 2010. They definitely didn't "architect/design ARM" for nearly a couple of decades after founding ARM, yeah, but they did use them.
- NobodyNada 2y agoApple cofounded ARM for use in the Newton product line; they released new Newton products from 1993-97 and discontinued them in 1998. They then used ARM again for the iPod, released in 2001.
- astrange 2y agoThey didn't design the iPod ARM SoC though, nor the ones in iPhones for quite some time, and the microcontrollers in Macs for power management and such were not ARM. (I mean, some of them might've been, but the one I'm thinking of was SH or something.)
- ksec 2y ago