13 ms·
So ARM64 has dedicated instructions for CRC32, but implementing it by hand using SIMD is still faster. Score another point for RISC.
by ntoskrnl 4y ago
So ARM64 has dedicated instructions for CRC32, but implementing it by hand using SIMD is still faster. Score another point for RISC.
- MichaelZuo 4y agoIt's very impressive someone messing around for a few hours could get the m1 chip to more than 2x the performance. Easy gains like that really shouldn't be possible, assuming Apple's silicon team are competent. Maybe there's some hidden gotcha here?
- naniwaduni 4y ago> Easy gains like that really shouldn't be possible, Easy gains are everywhere. The "gotcha", if you can call it that, is that optimizing particular operations comes with space tradeoffs that are more expensive when you do them in hardware.
- deleted 4y ago[deleted]
- dzaima 4y agoCRC32X works on 8 bytes at a time and has a throughput of one invocation per cycle, whereas the SIMD operates on blocks in parallel (the chromium code does 64 bytes an iteration, with a lot of instruction-level parallelism too). Theoretically M1 could have thrown more silicon at it to allow more than one CRC32X invocation per cycle, but that's not very useful if you can achieve the same with SIMD anyway.
- AlotOfReading 4y agoIntel's algorithm is very clever and would take a lot of space to implement in hardware. The implementation underlying the CRC32** instructions is probably some set of shift registers. That's a pretty good space/speed tradeoff to make. My largely uninformed guess is that they added the instructions to get fast CRCs for the filesystem 'for free'. There aren't many other cases where software CRC can be a bottleneck that also use these polynomials.
- nicoburns 4y agoI feel like CRC32 may be simple enough (and close enough to the kind of operation like adding and bit-shifting that general-purpose CPUs are good at anyway, that perhaps it doesn't benefit as much from dedicated silicon as other algorithms would.
- Sirened 4y agoThis is way more common than you'd think, and it's not by accident. Engineering teams optimize the paths that are heavily used to get the biggest improvement across the platform as a whole. CRC32X is certainly not as heavily used as NEON and so if you're forced to decide between spending area on being able to fuse extra instructions for NEON and slightly improving throughput for CRC32X, the obvious choice is NEON. You see this way more obviously on Intel's x86-64 cores where many of the highly used instructions are fast path decoded but some of the weirder CISC instructions that nobody really uses are offloaded to very slow microcode.
- bee_rider 4y agoI wonder -- could CRC32X be something that would also, specifically, not as interesting for Apple? They are mostly optimizing for desktop workloads. I wonder if worrying about checksuming, especially maximizing the throughput of checksum operations, is more of a server thing. (Like we have to checksum when we download things on desktop, but that's a one-off, and I guess things get checksummed in the filesystem, but even the nice NVME drives are pretty slow from the CPUs point of view).
- astrange 4y agoMachO uses codesigning with adhoc signatures as a form of checksumming, and there’s also TCP and whatever drives do. So it’s the converse, it’s so common the dedicated hardware does it instead of the CPU. And it’s not all the same algorithm.
- Sirened 4y agoI think it's a little bit of that and also I suspect they have far more pressing concerns than CRC32X being relatively slow (it is still a throughput of one per clock which isn't at all bad). Branch prediction and prefetching seems to be the really important problem at least for Apple due to their very deep ROB [1]. A mispredicted branch being resolved late (i.e. a branch dependent on an outstanding DRAM fetch) can lead to hundreds of executed instructions being discarded (wasting tons and tons of power and cycles). I don't quite remember the exact figure, but I've heard a good metric in CPU arch is that about one of every six instructions is a control flow instruction in general purpose programs (i.e. non-scientific/ calculation heavy). Being just a little bit faster on CRC32X calculation may not have been worth it when they could spend that precious power budget elsewhere. It's really just design choices all the way down. They may very well be doing a lot of CRC32s but they're almost certainly doing more of everything else than CRC32s. [1] https://www.anandtech.com/show/16226/apple-silicon-m1-a14-deep-dive/2 https://www.anandtech.com/show/16226/apple-silicon-m1-a14-de...
- interestica 4y agoSaving it for M2 to have something to show off?
- d_tr 4y agoI am not taking any hard stance on the usefulness of the specialized instruction, but M1 is a very wide and powerful core, so this won't be true everywhere. The single instruction might also be more power-efficient and keep other resources free for other stuff.
- athrowaway3z 4y agoAlso not taking hard stances, but both cases are suspect. Power efficiency because being 3 times faster means you're done 3 times earlier. Keeping other resources free because I suspect a CRC calculation is generally followed by an `if eq` statement. ( Even with out-of-order or speculative execution this creates a bottle neck that is nice to remove 3x faster )
- stingraycharles 4y agoIf you’re writing optimized code, hardly ever would you evaluate one CRC check at a time. You would process them in chunks, as the OP stated, but would just let a compiler do the auto-vectorization. This is even more true in the case of CRC, where there’s clearly almost always one branch that wins: this is perfect for branch prediction, which would mean the whole “if eq” condition is preemptively skipped.
- saagarjha 4y agoThe compiler probably isn’t going to be able to autovectorize a CRC unless you help it out.
- dottrap 4y agoI think power efficiency has a lot more variables now so it is not easy to know if consumption is linear with time. CPUs now dynamically throttle themselves, plus now Apple has advertised that its M1 cores are divided up between high-performance and high-efficiency efficiency cores, let alone how the underlying chip itself may consumer power differently for implementing different instructions. So for a hypothetical example, it could be that using general purpose SIMD triggers the system to throttle up the CPUs and/or move to the high performance CPUs, whereas the dedicated CRC instructions might exist on the high-efficiency cores and not trigger any throttling. I've forgotten all my computer architecture theory, but if I look back at Ohm's law and look at power, the equation is P = I^2 • R. Handwaving from my forgotten theory a bit here, ramping up the CPUs increases current, and we see that it is a squared factor. So by cutting the time by say a factor of 3 does mean you are done 3 times faster (which is a linear component), you still have to contend that you have a squared component in current which may have been increased. I have no clue if the M1 actually does any of this, but merely stating that it is not obvious what is happening in terms of power efficiency. We've seen other examples of this. For example, I've read that Intel's AVX family instruction generally increases the power consumption and frequency of when utilized, but non-obviously, it often runs at a lower frequency when in 256 or 512 wide forms compared to the lesser widths (which then requires more work on the developer to figure out what is the optimal performance path as wider isn't necessarily faster). And as another example, when Apple shipped 2 video cards in their Macbooks, some general purpose Mac desktop application developers who cared about battery life were tip-toeing around different high level Apple APIs (e.g. Cocoa, Core Animation, etc.) because some APIs under the hood automatically triggered the high performance GPU to switch on (and eat power), while these general purpose desktop applications didn't want or need the extra performance (at the cost of eating the user's battery).
- dragontamer 4y agoSIMD is a very powerful parallelization technique, with marvelous gains whenever I see it used. It seems like a fundamentally more efficient form of compute, but is very difficult to design algorithms for. I'd argue against "SIMD" as being "RISC", since you need all sorts of complicated instructions (ex: gather/scatter) to really support the methodology well in practice.
- tremon 4y agoBut scatter/gather is a primitive operation for SIMD, so if you want a RISC-based version of it, that's exactly what you would provide. Having dedicated instructions for specific operations (whether for crc/aes/nnp or whatever) feels like a CISC-based approach, so I think I agree with the GP. RISC vs CISC is about the simplicity of the instruction set, not about whether it's easy to use.
- mhh__ 4y agoThese days I'd argue risc vs cisc is more about regularity and directness than the size of the ISA as per se. I'd argue AArch64 isn't particularly RISC by the standards of the past but it sets the bar and tone for RISC today.
- dragontamer 4y agoAnd which SIMD instruction set should we be talking about? NEON-instructions or with the SVE instruction set? And if we're talking about multiple instruction-sets designed for the same purpose, is this thing really RISC anymore? Or do you really mean "just not x86" when you say RISC ??
- mhh__ 4y agoThat depends on how precisely you define the purpose. NEON and SVE seem to be aimed at different intensities of work.
- dragontamer 4y ago
- tlb 4y agoSimilarly, I wish that on x86, REP STOSB was the fastest way to copy memory. Because it only takes a few bytes in the icache. But fast memcpys end up being hundreds of bytes, to work with larger words while handling start and end alignment.
- saagarjha 4y agoWith ERMS it’s definitely not going to be slow, so it’s a good choice when you’re in a constrained environment (high instruction cache pressure, can’t use vector instructions).
- userbinator 4y agoIt still is in general situations (i.e. not the microbenchmarks where the ridiculously bloated unrolled "optimised" implementations may have a very slight edge.) I believe the Linux kernel uses it for this reason.
- jabl 4y agoThe kernel is a bit of a special case since very likely a syscall starts off with a cold I$, and also there's a lot of extra overhead if you insist on using SIMD registers. In general I agree with you though, optimizing memcpy implementations only against microbenchmarks is dumb.
- stncls 4y agoAlso, using SIMD registers is (generally) forbidden in kernel code, which heavily narrows down the competitors to "rep stosb".
- StillBored 4y agoThe real problem (and with the crc above) is that the fastest version for any given cpu may not be the fastest for any other. Its really short sighted to not spend the area on some of these features (aes, crc32, memcpy) because invariably one ends up with a long term optimization problem where in 5-10 years any given application has to run on on of a half dozen diffrent CPUs and optimizing for -mtune=native, and it likely results in suboptimal perf one the lastest CPUs because the micoarch designers can't be constrained to assuring that the the newer version runs any given instruction sequence proportionally faster than the previous. (aka the overall perf may go up but maybe something like the nontemporal store, or the polinomial mul doesn't keep up). And this is really the CISC vs RISC argument and why all these RISC cpus have these CISC like instructions. You want top perf in general code you assure the rep sto and mov sequences (or whatever) run the fastest microcoded version possible on a given core. But intel sorta messed this up in the p6->nehalem timeframe (IIRC when they added the fast string flag) until they rediscovered this fact. IIRC Andy Glew admitted it was a bit of an oversight combined with an release/area issue on the original PPro they intended to fix, but then it took 10 years.
- pclmulqdq 4y agoIt is not faster to use SIMD by hand - it is faster to use the vector unit alongside the integer unit, using both paths at the same time.
- userbinator 4y agoThat's very shortsighted thinking. The dedicated instruction could be optimised by the hardware in a future revision to become much faster.
- adrian_b 4y agoNo, the faster implementation uses another dedicated instruction, which happens to be more general than CRC32, i.e. the multiplication of polynomials having binary coefficients. So this has little to do with RISC, except the general principle that the instructions that are used more frequently should be implemented to be faster, a principle that has been used by the M1 designers and by any other competent CPU designers. In this case, ARM has added the polynomial multiplication instruction a few years after Intel, with the same main purpose of accelerating the authenticated encryption with AES. There is little doubt that ARM was inspired by the Intel Westmere new instructions (announced by Intel a few years before the Westmere launch in 2010). The dedicated CRC32 instruction could have been made much faster, but the designers of the M1 core did not believe that this is worthwhile, because that instruction is not used often. The polynomial multiplication is used by many more applications, because it can implement CRC computations based on any polynomial, no only that one specified for CRC32, and it can also be used in a great number of other algorithms that are based on the properties of the fields whose elements are polynomials with binary coefficients. So it made sense to have a better implementation for the polynomial multiplication, which allows greater speeds in many algorithms, including the CRC computation.
- saagarjha 4y agoAmusingly the Rosetta runtime uses crc32x
- IshKebab 4y agoThat sounds like a point for RISC to me? Ok maybe it is just a point against really complex instructions. There's clearly an optimum middle ground.
- astrange 4y agoComplex instructions are often good ideas. They’re best at combining lots of bitshifting (hardware is good at that and it can factor things out) but even for memory ops it can be good (the HW can optimize them by knowing things like cache line sizes). They get a bad rap because the only really complex ISA left is x86 and it just had especially bad ideas about which operations to use its shortest codes on. Nobody uses BOUND to the point some CPUs don’t even include it. One point against them in SIMD is there definitely is an instruction explosion there, but I haven’t seen a convincing better idea, and I think the RISC-V people’s vector proposal is bad and shows they have serious knowing what they’re talking about issues.
- rowanG077 4y agoIt's not really a fair comparison. There is only one CRC32 unit which means it can't make use of superscalar (at least if I understand the article correctly). If it would have more CRC32 units that would be the most efficient.