4 ms·
Interesting. I thought it would be interesting to compare the behaviour of (very) different AArch64 processors on this code. I ran your code on an Oracle Clou
by error503 4y ago
Interesting.
I thought it would be interesting to compare the behaviour of (very) different AArch64 processors on this code.
I ran your code on an Oracle Cloud Ampere Altra A1:
sum_slice time: [677.45 ns 684.25 ns 695.67 ns]
sum_ptr time: [689.11 ns 689.42 ns 689.81 ns]
sum_ptr_asm_matched time: [1.3773 µs 1.3787 µs 1.3806 µs]
sum_ptr_asm_mismatched time: [1.0405 µs 1.0421 µs 1.0441 µs]
sum_ptr_asm_mismatched_br time: [699.79 ns 700.38 ns 701.02 ns]
sum_ptr_asm_branch time: [695.80 ns 696.61 ns 697.56 ns]
sum_ptr_asm_simd time: [131.28 ns 131.42 ns 131.59 ns]
It looks like there's no penalty on this processor, though I would be surprised if it does not have a branch predictor / return stack tracking at all. In general there's less variance here than the M1. The SIMD version is indeed much faster, but by a smaller factor.
And on the relatively (very) slow Rockchip RK3399 on OrangePi 4 LTS (1.8GHz Cortex-A72):
sum_slice time: [1.7149 µs 1.7149 µs 1.7149 µs]
sum_ptr time: [1.7165 µs 1.7165 µs 1.7166 µs]
sum_ptr_asm_matched time: [3.4290 µs 3.4291 µs 3.4292 µs]
sum_ptr_asm_mismatched time: [1.7284 µs 1.7294 µs 1.7304 µs]
sum_ptr_asm_mismatched_br time: [1.7384 µs 1.7441 µs 1.7519 µs]
sum_ptr_asm_branch time: [1.7777 µs 1.7980 µs 1.8202 µs]
sum_ptr_asm_simd time: [421.93 ns 422.63 ns 423.30 ns]
Similar to the Ampere processor, but here we pay much more for the extra instructions to create matching pairs. Interesting here that the mismatched branching is faster than the single branch.
I guess absolute numbers are not too meaningful here, but a bit interesting that Ampere Altra is also the fastest of the 3 except in SIMD where M1 wins. I would have expected that with 80 of these cores on die they'd be more power constrained than M1, but I guess not.
Edit: I took the liberty of allowing LLVM to do the SIMD vectorization rather than OP's hand-built code (using the fadd_fast intrinsic and fold() instead of sum()). It is considerably faster still:
Ampere Altra:
sum_slice time: [86.382 ns 86.515 ns 86.715 ns]
RK3399:
sum_slice time: [306.94 ns 306.94 ns 306.95 ns]
- sakras 4y agoIf you’re only heavily using one of the cores, that core is free to use a lot more power, and can probably push its clock speed much higher than if this were an all-core workload. So I’d actually expect the opposite, that the ampere would be allowed to use a lot more power than the M1 (since it’s not a laptop).
- error503 4y agoAmpere's processors are more or less fixed clock, which makes sense to me in a processor designed for cloud servers. You don't really want unpredictable performance, and in a multi-tenant situation like Oracle Cloud where I ran it, you don't want customer A's workload to affect customer B's performance running on the same physical processor. With 10x the number of cores as the M1 Max and a TDP of 250W (maybe 5x? Apple doesn't publish numbers), the average power limit per core is likely significantly less, and the M1 might be able to leverage 'turbo' here for this short benchmark. Still, this is not really a meaningful benchmark, just interesting.
- sweetjuly 4y ago>It looks like there's no penalty on this processor, though I would be surprised if it does not have a branch predictor / return stack tracking at all One potential explanation is that a lot of processors have "meta" predictors. That is, they have multiple distinct predictors that they then use a second level predictor to decide when and how they use it. This is really useful because some predictors perform very well in certain cases but very poorly in others. Therefore, what may be happening is that the RAS is getting overridden by another structure since the predictor detects that the RAS is frequently wrong but another predictor is frequently right.