3 ms·
It'd be more interesting to see a test with more branching as that is the one big area where x86 CPUs tend to shine compared to "simpler" architectures (or thin
by largote 11y ago
It'd be more interesting to see a test with more branching as that is the one big area where x86 CPUs tend to shine compared to "simpler" architectures (or things like GPUs).
Also, I don't think this exercises the Floating-Point units for these CPUs (haven't looked at the code though).
- twotwotwo 11y agoYeah. This was another benchmark that found that an earlier Apple CPU got impressive instructions-per-clock on a couple of cryptographic functions (that aren't special-cased like AES): https://zerobyte.io/blog/2014/04/29/benchmarking-symmetric-crypto-on-the-apple-a7/ https://zerobyte.io/blog/2014/04/29/benchmarking-symmetric-c... But those functions also don't really stress-test branch prediction, etc. They use instruction-level parallelism but the control flow and memory accesses are predictable. (That's almost necessary for software crypto; if execution isn't ~constant-time, you risk timing attacks.) Looking at transistor counts: Apple said the three-core A8X had three billion transistors total. Given Apple's focus on GPU and bringing other components onto the SoC, many of those are not spent on the cores; still, in raw count, it's right up there with Intel's Core-branded CPU+GPU dies, going by a table in Wikipedia. AnandTech put up a lot of numbers about the A9, including specs like cache sizes and memory bandwidth and benchmarks like Geekbench and SPEC: http://www.anandtech.com/show/9686/the-apple-iphone-6s-and-iphone-6s-plus-review/4 http://www.anandtech.com/show/9686/the-apple-iphone-6s-and-i... Seems like, at a minimum, if your intuitions about mobile performance come from early in-order chips, they may not apply to Apple CPUs or even the many Cortex-A57-based SoCs out there. But it also doesn't seem like you can say those chips are up there with Intel ones in general. I guess the evergreen-but-not-wrong conclusion is that to really know your code's performance, you wan to test it on hardware as similar as possible to what it'll really run on.