10 ms·
Instructions per cycle: AMD Zen 2 versus Intel
- reitzensteinm 7y agoThe second example is just a benchmark of tzcnt, added in BMI1. It's a very specific and very bizarre benchmark to do when you could just look up the reciprocal throughput (unfortunately Zen 2 has not yet been added). https://www.agner.org/optimize/instruction_tables.pdf https://www.agner.org/optimize/instruction_tables.pdf Edit: This is wrong as BeeOnRope points out below. The first is SIMD heavy, so Zen 2 mostly closing the gap with Intel in one of the areas where Zen 1 was very weak is a good thing.
- BeeOnRope 7y agoZen2 is on uops.info, it's 2L0.5T on Zen, 3L1T on Intel, so slight theoretical edge for AMD (2 vs 1 uops tho). That said, I don't agree it's a tzcnt benchmark - there are about 9 instructions only one of which is tzcnt. I'm not sure why Zen2 is worse here.
- reitzensteinm 7y agoYou're right, I messed that up (though I'll leave it for posterity). I went into it with a bias thinking BMI was slow on Zen, since PDEP is 18 cycles vs 1 on Skylake, much to my disappointment back in the day. After reviewing the example again, there's no obvious reason why Zen 2 is slower, although it's likely a rare edge case. Too bad there's nothing decent like VTune on AMD platforms. I remember one session where my choice of temporary register significantly impacted throughput while implementing an unrolled int[] hash fn on my Kaby Lake processor. I never figured out exactly why, but sharp edges do exist even on Intel chips.
- BeeOnRope 7y agoThis benchmark heavily stresses branch misprediction recovery, so that could be worse on Zen. Also, I could not reproduce Daniel's results: I got IPC of 1.77 (SKX) or 2.00 (SKL) compared to Daniel's reported 2.80 (SKL, I think), so Intel still better but by a smaller margin. Waiting for clarification on that one.
- Const-me 7y agoI wonder how reliable are these Linux syscalls? Found this http://manpages.ubuntu.com/manpages/trusty/man2/perf_event_open.2.html http://manpages.ubuntu.com/manpages/trusty/man2/perf_event_o... and that article doesn't instill much confidence in the reliability of these counters. Comment for CPU_CYCLES says "Be wary of what happens during CPU frequency scaling", comment for INSTRUCTIONS says "these can be affected by various issues, most notably hardware interrupt counts", BRANCH_INSTRUCTIONS says "Prior to Linux 2.6.34, this used the wrong event on AMD processors" and so on. If I wanted to measure what OP was measuring, I would disable frequency scaling (probably doable on overclocker-targeted motherboards, also search finds some utilities which claim to do that, both windows and linux ones), measure time, then divide by frequency.
- amluto 7y agoCPU_CYCLES counts cycles. This means that the time per cycle varies with frequency. If you're trying to see how many cycles something that fits in L1 takes, CPU_CYCLES is the right thing to measure.
- reitzensteinm 7y agoParent is pointing to documentation suggesting that it's measuring time and dividing it by frequency, and perhaps not perfectly in the case of dynamic scaling. They seem aware of what CPU_CYCLES is supposed to do.
- amluto 7y agoThe documentation is not the best. CPU_CYCLES is genuinely counting cycles. perf is all about reading actual hardware counters. It's awesome for this. There is essentially nothing made up about perf's output, except to the extent that the hardware itself reports inexact output. (For example, perf annotate may attribute events to an instruction near the instruction in question on older hardware, because older hardware has a small amount of skew when sampling.)
- chucklenorris 7y agoHeh, I'm curious if he used the mitigations for all the side channel flaws for the intel processors.
- BeeOnRope 7y agoThe mitigations don't affect CPU bound benchmarks [1] which don't call into the kernel or use specific user-space mitigations, so it won't matter here. [1] There are some rare exceptions, such as https://travisdowns.github.io/blog/2019/03/19/random-writes-and-microcode-oh-my.html https://travisdowns.github.io/blog/2019/03/19/random-writes-... , but it is unlikely to matter here.
- fulafel 7y agoSmt on/off has a large effect.
- BeeOnRope 7y agoIt's a single threaded test, so I don't think that matters here.
- fulafel 7y agoTrue. But generally it affects cpu bound benchmarks.
- NullPrefix 7y agoSMT off might mean not enough spare threads to run the OS telemetry.
- BeeOnRope 7y agoThis is Linux but I don't think that would be true even on Windows.
- amluto 7y agoThat may have been true, but it is rather dramatically false with the new JCC erratum workaround. It’s also false if you’re using a hypervisor that mitigates the iTLB multihit issue.
- yifanlu 7y agoAssuming both Intel and AMD implement performance monitors the same (i.e. same notion of instructions executed, which may be hard to measure with speculative execution), the comparison is still flawed because it doesn’t matter if Intel can do more instruction per cycle if AMD can produce more cycles in a span of wall time. > However, it is not clear whether these reports are genuinely based on measures of instruction per cycle. Rather it appears that they are measures of the amount of work done per unit of time normalized by processor frequency. That’s precisely why nobody really uses IPC as a way to compare processors. “How much work done per unit of time” is a much better measurement and I guess for historical reasons, people conflate it with IPC. But real textbook IPC is useless for comparison.
- BeeOnRope 7y agoI'm this case the frequencies are similar and so wall clock time reflects the IPC difference (also, the two CPUs take the same code path, so the I is the same in this case, which isn't always true).
- mping 7y agoBut on these processors, I believe the frequency is rarely sustained right? Due to thermal throttling and other factors.
- ShinyRice 7y agoThat only really happens on laptops, which can't dissipate as much heat as desktop systems due to size constraints. On a desktop, if you're using even AMD's stock cooler, you won't thermal throttle. That is, if you don't overclock.
- AstralStorm 7y agoIt's not about throttling. What will happen is that the CPU won't automatically clock up dynamically as much if you have worse cooling. They behave like GPUs more and more with regards to clocks.
- NohatCoder 7y agoIn case anyone is not aware: This is a very small sample of microbenchmarks. When benchmarking very simple tasks like these performance tend to vary wildly between architectures. For instance instructions are assigned to one of a handful of ports when executed, certain instructions may only be assigned to certain ports, what ports an instruction may be assigned to differ between architectures. If an inner loop use only a few different instructions one architecture may be unlucky in that most of the instructions need the same ports, and so it can execute fewer instruction overall. For real benchmarking use lots of different complicated jobs. It is not perfect, but it is the best way we have of comparing different processors head to head.
- mrb 7y agoIndeed. Back in 1999 the AMD K7 was a full 3 times faster than Intel on microbenchmarks measuring the performance of ROR/ROL instructions, because the throughput per clock of these rotate instructions was exactly 3 times higher than on Intel. Obviously this did not mean that AMD was 3 times faster than Intel. Picking 1 or 2 random microbenchmarks like the blog post author did is not useful to categorize overall performance across all real-world workloads. If he had picked different ones, they might have shown AMD twice faster than Intel.
- BeeOnRope 7y agoExamples like that still exist: AMD popcnt throughput is 4x Intel's, for example (4/cycle vs 1).
- ncmncm 7y agoThe author appears to be benchmarking the specific operations that bottleneck their json parsing library when running on an Intel chip, which seems reasonable, on its face. It can fail if the library is limited by a different set of operations, on the different machine. But that is unlikely if the specific operations tested are slower.
- alecmg 7y agoUseless, strictly academic interest. There is more than execution ports in design of processors. Not every task can be SIMD optimized to extent of approaching theoretical IPC limits, most will be bottlenecked by memory access or even IO. I prefer the "fake" but real-world IPC. Same clocks, same real world task, measure time to finish.
- fluffything 7y agoI think recommending people to to prefer {insert your favourite benchmark here} is very bad advice, and disproving your claim that Lemire's benchmarks are useless because YOU don't care about them is as simple as showing that they are useful for Lemire, which is something this post shows. If you care enough about a particular CPU to do benchmarks, you should benchmark what YOU care about. Lemire's job is to improve the implementation of particular algorithms to make optimal use of the hardware. Knowing the different theoretical hardware limits tells you how good an implementation is doing along different axes, and benchmarking those limits is a critical part of doing Lemire's job correctly. You probably have a different use case for computers than Lemire, and it is therefore completely reasonable for you to care about different benchmarks.
- Erwin 7y agoI think this was more of a response to the linked benchmark at guru3d which said: > Instructions per cycle (IPC) > For many people, this is the holy grail of CPU measurements in terms of how fast an architecture per core really is. Based on his work with simdjson, professor Lemire seems to be quite aware of microbenchmarks being problematic. But general articles out here and on HN are proclaiming Intel is doomed and can never recover, due to mitigations/lack of cores/lack of chiplets. Those concerns have yet to be reflected in the stock price.
- pjc50 7y agoIntel are behind. They have a pretty big cash buffer and a solid sales channel, as well as being pretty entrenched in OEMs. So they are a very long way from being doomed, even if it takes them a long time to turn the ship around (like 00's Microsoft).
- _ph_ 7y agoWhile only being part of the performance equation, analyzing IPC can be quite interesting in understanding the design of the processor and how performance might be achieved. One thing itches me with the presented comparison: it is running very few benchmarks generated with the same compiler. For a thorough IPC analysis, shouldn't the tests rather being programmed in assembly to exclude any influence by the compiler choice? Also probably a wider range of algorithms should be checked, as IPC on modern processors depends less on how many cycles a certain instruction takes (you should be able to find that in the manuals), but how well multiple components of the processor can be utilized at the same time. Which extremely depends on the actual program to be run.
- nabla9 7y agoIn more comprehensive single thread benchmarks (single thread POV Ray) Intel can still beat Zen 2 architecture sometimes. This test seems to indicate the reason why.
- qxnqd 7y agoITT: AMD apologists. Sorry guys but Intel is still king of single core performance. But that's not a problem because I'm sure by 2050 most desktop applications and games will correctly make use of many cores, then AMD will reign
- tempguy9999 7y agoWorth responding to blatant troll to point out it's not about performance but performance by price for 99% of uses.
- eyegor 7y agoOr performance per watt in server land. Which is a metric that zen 2 dominates in. Very few applications truly care about maxing performance at all costs.
- deleted 7y ago[deleted]
- ncmncm 7y agoBut, indeed, some do. They will provide as much power and as much cooling as they need to get that performance.
- ncmncm 7y agoBy 2050 we might well be more concerned with which rocks can be slung farther. Assuming civilization will survive until then, given current political trends, is rash.
- tempguy9999 7y agoI'm rather surprised at the claim that "but it might easily execute 7 billion instructions per second on a single core". I'd even question it except the author's an expert. If you can keep it fed then ok but one cache miss to main mem, either instruction or data, will allow the instruction buffers to completely empty and stay empty for quite a long time. I don't think you can control placement to reasonably assure cache hits always for anything but the most trivial code, am I missing something? Also if you could keep a consistent throughput like this I wonder if thermal throttling might have to kick in. I mean you're doing a lot of work...
- touisteur 7y agoI can't find it back but in a recent article I read that it was useful to have an idea of the upper-boundary abilities of an arch+algorithm, so that you 'know' what you're aiming for, but it might not be attainable practically without huge human or decades of superoptimizer effort... Yes if your algorithm reaches for cold data, you'll get hit. Can you get around that? Do you really need to hit the cache when you're computing the seven-billionth decimal of pi or factoring numbers ? This work is quite interesting, if only for compilers or superoptimizers.
- eyegor 7y agoI think the only real way to compare IPC is to actually talk to the architects. Trying to write microbenchmarks is a fools errand when you aren't aware of how the cpu processes the instructions you give it. Are you actually stressing the fpu, or is the cpu speculatively executing and then branch predicting the workload (common for micro loops)? If it is, is that what you meant to test? Are you trying to compare like for like (in which case you have to write assembly), or are you trying to write performance benchmarks (and then the only meaningful metric is cpu time)? This is an interesting idea, but I'm not sure how you could derive meaning from comparing two vastly different architectures at such a high level.
- zippie 7y agoIPC microbenchmarks do not properly reflect the complex workloads running on post Zen2 microarchitecture. Zen2 upends microarchitecture schematics enough to warrant a different metric. IPC MB’s, in my experience, tend to benchmark best case scenarios and that is probably the exception rather than the rule for application workloads in modern MA’s. Case in point, microbenchmarks showed significant improvements in IPC for Zen2 in lieu of Skylake yet for the application workload (CPU data bound), Skylake held up neck and neck. The more appropriate benchmarking metric for post-Zen2 processors is CPI [0]. [0] https://john.e-wilkes.com/papers/2013-EuroSys-CPI2.pdf https://john.e-wilkes.com/papers/2013-EuroSys-CPI2.pdf
- mmrezaie 7y agoBut isn't CPI is just reverse of IPC, and CPI just makes the IPC score being bounded between 0..1?
- jonstewart 7y agoIt’s depressing how many comments here are quick to dismiss the benchmarking/article. Yes, yes, memory bandwidth, I/O, and cache hierarchies are all important, but Daniel Lemire is one of the top people in the world when it comes to optimizing algorithms for modern CPUs. Do you like search engines? Lemire has made them significantly faster. He is often able to take code/algorithms that already seem fast, and make them much faster. He’s recently branched out beyond search engine core algorithms into some aspects of string processing (base64, UTF-8 validation, JSON parsing). In this blog post, he’s paying attention to IPC because he’s typically working with inner loops where the data’s being delivered from RAM to L1 as efficiently as possible.
- BeeOnRope 7y agoI have plenty of respect for Daniel (and you can even find me below in this discussion defending some aspects of this test), but I too find some fault with this article. The main problem I have is that the claim in dispute seems to be that Zen 2 has comparable (perhaps slightly higher) IPC to Skylake, and then Daniel picks out two benchmarks and shows that Skylake has higher IPC than Zen 2... proving what exactly? Contradicting people who said that Zen 2 had a higher IPC on every benchmark? Yes, those people were wrong, but it's easy to prove a point if you pick an argument almost no one was making it in the first place. In the same (second) benchmark that he selected the "basic_decoder" sub-benchmark, but there is also another benchmark "bogus" which tests the empty function calling time, and this case I measure a reversed scenario: Intel at IPC 2.25 and AMD at 3.43. So should we now say that Intel IPC is "quite poor"?
- jonstewart 7y agoHa, I’m not referencing _your_ comments here, and I am curious about how you couldn’t reproduce his results; he’s quick to publish and seems happy to correct so we’ll see. I’m referencing more the other comments here saying things about the benchmarks not being realistic because good benchmarks need to have a mix of tasks, like memory and I/O—this ain’t a Phoronix post, folks. I started reading through your blog last night. I’m slowly trying to learn how to go from being a programmer who doesn’t write slow code, to one who writes fast code, so absorbing a lot about vectorization and ILW, etc.