4 ms·
The problem with cache misses specifically is that perf can't point an event to its exact location so you may get a high rate of cache misses on instruction tha
by vient 3y ago
The problem with cache misses specifically is that perf can't point an event to its exact location so you may get a high rate of cache misses on instruction that does not even work with memory, like jump, so you need to look what gets executed before that instruction and what looks like probable source there. I am interested if there are any more efficient approaches than this.
- soulbadguy 3y agoPerf can. But you need to massage it a bit. Out of the box , perf uses a bunch of generic performance counter which maybe not be the best. Intel cpu's have a a collection of "precisise" events which record the exact originating/offending instructions. If you use those it should fix your problem Pmu-tool by default uses precise events when available
- vient 3y agoOh, I did not know that perf by default may use some uncore events that are less precise than some Intel-specific events. Thanks for the info!
- soulbadguy 3y agoYou might want to look at https://icl.utk.edu/papi/ https://icl.utk.edu/papi/ . It's a generic library which abstract the details of specific cpu arch and aims to provide a single unify view of performance counters. I think somewhere in there you can see how perf translate counters to intel (or AMD) specific one.