3 ms·
Yes but this level of info is too coarse. It is not sufficient to know your level of L1/L2/L3 misses if your app has megabytes of code and gigabytes in heap. As
by vient 3y ago
Yes but this level of info is too coarse. It is not sufficient to know your level of L1/L2/L3 misses if your app has megabytes of code and gigabytes in heap. As I said in adjacent response, I tried to record cache misses with `perf` but the info produced by `perf report` is confusing at best - I did not manage to find any good tutorial on debugging cache misses in that way.
- dragontamer 3y agoYou know about the --control=fifo option of perf, right? You can turn perf on, and off, using that fifo. You can use this to programmatically enable the profiler for individual sections of code, and then disable it for the other sections. ----------- Are you coarse-profiling your code first? Have you run gprof first and figured out which section of code you're focusing on? EDIT: Using perf alone (though IMO a bad idea), you can also get instruction counts. You should be focusing on the instructions that are run the most. But because instructions take a variable amount of time, what you really want to be doing is profiling using a timer-based method (ex: gprof) and narrowing your search through that instead. But sometimes, the instruction-based approach is also useful.
- vient 3y agoManually enabling/disabling perf is too slow when you want to profile a part that takes <1ms to execute, but it is an option if nothing better exists. It would be great if you could just isolated interesting code in microbenchmark but we all know that this will skew results. About code execution profiling, I usually use Intel PT so yes, I know where I want to look.
- soulbadguy 3y ago> I did not manage to find any good tutorial on debugging cache misses in that way. Saddly this is very true. There is a lack of good performance investigation information out there. And pref while very useful is not really a beginner tools. In your case, you want an "annotated" profile. Where perf will annotate the source code with the even, and basically pin-point which part of your code is producing the offending events (in this case the cache misses). If you really want to dig deep, i can't recommend intel Vtune. Only works with intel cpu (i think ?), but is the best tool in you want to understand performance at a deeper level.
- vient 3y agoThe problem with cache misses specifically is that perf can't point an event to its exact location so you may get a high rate of cache misses on instruction that does not even work with memory, like jump, so you need to look what gets executed before that instruction and what looks like probable source there. I am interested if there are any more efficient approaches than this.
- soulbadguy 3y agoPerf can. But you need to massage it a bit. Out of the box , perf uses a bunch of generic performance counter which maybe not be the best. Intel cpu's have a a collection of "precisise" events which record the exact originating/offending instructions. If you use those it should fix your problem Pmu-tool by default uses precise events when available
- vient 3y agoOh, I did not know that perf by default may use some uncore events that are less precise than some Intel-specific events. Thanks for the info!
- soulbadguy 3y agoYou might want to look at https://icl.utk.edu/papi/ https://icl.utk.edu/papi/ . It's a generic library which abstract the details of specific cpu arch and aims to provide a single unify view of performance counters. I think somewhere in there you can see how perf translate counters to intel (or AMD) specific one.