4 ms·
What off-by-one on crashing instructions are you referring to? Some stack walkers intentionally do off-by-one byte offsets when doing symbol/source lookup on pa
by brucedawson 7y ago
What off-by-one on crashing instructions are you referring to? Some stack walkers intentionally do off-by-one byte offsets when doing symbol/source lookup on parent functions, but not on the crashing instruction.
Crashing is (with the exception of floating-point exceptions) precise. A particular instruction crashes, the exception record points there, the instructions afterwards are discarded with no side effects. This is necessary to support things like restarting execution.
Sampling, on the other hand, is not even well defined. There are hundreds of instructions in flight, many completing simultaneously, and when an external interrupt happens the CPU has to decide which ones to commit and which to discard. The linked article gives many more thoughts about how the CPU draws the line in the silicon.
- Const-me 7y ago> What off-by-one on crashing instructions are you referring to? Here’s an example for GCC on ARM Linux: https://github.com/dotnet/corert/issues/7826 https://github.com/dotnet/corert/issues/7826 I think I have observed similar symptoms on Windows, too. > Sampling, on the other hand, is not even well defined A crash by e.g. RAM access violation, and interrupt generated by CPU to collect sample for a profiler, are pretty similar, IMO.
- rrss 7y agoaccess violations are generated internally by a single instruction, and all mainstream CPUs guarantee precise exceptions. PMU interrupts for sampling are external so the CPU picks wherever it wants to stop in the program.
- Const-me 7y ago> all mainstream CPUs guarantee precise exceptions Yes, and precise interrupts, too. > CPU picks wherever it wants to stop in the program. I've used profilers quite a lot, and based on my observations they're quite accurate, to exact instruction.
- rrss 7y ago> Interrupt-based sampling introduces skids on modern processors. That means that the instruction pointer stored in each sample designates the place where the program was interrupted to process the PMU interrupt, not the place where the counter actually overflows, i.e., where it was at the end of the sampling period. In some case, the distance between those two points may be several dozen instructions or more if there were taken branches. https://perf.wiki.kernel.org/index.php/Tutorial#Event-based_sampling_overview https://perf.wiki.kernel.org/index.php/Tutorial#Event-based_...
- brucedawson 7y agoBut what does that even mean? Seriously. An interrupt fires, on a particular clock tick. At that point there are, let's say, 130 instructions in flight. In the case of a loop like this one there may be seven instructions being retired per clock-cycle. So, you end up with patterns. I linked to some detailed reverse engineering of which instructions are likely to end up being the victim. One common pattern is that the instruction after an expensive one will have the samples assigned to it, but there's more to it than that - I recommend reading it. TL;DR - I'm not saying you're wrong, it's just that you're not saying anything specific enough for write/wrong to apply. "accurate, to exact instruction" has not been meaningful for sampling profilers for more than 2.5 decades.