4 ms·
I have always had issues with the perf call trace sampling with frame pointers, even when virtually everything in userspace compiled with fno-omit-frame-pointer
by fooblaster 2y ago
I have always had issues with the perf call trace sampling with frame pointers, even when virtually everything in userspace compiled with fno-omit-frame-pointer. It doesn't look like any of the failure modes listed in the article to me though. Shrug.
FYI, if you happen to be running on an intel cpu, --call-graph lbr uses some specicalized hardware and often delivers a far superior result, with some notable failure modes. Really looking forward to when AMD implements a similar feature.
- rwmj 2y agoThe problem with Intel LBR (last branch records) is that the depth of the call stack is relatively limited. It depends on the generation of CPU, but LWN has a table here: https://lwn.net/Articles/680985/ https://lwn.net/Articles/680985/ Anything less than 32 is fairly useless for profiling from the kernel through to userspace.
- fooblaster 2y agoYeah, right. I still seem to get far more "comprehensible" traces when using it, even with this limitation. It's often really easy to localize where a trace is coming from, even when truncated. It probably breaks flamegraphs though.
- Sesse__ 2y agoI've tried --call-graph lbr a bunch of times, but often, it… just returns junk? I don't fully understand why, it sometimes returns wild pointers even if you don't have deep stacks.
- fooblaster 2y agoI often get junk when sampling without lbr. Which kernel are you running? The quality of perf and the associated perf_events varies wildly across kernel versions.
- Sesse__ 2y agoA variety of kernels over the last five years, on a multitude of Intel CPUs. :-) I last tested this on 6.10, I think. It's certainly true that there can be junk in --call-graph fp, too.