52 ms·
> perf has an option to use DWARF data for stacks, but IMHO it simply does not work. Could DWARF data be fixed/improved so that this works reliably? If so, we
by wizeman 4y ago
> perf has an option to use DWARF data for stacks, but IMHO it simply does not work.
Could DWARF data be fixed/improved so that this works reliably?
If so, we could have our cake and eat it too (in the future, not right now, of course).
- eklitzke 4y agoDWARF perf sampling works fine but has really high overhead, requires reading a large amount of stack space to find each prior frame, and the current tooling for generating perf reports from DWARF samples is really slow. You could definitely improve the last issue, I suspect no one has because it's not widely used. I don't think you can fix the first two issues, or at least you can't make it as fast as frame pointers. Also, frame pointers just use up one extra register. The actual overhead of compiling with -fomit-frame-pointer vs -fno-omit-frame-pointer is small enough for most programs that it's definitely worth including frame pointers if you think you'll ever need to profile a program running as a release build. I'll also add frame pointers are very useful for debugging if you don't have debug symbols, which is more often than you might think.
- kentonv 4y ago> Also, frame pointers just use up one extra register. This comment suggests it's more than that: https://lwn.net/Articles/920165/ https://lwn.net/Articles/920165/ My non-expert understanding: It seems that when frame pointers are enabled, the compiler prefers to address all stack variables using offsets from the frame pointer, whereas when frame pointers are disabled, stack variables are addressed at offsets from the stack pointer. But, it turns out frame-pointer-based addressing is slightly less efficient than stack-pointer-based addressing. The problem is that the offset from the stack pointer / frame pointer is encoded into the instruction as either 8 bits or 32 bits -- 8 bits if that's enough, 32 bits otherwise. But it turns out that for large stack frames, the stuff close to the frame pointer is more stuff that isn't actually used during the function body, such as saved register values, whereas the stuff close to the stack pointer is typically the data that's currently being operated upon. So, addressing based on the stack pointer is less likely to spill over into 32-bit offsets. However, this all sounds like something that could be fixed in the compilers. They could still address off the stack pointer even when frame pointers are available. In fact, if they could choose on an instruction-by-instruction basis, maybe they could even save bytes. But for some reason they don't currently do this. It's not clear to me why -- maybe just because omitting frame pointers has been such a standard optimization for so long that no one has bothered optimizing the other case?
- eklitzke 4y agoThat makes sense and isn't something I had considered, but a 32-bit displacement vs an 8-bit displacement just leads to bloat in binary size, it doesn't affect how many cycles your movs jumps etc. take. There are some second order effects of things where larger code size can cause you to get a worse hit rate in the instruction cache, but usually those effects are miniscule when you actually benchmark code. There are going to be some pathological programs where the instruction cache hit rate gets way worse with the large binary size, but I can't really think of many programs I've seen that are mostly stalled on instruction decoding/fetching. Here's my hypothesis of why stack-pointer offsets haven't been implemented when compiling with frame pointers. At companies like Google/Facebook that compile huge C++ binaries (usually 100MB+ stripped) it's common that production binaries would be compiled with something like -fno-omit-frame-pointer (so system-wide profiling can be done on all C++ code in production) and also generate split debug symbols using -gline-tables-only. The latter compiler flag basically just generates enough DWARF information to recover the mapping of pc -> source code line, so if you get a core dump you can figure out which line of C++ code was actually being executed in each frame of the call stack. If you're doing frame-pointer offsets it means that you can also recover the value of local variables in each frame (assuming they haven't been optimized out by the inliner) just based on the offset of the variable from the frame-pointer. So basically -fno-omit-frame-pointer and -gline-tables-only give you enough debug data to get the full call stack with line numbers and the values of local variables in each frame (except the inlined ones), while also minimizing the cost of generating/storing debug data.