4 ms·
How do flame graphs handle the case where most of the time is spent in some leaf function that is called from all over the program? In this case, each individua
by zlurkerz 2y ago
How do flame graphs handle the case where most of the time is spent in some leaf function that is called from all over the program? In this case, each individual stack would not take much time but in aggregate, a lot of time is spent in the function at the top of all of the call stacks. This should not be that uncommon to have hotspots in things like copying routines, compression, encryption etc that are not associated with any particular stack.
pprof from https://github.com/google/pprof https://github.com/google/pprof can produce a DAG view of a profile where nodes are sized proportional to their cumulative time, e.g.,
https://slcjordan.github.io/images/pprof/graphviz.png https://slcjordan.github.io/images/pprof/graphviz.png and such a view would seem to cover the case above and subsume the usual use cases for a flame graph, would it not?
Although I guess a flat text profile of functions sorted by time would also highlight these kinds of hot spots. Still, if we want a single graphical view as a go-to, it's not clear that flame graphs are all that much better than pprof DAGs.
- mandarax8 2y agoMy flame graph tool (KDAB hotspot) has a bottom-up flame graph for that purpose.
- zlurkerz 2y agoRight, but then you have to know to consult both top-down and bottom-up. Seems like a DAG is a union of these views.
- nemetroid 2y ago> How do flame graphs handle the case where most of the time is spent in some leaf function that is called from all over the program? In my experience, they don't really. They're very good for finding easy wins where a high/medium-level operation takes much longer than you'd expect it to, not so much for finding low-level hotspots in the code. > such a view would [...] subsume the usual use cases for a flame graph, would it not? I don't think it does. A flame graph gives you a good idea of the hierarchical structure (and quickly). Sometimes you can work that out from a DAG, but not necessarily: Let's say that A and B call M, and M calls X and Y. All four edges have roughly the same value. A flame graph can show you that 90% of the time spent in X started with A and that 90% of the time spent in Y started with B, but the DAG can't.
- tanelpoder 2y agoI had the same experience/question at some point when I noticed (and investigated), why didn't a new Ubuntu Linux install break out interrupt usage in vmstat anymore. [1] Perf+flamegraphs were able to catch in-interrupt-handling samples as I was on a bare metal machine with PMU counters accessible. But the various NVMe I/O completion interrupts happened all over the place, even when the CPUs were in random userspace code sections. Neither the bottom-up & top-down FlameGraph approach made it visually clear in this case. But since interrupt time did show up and since the FlameGraph JS tool had a text search box, I searched for "interrupt" or "irq" in the search box - and the JS code highlighted all matching sections with purple color (completely distinct from the classic color range of flamegraphs). And seeing that purple color all over the place made me smile and gave me a strong visual "a-ha" moment. Probably with some extra JS tinkering, these kinds of "events scattered all over the place" scenarios could be made even clearer. [1] https://news.ycombinator.com/item?id=26139611 https://news.ycombinator.com/item?id=26139611 (2021)
- yxhuvud 2y agoThey don't. The best way to visualize that, that I've seen is the DHAT tool of Valgrind, that basically builds a trie from the roots based on how much allocations happens in them (the tool measures allocations, but the visualization could just as well be used for time spent).
- foota 2y agoNot sure it's a common feature, but I've seen a flame graph tool that you can pivot around some function, so that it filters to only stacks involving that function, and then goes both up and down from there.
- vient 2y agoI like speedscope.app for viewing flamegraphs. It is more interactive than traditional SVG flamegraphs, and what is relevant here is a "sandwich" view - basically a sorted list of all functions, you see what function was spent the most time in, click on it and see all calling stack traces like a mini flamegraph, filtered and centered on this function. Speedscope supports several popular trace formats, really useful.
- nox101 2y agoI like perfetto https://perfetto.dev/ https://perfetto.dev/
- dangets 2y agoJava Mission Control [0] has a button to toggle for displaying the profile as thread roots or method roots for this purpose. I am not imaginative enough to come up with a visualization that shows both (maybe utilize background color or another indicator to show the leaf function's relative frequency in the other direction?). Either way, both directions have their use case when investigating. [0] https://adoptium.net/jmc/ https://adoptium.net/jmc/