4 ms·
Nice write up! Using BPF to trace malloc/free is good example of the tool’s power. Unfortunately, IME, this approach doesn’t scale to very high load services. O
by cranekam 5y ago
Nice write up! Using BPF to trace malloc/free is good example of the tool’s power. Unfortunately, IME, this approach doesn’t scale to very high load services. Once you’re calling malloc/free hundreds of thousands of times a second the overheard of jumping into the kernel every time cripples performance.
It would be great if one could configure the uprobes for malloc/free to trigger one in N times but when I last looked they were unconditional. It didn’t help to have the BPF probe just return early, either — the cost is in getting into the kernel to start with.
However, jemalloc itself has great support for producing heap profiles with low overhead. Allocations are sampled and the stacks leading to them are recorded in much the same way as the linked BPF approach:
https://github.com/jemalloc/jemalloc/wiki/Use-Case:-Heap-Profiling https://github.com/jemalloc/jemalloc/wiki/Use-Case:-Heap-Pro...
- kouteiheika 5y ago> Once you’re calling malloc/free hundreds of thousands of times a second the overheard of jumping into the kernel every time cripples performance. Shameless plug in case you (or anyone else) is interested, I wrote a memory profiler for exactly this usecase: https://github.com/koute/bytehound https://github.com/koute/bytehound It's definitely not perfect, but it's relatively fast, has an okay-ish GUI, and it's even scriptable: https://koute.github.io/bytehound/memory_leak_analysis.html https://koute.github.io/bytehound/memory_leak_analysis.html
- cranekam 5y agoInteresting! What is the overhead of this? We found that jemalloc's heap profiling had a small (perhaps 1-2%? It's been a while) CPU penalty and, depending on the complexity of the code being profiled and the sample rate, potentially a few hundred MB of extra RAM use on very large, complex binaries. I'd assume the RAM cost is similar given the data is the same (i.e. backtraces).
- kouteiheika 5y agoIt widely depends on the program's allocation patterns. It works by continuously tracking allocations and by default it emits every allocation as it happens, with a full stack trace attached. It certainly adds certain overhead, but usually it's not significant enough to be a dealbreaker (especially if you compile the current master and configure it to automatically strip out all short lived temporary allocations). I've used it to profile a bunch of time-sensitive software, e.g. telecom software and blockchain software, where other profilers were too slow to be usable. Not sure how it compares to jemalloc's profiler, but it should be orders of magnitude faster than Valgrind while at the same time gathering a lot more data.