3 ms·
The reality is that a lot of profilers cause a sort of Heisenberg effect when running in production where it can slow down your code so much that it's not as me
by devonkim 6y ago
The reality is that a lot of profilers cause a sort of Heisenberg effect when running in production where it can slow down your code so much that it's not as meaningful anymore. This is now the time your ops engineers like me will be wagging their finger saying the application should have had instrumentation or APM support built-in like via OpenTracing or that ebpf and friends could be useful if your production machines were reasonably up to date. In most occasions I've seen, the majority of engineering teams outside massive scale companies with gobs of resources even today wind up debugging performance problems like it's 1999 with application-external hypotheses and checking for smoking guns like high page fault rates, dropped packets, etc. (eg. USE methodology).
- jeffbee 6y agoSampling cycle count or LBR with linux perf events is almost invisible to performance, maybe a 1% hit to throughput as a general rule. The problems come from languages where the PC is irrelevant, like Python, but nobody uses Python because it's fast, so using a python profiler like pyflame should be fine, even in production.
- devonkim 6y agoProfiling in production with the JVM using something like YourKit or JProfiler is the typical case for myself. Ironically, I've found profiling with Python in production easier for the reasons you've mentioned. If something is down or running slowly already, adding another 3%+ latency is hardly going to be an issue. Architecturally, with big monolithic programs that do too many things attaching a profiler to try to analyze 1% of the program's responsibilities or surface area becomes a risk to other production operations unfortunately. In most cases slowdowns happen because of resource saturation, things timing out, blocking on shared resources. In the first scenario, trying to run a profiler can exacerbate the problem or even fail to start, so the only way forensics can be done there is by emitting observability data prior to the failure point. Other approaches taken have been the more Erlang style "let it fail" methodology which is fine for newer projects but represents a rewrite for most systems in practice and is thus far, far beyond profiling discussions.
- cozzyd 6y agoBefore reaching for a "real profiler" there's always the poor man's sampling profiler: gdb -p $PID ctrl-c bt c ctrl-c bt c crl-c bt c ...
- viraptor 6y agoLess keyboard bashing: `while true ; do gdb -p $PID --batch -ex bt ; done`