16 ms·
Google XRay: A Function Call Tracing System [pdf]
- 4ad 10y agoSeems much more restricted than perf or DTrace, or all the other tracing tools. I don't understand the point of this at all. NIH? I'd love to heard brendangregg's or bcantrill's thoughts on this. (Does HN give users some kind of alert when they are mentioned?).
- bitmapbrother 10y agoDoes perf or Dtrace have the same functionality? >XRay allows you to get accurate function call traces with negligible overhead when off and moderate overhead when on, suitable for services deployed in production. XRay enables efficient function call entry/exit logging with high accuracy timestamps, and can be dynamically enabled and disabled.
- 4ad 10y agoYes?
- DannyBee 10y agoFirst, i always assume dtrace can do something. Dtrace is not that low overhead in the situation that xray is describing. http://dtrace.org/blogs/brendan/2011/02/18/dtrace-pid-provider-overhead/ http://dtrace.org/blogs/brendan/2011/02/18/dtrace-pid-provid... When they traced everything, it was two orders of magnitude slower. At 100k event/s (quite possible in the situations x-ray is used), the app would be about 60% slower. For just a simple probe, not full on tracing. Does this make it better than dtrace? No. It just serves a different set of use cases.
- dap 10y ago> When they traced everything, it was two orders of magnitude slower. > At 100k event/s (quite possible in the situations x-ray is used), the app would be about 60% slower. For just a simple probe, not full on tracing. I think you've misunderstood the post. The actual time reported in that post is 600 ns per probe, with sanity checks reporting as much as 2000ns, which he concluded was in the same ballpark. He measured that by observing a program that just calls two functions in a tight loop, one of which is strlen() on a dozen-character string. That's basically the worst possible case for any function call tracing system -- I don't think it's fair to say that's "a simple probe, not full on tracing". The "60%" slower conclusion is only by construction: if you have N events per second, and the framework adds overhead of T nanoseconds per probe, then you can pick N such that the tracing framework adds whatever percentage overhead you like. In this case, Brendan picked N=100K and came up with 60% overhead for that case, but I think there's a math error in there. For that calculation, he assumes a 6us probe time instead of 600ns. I think the overhead would be 6% for 100,000 events, not 60%.
- DannyBee 10y agoI didn't misunderstand. The problem with this is that you assume that if you probe every function instead of just strlen that you get the same time bounds. The only way to do this is, afaik, something similar to http://docs.oracle.com/cd/E19253-01/819-5488/gcgmc/index.html http://docs.oracle.com/cd/E19253-01/819-5488/gcgmc/index.htm... This really just is going to expand to putting a probe on every function. I'd really like to see numbers on how dtrace handles many millions of probes at once (which is what xray is handling). It's not uncommon to have 10 or 100 million functions in some of these programs. I have strong doubts that dtrace has the same overhead given 200 million probes. (since the probe data structures alone will likely take up gigabytes of memory, accessing them is unlikely to be cache friendly, etc) AFAICT, there is no more generic function entry probe than what that blog post describes. But i'd love to be wrong, and understand how dtrace is going to determine what instructions are a function entry in several ns :P TL;DR dynamic probing infrastructures are not a panacea
- dap 10y ago> I'd really like to see numbers on how dtrace handles many millions of probes at once (which is what xray is handling). It's not uncommon to have 10 or 100 million functions in some of these programs. I have strong doubts that dtrace has the same overhead given 200 million probes. I see now. That's a fair question, and I'm not aware of data either way. I just tried a pretty simple experiment inspired by Brendan's that suggests that on my machine, the overhead is about 1450ns per probe for as many as 140,000 probes: https://gist.github.com/davepacheco/a12a0d45d55f0d7a28c312c2cd3cf234 https://gist.github.com/davepacheco/a12a0d45d55f0d7a28c312c2... > AFAICT, there is no more generic function entry probe than what that blog post describes. But i'd love to be wrong, and understand how dtrace is going to determine what instructions are a function entry in several ns :P Well, DTrace as architected is always going to pay the cost of a context switch into the kernel for each probe, and I think it's fair to take Brendan's result of 600ns as a lower bound of the per-probe overhead, at least on his machine. However, once in the kernel, for a typical native program (i.e., not JIT), I expect DTrace would only record the current userland thread instruction pointer. Names are typically resolved asynchronously by the consumer. So I would be surprised if it really was much slower, especially given the result above, but I too would like to see data. I'm not saying that DTrace solves all problems or even that the OP should have used it instead. It's certainly true that for the special case of userland function boundary tracing, one might expect to do better by skipping the context switch (at the expense of much functionality, including any ability to correlate with broader system activity). But since DTrace was brought up, I wanted to help clarify the uncertainty about what it can do and what its overhead is.
- teraflop 10y agoI don't know about DTrace, but perf is a sampling profiler, not an accurate call tracer.
- 4ad 10y agoPerf is much more than a sampling profiler: https://perf.wiki.kernel.org/index.php/Main_Page https://perf.wiki.kernel.org/index.php/Main_Page. It can use uprobes and kprobes. There are many ways to do dynamic tracing in Linux, apart from perf. Ftrace, raw kprobes/uprobes, eBPF, bcc compiler for eBPF, etc.
- teraflop 10y agoAfter perusing the perf man page, the only way I could figure out how to make it accurately count userspace function calls was using hardware breakpoints, e.g.: "perf stat -e mem:0xADDRESS:x" Obviously that's not a very good approach, because you're limited to the number of breakpoints that your CPU can handle simultaneously (4 on my machine) and there's a lot of overhead. If you know of a better way to accomplish the same thing with perf, I'd be happy to hear it.
- wmf 10y agoCheck out http://www.brendangregg.com/perf.html#DynamicTracingEg http://www.brendangregg.com/perf.html#DynamicTracingEg
- cthalupa 10y agohttps://lwn.net/Articles/499190/ https://lwn.net/Articles/499190/ & https://gnu.wildebeest.org/blog/mjw/2012/05/24/pull-user-space-probe-instrumentation/ https://gnu.wildebeest.org/blog/mjw/2012/05/24/pull-user-spa... http://www.brendangregg.com/blog/2015-06-28/linux-ftrace-uprobe.html http://www.brendangregg.com/blog/2015-06-28/linux-ftrace-upr...
- teraflop 10y agoAh, cool. I just tested it out and it seems to work as documented. Unfortunately it requires root access, and incurs about 1 microsecond of overhead per function call on my machine.
- asuffield 10y ago(Tedious disclaimer: my opinion only, not speaking for anybody else. I'm an SRE at Google) I recommend reading page 2 of the paper, which discusses the specific set of features that XRay offers. How do I get perf or DTrace to give me the six things listed there? I can only think of ways to get a couple of them.
- dap 10y ago> The cost is acceptable when tracing and barely measurable when not tracing. "Acceptable" is obviously relative, but with DTrace's pid provider, the cost is zero when not tracing, and about the cost of a fastcall per probe point when enabled. > Instrumentation is automatic and directed towards functions that are important for understanding the binary’s execution time. I'm not sure what this means, but with DTrace, you enumerate the functions or binary objects (with wildcards and such) that you want to instrument, and the framework takes care of reliably instrumenting them, no matter the state of the process. Is that "automatic" and "directed"? I need to read the rest of the paper more closely. > Tracing is efficient in both space and time -- only recording what is required and what matters. DTrace records exactly what you ask it to. It supports in-situ aggregation for cases where it's not tenable to record a complete log of all interesting events. This is an important part of the design. > Tracing is configurable with thresholds for storage (how much memory to use) and accuracy (whether to log everything or only function calls taking at least some amount of time). With DTrace, it's pretty easy to filter on function execution time. The buffer size is configurable. There are also multiple buffer policies for different use-cases (e.g., ringbuffer of the last N events leading up to some other event). > Tracing does not require changes to the operating system nor super-user privileges. If they're running Linux, as I imagine they are, DTrace isn't necessarily an option. Several other platforms have just ported it. Using it to record user-level state on your own processes does not require superuser privileges. > Tracing can be turned on and off dynamically without having to restart the server. Absolutely -- that's what the "D" is for. I'd strongly recommended checking out the DTrace paper: https://www.usenix.org/legacy/event/usenix04/tech/general/full_papers/cantrill/cantrill_html/ https://www.usenix.org/legacy/event/usenix04/tech/general/fu... There may be good reasons not to use DTrace for this, but I'm not sure which of those six goals would be the sticking point other than OS availability. (edit: I also haven't read beyond that yet!)
- evmar 10y agoIt's definitely the sort of paper that would get reviewer feedback if it was submitted to a conference, in that it failed to compare any related work. (But it's unfair to grade it by that metric because it reads like more of a "here's what we did" blog post, not a research paper.) From reading it, I believe no other tool is both (1) non-sampled and (2) supports instrumenting all functions simultaneously. (At least none of the other comments in this thread point at tools that do this.)
- cthalupa 10y ago>From reading it, I believe no other tool is both (1) non-sampled and (2) supports instrumenting all functions simultaneously. (At least none of the other comments in this thread point at tools that do this.) Perf supports dynamic tracing of probepoints, and probing multiple points simultaneously. Perhaps I'm misunderstanding things, but I don't see how perf doesn't fit those two requirements. It does a lot more than just sampling
- DannyBee 10y ago1. Errr, perf definitely does not support inserting N probepoints, where N is the number of function entry + exit in a program. 2. Perf is 14x slower with tracepoints, and is definitely not that low overhead with tons and tons of probe points. See http://events.linuxfoundation.org/sites/events/files/slides/perf-collabsummit-2015.pdf http://events.linuxfoundation.org/sites/events/files/slides/...
- cthalupa 10y agoIs there something I'm missing in the XRay paper that mentions speed? The only overhead I can find mentioned in the paper is CPU and memory utilization. If XRay is able to trace millions of probe points with less overhead than uprobe + frontends, that's extremely impressive... But I just don't see any numbers so far that have said this is the case?
- DannyBee 10y ago
- DannyBee 10y agoSo, to be fair to the xray people, what they have done meets a very specific set of use cases you can't meet with most other tracing tools (I haven't looked at dtrace hard enough. at a glance, it looks like it would not meet several things, but dtrace is not a sane option for them for a variety of reasons). Past that, the mechanism used to do this, and the system itself, is not uncommon to build. I don't see any claims that say they think they've broken new ground. Just that they built it :)
- georgehm 10y agoSame idea, different use? https://blogs.msdn.microsoft.com/oldnewthing/20110921-00/?p=9583 https://blogs.msdn.microsoft.com/oldnewthing/20110921-00/?p=...
- kabdib 10y agoSeems unnecessary to do anything to registers in the call edge glue. We did this (oh, decades ago...) on the 68000, on the Macintosh, and really all we needed was a distilled link map and some return addresses. Maybe there are some complications we didn't have, though.
- 4ad 10y agoI agree. This technique is as old as programming itself.
- wslh 10y agoI am never tired of this shameless plug: our state of the art open source instrumentation engines for Microsoft Windows where you can hook applications without even knowing about the complexities of hooking. The most programmer friendly is https://github.com/nektra/Deviare2 https://github.com/nektra/Deviare2 while the Microsoft Detours competitor is: https://github.com/nektra/Deviare-InProc https://github.com/nektra/Deviare-InProc It has an embedded disassembler to smartly hook functions even if they have a jmp in the prologue.
- 05 10y agoYeah but what if they have a branch target in the prologue?
- Keyframe 10y agoYes.
- mxmauro 10y agoWhen the stub for calling the original function is created, most hooking engines assumes the prologue contains the standard "mov edi,edi/push ebp/mov ebp,esp" and it is wrong. If the prologue contains, for e.g., a relative jmp, copying opcodes is not enough. You must convert it to an absolute jump in the generated stub. The same applies to several instructions doing relative/indirect addressing.
- dang 10y agoUrl changed from https://research.google.com/pubs/pub45287.html https://research.google.com/pubs/pub45287.html, which points to this. This project was discussed recently at https://news.ycombinator.com/item?id=11595287 https://news.ycombinator.com/item?id=11595287, but perhaps the current post adds more information.
- brendangregg 10y agoNOT ANOTHER TRACER!! I'm sure it's impressive engineering work, but why oh why... How does it compare to Linux uprobes, which are built into Linux mainline? Bear in mind there are different front ends for uprobes (ftrace, perf_events, bcc, ...), and these are also still in development, so if one lacked certain features they needed, such features could be added. There's been a LOT of work in this area in the past 6 months, as well (see lkml). If the goal was lowest performance, then why compile with no-op sleds ("negligible overhead") instead of using dynamic tracing (literally "zero overhead")? Or, if the existing kernel-based dynamic tracers benchmarked poorly, then why not something like LTTng? How does it compare to DTrace, as well? (Doesn't Google have some FreeBSD?). All the tracers I mentioned can not only do dynamic tracing, but also instrument all user and kernel code, without special recompilation.
- deleted 10y ago[deleted]
- deadmutex 10y agoUnfortunately, I think the atleast one of developers was not aware of uprobes. I saw a post from the developer that said he was looking into uprobes after this point, and he was not aware of it. He seems like a nice guy, and just seems like an oversight.
- deleted 10y ago[deleted]
- moyix 10y agoIs there any documentation on how the dynamic tracing functionality actually works? In particular, how does it avoid the traditional problem with hooking via patching (i.e., that you have to get the disassembly exactly correct or you risk placing a hook on top of a jump target)?
- davidtgoldblatt 10y agoMost of XRay was written before uprobes was merged into the Linux kernel (and well before such kernels were widely available). I don't think any of the alternatives you mentioned are Pareto superior to XRay when considering all of "speed while tracing", "speed while not tracing", and "flexibility". E.g.: - In "speed while tracing", anything that takes a context switch per traced function will probably be dramatically slower. Even if there's some fast dispatch mechanism you have in mind that I'm not familiar with when you say dynamic tracing, if it doesn't insert the moral equivalent of a nop-sled, it will have to either choose between logging the whole PC (spending data, which means spending RAM and disk time) or figuring out how to map it to a function-specific unique int (spending cycles). - In "speed while not tracing", anything much more expensive than nop-sleds will be too slow to run in production. - Anything that doesn't have a compile time component probably won't be able to completely hook functions that get inlined, or whose source you aren't able to change, won't be able to pick out information the runtime wants to summarize from function arguments, etc. To me, the neat thing about XRay isn't so much the "function patching" aspect, except insofar as it serves as a mechanism to execute arbitrary code at function entry or exit in a way that's runtime-customizable and very low overhead when you want it to be.
- israrkhan 10y agoThe fact that it relies on compiler instrumentation makes it less interesting for people outside Google. Given that there are other instrumentation systems that do not require changes to compiler
- Mister_Snuggles 10y agoI would love to see a tracer that works across processes and machines. For example, a request hits the web server, the web server tickles an application server, the application server makes a number of database queries, and the results propagate back the way they came. I'd love to see something that could follow this request through all of the layers and across all of the machines, without the luxury of having the source code for most of the components. I don't see why this wouldn't be possible, but I certainly see why it wouldn't be EASY!
- ariwilson 10y agoe.g. Google Dapper? http://static.googleusercontent.com/media/research.google.com/en//pubs/archive/36356.pdf http://static.googleusercontent.com/media/research.google.co...
- kyrra 10y agoGoogle published a paper similar to this 6 years ago: http://research.google.com/pubs/pub36356.html http://research.google.com/pubs/pub36356.html Dapper will follow the call path of an RPC and report the details of it.
- bg451 10y agoSeeing that people have already mentioned Dapper, I'd add that there are already a few open source implementations of Dapper available. Appdash[1] is a very lightweight tracer that isn't too fancy, but gets the job done. There's also the much bigger, more popular tracer Zipkin[2], which was built by Twitter. [1] https://github.com/sourcegraph/appdash https://github.com/sourcegraph/appdash [2] https://github.com/openzipkin/zipkin https://github.com/openzipkin/zipkin
- Mister_Snuggles 10y agoIt looks like both of these require changes to the application. Dapper appears to have the same issue. My use case is for understanding a closed source application which has a number of separate pieces. Since I don't have the source code, these don't look like viable options.
- erikpukinskis 10y agoMore and more I think call tracing is a core process of programming, and am re-architecting the code I write to be easily traceable. I feel like this approach of using layer-cake architectures where you function call has to plumb through a dozen layers that you didn't write and then trying to make sense of that with data analytics is the wrong approach. Instead, I have been ditching layer-cake libraries for vertically integrated libraries that do one thing and take full responsibility for it. This requires architecting your application in a different way... it generally means more boilerplate. Libraries do the heavy lifting, but no sexy DSLs that turn your boilerplate into terse method chains and such. But the ability to simply put a breakpoint anywhere in the system and have the stack be a good representation of which pieces of code are doing something right now vs just hanging around because this-kind-of-thing might need that-kind-of-interface later on.
- compudj 10y agoI don't see any mention of Intel's errata on cross-modifying code on SMP in the paper. I wonder how the authors handle this ? See "Unsynchronized Cross-Modifying Code Operations Can Cause Unexpected Instruction Execution Results" Ref. http://www.intel.com.tr/content/dam/www/public/us/en/documents/specification-updates/xeon-5400-spec-update.pdf http://www.intel.com.tr/content/dam/www/public/us/en/documen... AX72. This is one of the main challenges to cross-modifying code, and one key reason why LTTng-UST does not use a nop-slide today. One possible approach to this is to SIGSTOP the entire process while doing the code modification, which is unwanted in real-time systems. Another approach would be to integrate with uprobes and do a temporary breakpoint bypass, similarly to what is done in the Linux kernel today for jump labels.