3 ms·
Except, you don't necessarily know that it's the database that's the problem. It might be trying to do a DNS query, and that's timing out. Or spinning in a loop
by lambda 12y ago
Except, you don't necessarily know that it's the database that's the problem. It might be trying to do a DNS query, and that's timing out. Or spinning in a loop trying to read from a socket that has no data available. Or blocking on a filesystem operation on a networked filesystem that has become disconnected. Or any number of other things.
The nice thing about strace is that it pretty quickly tells you exactly what the program in question is doing with the outside world, which many times can provide you the quick clue you need about what's going on.
strace is a pretty simple tool, at its most basic usage it's just "strace <pid>", and the nice thing is that it's pretty universal; it comes in handy many, many times, for solving many different kinds of problems. Have some third party code that seems to exhibit different behavior on different machines, but don't know what is different about its environment that's making it behave differently? strace it and see what files it opens for configuration data. Your program not making any progress, and it doesn't seem to be CPU bound as it's only using a few percent of CPU? strace it and see what system calls its blocking on.
Note that for this kind of use case, where the process is hanging, you generally don't get way more output than you can handle; the fact that it's hanging on this particular call means that that's the call that you will see at the end of the output for several seconds until it times out.
It's certainly not the only tool in your arsenal; there will be other cases where more specific tools are more appropriate. But it's a really useful general purpose tool.
My general purpose debugging arsenal, when I need to debug something that I haven't written, generally consists of tcpdump/wireshark to capture and analyze any network traffic, strace to see how the program is interacting with the world outside of it, turning up logging on any relevant systems to their maximum value and looking through that output, and then finally if all else fails attaching to it with GDB and stepping through the code.
Sure, it's always better to have more specific tools, to have turned on appropriate monitoring, and so on. But you don't always have that luxury. Sometimes more specific tools haven't been written. Sometimes your monitoring may find that your database is up, as it is accepting connections, but its actually hanging when you try to do anything, or it's succeeding for reads but hanging on writes, or the like. Monitoring code is never going to be as thorough as you like, there will always be something that you miss, and so you need general-purpose tools that can be used in any situation to get a handle on what is going on quickly so that you can narrow down on the problem.