20 ms·
Linux Crisis Tools
- prydt 3y agoLove the list and the eBPF tools look super helpful.
- FridgeSeal 3y agoThis is a handy list. > 4:07pm The package install has failed as it can't resolve the repositories. Something is wrong with the /etc/apt configuration… Cloud definitely has downsides, and isn’t a fit for all scenarios but in my experience it’s great for situations like this. Instead of messing around trying to repair it, simply kill the machine, or take it out of the pool. Get a new one. New machine and app likely comes up clean. Incident resolves. Dig into machine off the hot path.
- throw5323446 3y ago> Instead of messing around trying to repair it, simply kill the machine, or take it out of the pool. Get a new one. "4:10pm the new machine still has the same performance issue"
- FridgeSeal 3y agoSure, but more often than not - esp in cloud scenarios, sometimes you just get a machine that is having a bad day and it’s quicker to just eject it, let the rest of the infra pick up the slack, and then debug from there. Additionally if you’ve axed a machine, and got the same issue, you know it’s not a machine issue, so either go look at your networking layer or whatever configs you’re using to boot your machines from…
- tjoff 3y ago> esp in cloud scenarios ... so the nice thing about the about the cloud is that you can workaround cloud-specific issues?
- deleted 3y ago[deleted]
- jandrese 3y ago4:20pm Turns out it was DNS
- Propelloni 3y agoThat made me laugh. Thank you. Of course, it is not DNS. DNS has become the new cabling. DNS is not especially complicated, but cabling is neither. Yet, during dot.com and subsequent years the cabling was causing a lot of the problems so that we get used to first check the cabling. But it only took a few more years to realize that it is not always cabling, actually failures are normally distributed. Is it wrong to check DNS first? No, but please realize that DNS misconfiguration is not more common than other SNAFUS.
- ninkendo 3y agoIt’s not DNS There’s no way it’s DNS It was DNS
- smackeyacky 3y agoCertificates are the new DNS for service breakages
- SerCe 3y agoThat's actually amazing, a reproducible problem is a 90% solved problem!
- Jedd 3y agoYou're describing one of the benefits of virtualised cattle, not necessarily or exclusively 'cloud'.
- userbinator 3y agoDig into machine off the hot path. Unfortunately, no one has the time to do that (or let someone do it) after the problem is "solved", so over time the "rebuild from scratch" approach just results in a loss of actual troubleshooting skills and acquired knowledge --- the software equivalent of a "parts swapper" in the physical world.
- patrick451 3y agoThe end state of a culture that embraces restart/reboot/clear-cache instead of real diagnoses and troubleshooting is a cohort of junior devs who just delete their git repo and reclone instead of figuring out what a detached HEAD is. I don't really fault the junior dev who does that. They are just following the "I don't understand something, so just start over" paradigm set by seniors.
- yosefk 3y agoTo be fair, with git, specifically, it's a good idea to at least clone for backup before things like major merges. There are lots of horror stories from people losing work to git workflow issues and I'd rather be ridiculed as an idiot who is afraid of "his tools" (as if I have anything like a choice when using git) and won't learn them properly than lose work thanks to a belief that this thing behaves in a way which can actually be learned and followed safely. A special case of this is git rebase after which you "can" access the original history in some obscure way until it's garbage-collected; or you could clone the repo before the merge and then you can access the original history straightforwardly and you decide when to garbage-collect it by deleting that repo.
- theptip 3y agoGit is a lot less scary when you understand the reflog; commit or stash your local changes and then you can rebase without fear of losing anything. (As a bonus tip, place “mybranch.bak” branches as pointers to your pre-rebase commit sha to avoid having to dig around in the reflog at all.) I would never ridicule anyone for your approach, just gently encourage them to spend a few mins to grok the ‘git reflog’ command.
- KingOfCoders 3y agoKill the machine might destroy evidence. It might be the case you have everything logged outside, but most often there is something missing.
- monkpit 3y agoTake it out of the pool then.
- pstuart 3y agoSounds like it's time to create a crisis-essential package group a la build-essential.
- yjftsjthsd-h 3y agoI have in the past created a package list in ansible/salt/chef/... called devops_tools or whatever to make sure we had all the tools installed ahead of time.
- kunley 3y agoBrendan Gregg as always with down to earth approach. Love the warroom example
- SamuelAdams 3y agoWould these tools still be useful in a cloud environment, such as EC2? Most dev teams I work with are actively reducing their actual managed server and replace it with either Lambda, or docker images running in K8. I wonder if these tools are still useful for containers and serverless?
- mdekkers 3y ago> Most dev teams I work with are actively reducing their actual managed server and replace it with either Lambda, or docker images running in K8. There are plenty of services that don’t fit on k8s or Lambda. Not all pegs fit in those holes.
- yla92 3y agoIt's still useful in EC2 (or any other VM-based environments) and Docker containers, as long as you can install the necessary packages (if they are not installed by default). Because after all, there are "servers" underneath, even for the serverless apps, I suppose. It's definitely harder for apps running in Lambda because we may not have access to the underlying OS. In such case, I kind of fallback to using the application level observability tools like Pyroscope (https://pyroscope.io https://pyroscope.io). It doesn't always work for all the cases and have some overheads/set up but it's still better than flying and more useful than the Cloud Provider's provided metrics.
- ranger207 3y agoIME there's always that one service that wasn't ever migrated to containers or lambdas is is off running on an EC2 somewhere, and nobody knows about it because it never breaks, but then the one time AWS schedules an instance retirement for it...
- cpuguy83 3y agoContainers are just processes running on the host where the process has a different view of the world from the "host". The host can see all and do all.
- reilly3000 3y agoIn such a crisis if installing tools is impossible, you can run many utils via Docker, such as: Build a container with a one-liner: docker build -t tcpdump - <<EOF \nFROM ubuntu \nRUN apt-get update && apt-get install -y tcpdump \nCMD tcpdump -i eth0 \nEOF Run attached to the host network: docker run -dP --net=host moremagic/docker-netstat Run system tools attached to read host processes: for sysstat_tool in iostat sar vmstat mpstat pidstat; do alias "sysstat-${sysstat_tool}=docker run --rm -it -v /proc:/proc --privileged --net host --pid host ghcr.io/krishjainx/sysstat-docker:main /usr/bin/${sysstat_tool}" done unset -v sysstat_tool Sure, yum install is preferred, but so long as docker is available this is a viable alternative if you can manage the extra mapping needed. It probably wouldn’t work with a rootless/podman setup.
- blueflow 3y agoIs there a situation where apt cant download and install packages but docker can fetch new containers? apt libs borked or something?
- Smar 3y agoI would just decompress the .deb in such case. As a last resort, even a .rpm might work. Of course handling dependencies by hand is annoying, but depending on situation it might be faster anyway.
- xyst 3y agoUnless you are in an air gapped situation. Good luck pulling “Ubuntu” image!
- supriyo-biswas 3y agoOn that note I'd largely prefer if `busybox` contained more of these tools, it'd be very helpful to have a 1MBish file that I can upload into a server and run it there.
- mmh0000 3y agoI was surprised that `strace` wasn't on that list. That's usually one of my first go-to tools. It's so great, especially when programs return useless or wrong error messages.
- brendangregg 3y agostrace is ok as a last resort, but "perf trace" and bpf tracing tools are the production-safe alternative. https://www.brendangregg.com/blog/2014-05-11/strace-wow-much-syscall.html https://www.brendangregg.com/blog/2014-05-11/strace-wow-much...
- slacka 3y agoWhy don't recommend atop? When a system is unresponsive, I want a I want a high-level tool that immediately shows which subsystem is under heavy load. It should show CPU, Memory, Disk, and Network usage. The other tools you listed are great, once you know what the cause is.
- brendangregg 3y agoMy preference is tools that give a rolling output as it let you capture the time-based pattern and share it with others, including in JIRA tickets and SRE chatrooms, whereas top's generally clear the screen. atop by default also sets up logging and runs a couple of daemons in systemd, so it's more than just a handy tool when needed, it's now adding itself to the operating table. (I think I did at least one blog post about performance monitoring agents causing performance issues.) Just something to consider. I've recommended atop in the past for catching short-lived processes because it uses process accounting, although the newer bpf tools provide more detail.
- CartwheelLinux 3y ago[dead]
- vram22 3y agofuser and lsof are useful too. https://man7.org/linux/man-pages/man1/fuser.1.html https://man7.org/linux/man-pages/man1/fuser.1.html https://en.m.wikipedia.org/wiki/Lsof https://en.m.wikipedia.org/wiki/Lsof
- SuperHeavy256 3y agoSo basically busybox?
- Linda231 3y ago[dead]
- randomgiy3142 3y agoI use zfsbootmenu with hrmph (https://github.com/leahneukirchen/hrmpf https://github.com/leahneukirchen/hrmpf). You can see the list of packages here (https://github.com/leahneukirchen/hrmpf/blob/master/hrmpf.packages https://github.com/leahneukirchen/hrmpf/blob/master/hrmpf.pa...). I usually build images based off this so they’re all there, otherwise you’ll need to ssh into zfsbootmenu and load the 2 gb separate distro. This is for home server, though if I had a startup I’d probably setup a “cloud setup” and throw a bunch of servers somewhere. A lot of times for internal projects and even non-production client research having your own cluster is a lot cheaper and easier then paying for a cloud provider. It also gets around when you can’t run k8s and need bare metal. I’d advised some clients on this setup with contingencies in case of catastrophic failure and more importantly test those contingencies but this is more so you don’t have developers doing nothing not to prevent overnight outages. A lot cheaper than cloud solutions for non critical projects and while larger companies will look at the numbers closely if something happened and devs can’t work for an hour the advantage of a startup is devs will find a way to be productive locally or simply have them take the afternoon off (neither has happened). I imagine these problems described happen on big iron type hardware clusters that are extremely expensive and spare capacity isn’t possible. I might be wrong but especially with (sigh) AI setups with extremely expensive $30k GPUs and crazy bandwidth between planes you buy from IBM for crazy prices (hardware vendor on the line so quickly was a hint) you’re way past the commodity server cloud model. I have no idea what could go wrong with such equipment where nearly ever piece of hardware is close to custom built but I’m glad I don’t have to deal with that. The debugging on those things work hardware only a few huge pharma or research companies use has to come down to really strange things.
- semi-extrinsic 3y agoOn compute clusters there are quite a few "exotic" things that can go wrong. The workload orchestration is typically SLURM, which can throw errors and has a million config options to get lost in. Then you have storage, often tiered in three levels - job-temporary scratch storage on each node, a distributed fast storage with a few weeks retention only, and an external permanent storage attached somehow. Relatively often the middle layer here, which is Lustre or something similar, can throw a fit. Then you have the interconnect, which can be anything from super flakey to rock solid. I've seen fifteen year old setups be rock solid, and in one extreme example a brand new system that was so unstable, all the IB cards were shipped back to Mellanox and replaced under warranty with a previous generation model. This type of thing usually follows something like a Weibull distribution, where wrinkles are ironed out over time and the IB drivers become more robust for a particular HW model. Then you have the general hardware and drivers on each node. Typically there is extensive performance testing to establish the best compiler flags etc., as well as how to distribute the work most optimally for a given workload. Failures on this level are easier in the sense that it typically just affects a couple of nodes which you can take offline and fix while the rest keep running.
- josephcsible 3y ago> and...permission errors. What!? I'm root, this makes no sense. This is one of the reasons why I fight back as hard as I can against any "security" measures that restrict what root can do.
- rr808 3y agoYou guys get root access? I have to raise a ticket for a sysadmin to do anything.
- zer00eyz 3y agoI am a consultant now so it's a new company every few months. There are groups of people you always make nice with. * Security people. The kinds with poorly fit blazers who let you into the building. Learn these peoples names, Starbucks cards are your friends. * Cleaning people. Be nice, be polite, again learn names. Your area will be spoltless. It's worth staying late every now and again just to get to know these folks. * Accounting: Make some friends here. Get coffee, go to lunch, talk to them about non work shit, ask about their job, show interest. If you pick the right ones they are gonna grab you when layoffs are coming or corp money is flowing (hit your boss up for extra money times). * IT. The folks who hand out laptops, manage email. Be nice to these people. Watch how quickly they rip bullshit off your computer or wave some security nonsense. Be first in line for every upgrade possible. * Sysadmins. These are the most important ones. Not just because "root" but because a good SA knows how to code but never says it out loud. A good sysadmin will tell you what dark corners have the bodies and if it's just a closet or a whole fucking cemetery. If you learn to build into their platform (hint for them containers are how they isolate your shitty software in most cases) then you're going to get a LOT more leeway. This is the one group of people who will ask you for favors and you should do them.
- Propelloni 3y agoSo true. If you want to know anything about an office, ask the sysadmins. Double-plus on being nice to the facility managers, cleaning people and security. Not only do they do a thankless job but they are often the most useful and resourceful people around if you need something taken care of. They know how to get shit done.
- rkachowski 3y ago> Starbucks cards are your friends like, how? are you straight up bribing people with coffee for security favors? or is it like, "hey man, thanks for helping me out I'd like to buy you a coffee but I'm busy with secret consulting stuff - here's a gift card" Is this something that only works for short lived external consultant interactions?
- zer00eyz 3y agoThe only thing I would add is nmap. Network connectivity issues aren't always apparent in some apps.
- sneak 3y agoscreen/tmux byobu pv rsync and of course vim.
- vram22 3y agodd, echo * as a poor man's ls if ls is accidentally deleted, busybox, cpio, fsck and fsdb. Used all of these and more, in Unix, not just Linux crisis situations.
- sneak 3y agoThose are already there. We are talking about diagnostic and recovery tools that should be installed by policy, in advance, so that they are already in place to aid in emergencies.
- vram22 3y agoOkay, my mistake. But busybox is not always already there, right? Installed, I mean? Not at a box right now.
- bostik 3y agoBusybox has one big downside: the tools it provides tend to have a rather ... limited set of options available. The easy stuff you can do in a standard shell might not be supported.
- devsda 3y agoNot all servers are containerized, but a significant number are and they present their own challenges. Unfortunately, many such tools in docker images will be flagged by automated security scanning tools in the "unnecessary tools that can aid an attacker in observing and modifying system behavior" category. Some of those ( like having gdb) are valid concerns but many are not. To avoid that we have some of these tools in a separate volume as (preferably) static binaries or compile & install them with the mount path as the install prefix (for config files & libs). If there's need to debug, we ask operations to mount the volume temporarily as read-only. Another challenge is if there's a debug tool that requires enabling a certain kernel feature, there are often questions/concerns about how that affects other containers running on the same host.
- Too 3y agoA better way is to build a second image including the debug tools and a root-user, then start it with the prod-containers pid-namespace and network-namespace mounted. Starting a second container is usually a good idea anyway, since you need to add a lot of extra flags like SYS_PTRACE capability, user 0 and --privileged for debuggers to work. This way you don't need to restart the prod-container either, potentially loosing reproduction-evidence. Remembering how to do all this in an emergency may not be entirely obvious. Make sure to try it first and write down the steps in your run books.
- devsda 3y ago> A better way is to build a second image including the debug tools and a root-user. That was our initial idea. But management and QA are paranoid enough that they consider these as new set of images that require running the complete test suite again even when they are built on top of certified images. Nobody is willing to test twice, so we had to settle for this middle.
- remram 3y agoIf an attacker can execute files from the filesystem, and all that's missing to run them is them being present on the filesystem, the attacker could just... write those files themselves? I really don't understand in what scenario this policy makes any sense, apart from "my organization misuses security scanners".
- donio 3y agoI always cover such tools when I interview people for SRE-type positions. Not so much about which specific commands the candidate can recall (although it always impresses when somebody teaches me about a new tool) but what's possible, what sort of tools are available and how you use them: that you can capture and analyze network traffic, syscalls, execution profiles and examine OS and hardware state.
- keepamovin 3y ago[flagged]
- b112 3y agoThis is a good example of why chatGPT is just useless for so many things. * Why even create a use-time script, when the whole point of the original article is how these tools should be pre-installed, because the system may be in a state where disk, or network io, or load, prevents easy install of anything, during times of instability/issues * Why bother with a 100+ line script, when you just "apt-get install thing" then use the blasted thing * Who's going to maintain this script down the road? Why are things on it? Why even spend time looking at it, debugging it? It's literally creating extra work, whilst doing nothing useful. Please people. If you don't understand what's going on, or why things are to be done, don't hand it off to chatGPT. Just learn, learn, learn or if you can't/won't learn, go find a job in a different field.
- keepamovin 3y agoAnd yet, so useful it is, despite flagging. You can have it all in one place, and have Chat GPT write it for you, so you don't have to type it all out. And you've saved yourself time. Is time not important to you? Hahahahaha! :)
- keepamovin 3y agoWow, that's really mean and narrow response. I get if you're mad about something right now that has nothing to do with this, but please don't bring that here. Rather constructively deal with the source of your feelings, instead of trying to take it out on people and places that don't deserve it! Haha :) In your quest for invective did you consider other possibilities? Perhaps: - you could learn by using these things, with an easy interface - perhaps you're not the expert who knows everything, and other people have use cases beyond what you'd consider? - you don't have to wait until things go wrong to familiarize yourself with these things, you can use the script now to ensure they're installed Personally I think that if you don't think ChatGPT is useful for very many things, you're not using it right! Haha :) In general it seems your comment would benefit from the HN guidelines advice of assuming a generous interpretation, rather than going in other direction! Haha! :) If you want to find "uselessness" or "stupidity" in what you are looking at, well, surely you can always find it if you try. But, truth is, you find there what you bring to it. So, try harder to find something good! Haha :) But also that comment itself could be an example of "not understanding what's going on" yet speaking like there there is that understanding. I understand if the job situation is difficult and you're unhappy with the number of people in the field now, but if you dislike appreciating a diversity of approaches beyond your own, perhaps it is you who ought "go find a job in a different field." Or, learn to be more welcoming and less arrogant in your quest to be "right". You could just apply yourself to finding something useful in whatever you're responding to. That, it seems, would be the way to really be right, and also to make a useful contribution, not just to the forum, but with that attitude, to the field. :)
- sargun 3y agoWhen I was at Netflix, Brendan and his team made sure that we had a fair set of debugging tools installed everywhere (bpftrace, bcc, working perf) These were a lifesaver multiple times.
- deleted 3y ago[deleted]
- logifail 3y agoDoesn't one increase a system's attack surface area/privilege escalation risk by pre-installing tools such as these?
- c0l0 3y agoUsually (not by design, but by circumstance), if someone gains RCE on your systems, they can also find a way to bring the tools they need to do whatever they originally set out to do. It's the old "I don't want to have a compiler installed on my system, that's dangerous, unnecessary software!"-trope driven to a new extreme. Unless the executables installed are a means to somehow escalate privileges (via setuid, file-based capabilities, a too-open sudo policy, ...), having them installed might be a convenience for a successful attacker - but very rarely the singular inflection point at which their attempted attack became a successful one. The times I've been locked in an ill-equipped container image that was stripped bare by some "security" crapware and/or guidelines and that made debugging a problem MUCH harder than it should have been vastly outnumber the times where I've had to deal with a security incident because someone had coreutils (or w/e) "unnecessarily" installed. (The latter tally is at zero, for the record.)
- logifail 3y ago> It's the old "I don't want to have a compiler installed on my system, that's dangerous, unnecessary software!"-trope (This is a genuine question). In what circumstances would you need to install/run dev tools in prod? Of course having a compiler installed isn't necessarily an issue... but it might well be a sign that there is an underlying problem! (FWIW, I used to build everything from source. Yes, also in prod. That was a while ago...)
- citrin_ru 3y agoHow do you see an escalation using one of listed in the article tool (unless a binary has suid bit which you shouldn’t set if worried about security). Many of these tools provide convenient access to /proc - if an attacker needs something there they can read/write directly to /proc. Though in case of eBPF - disabled kernel support would reduce attack surface and if it disabled in the kernel’s user mode tools are useless.
- sirwitti 3y agoRelated to that, I recently learned about safe-rm which lets you configure files and directories that can't be deleted. This probably would have prevented a stressful incident 3 weeks ago.
- kureikain 3y agoI don't see nmap, netstat, and nc being mention. They had saved me so many time as well.
- pjmlp 3y agoThe list is great, but only for classical server workloads. Usually not even a shell is available in modern Kubernetes deployments that take a security first approach, with chiseled containers. And by creating a debugging image, not only is the execution environment being changed, deploying it might require disabling security policies doing image scans.
- ur-whale 3y agoCan't imagine handling a Linux crisis without ssh [EDIT]: typo
- infofarmer 3y agosomewhat related: /rescue/* on every FreeBSD system since 5.2 (2004) — a single statically linked ~17MB binary combining ~150 critical tools, hardlinked under their usual names https://man.freebsd.org/cgi/man.cgi?rescue https://man.freebsd.org/cgi/man.cgi?rescue https://github.com/freebsd/freebsd-src/blob/main/rescue/rescue/Makefile https://github.com/freebsd/freebsd-src/blob/main/rescue/resc...
- washadjeffmad 3y agoAnd I haven't needed to use it in fifteen years. Over the past four or five years, I've ported what I can to a *BSD, for sanity reasons.
- js4ever 3y agoLet's add NCDU to the list, it's super usefull to find what is taking all the disk space
- anthk 3y agotmux, statically linked (musl) busybox with everything, lsof, ltrace/strace and a few more. Under OpenBSD this is not an issue as you have systat and friends in base.
- michaelhoffman 3y agoWhen would you need to use rdmsr and wrmsr in a crisis?