8 ms·
Why doesn't `kill -9` always work?
- Trufa 14y agoThat site is trying to murder my eyes! Go here http://www.readability.com/articles/zcqkmihi http://www.readability.com/articles/zcqkmihi and switch to Readability view!
- AYBABTME 14y agoI love this app: http://readable.tastefulwords.com/ http://readable.tastefulwords.com/ I've setup a bookmarklet using it; greatest thing since sliced bread.
- roryokane 14y agoI use a bookmarklet for that purpose – Zap Colors from https://www.squarefree.com/bookmarklets/zap.html https://www.squarefree.com/bookmarklets/zap.html. It resets colors to the browser default while keeping the page layout the same.
- glenstein 14y agoAnother (which I originally found because of a comment here a few years ago) is http://viewtext.org/ http://viewtext.org/. I like this one because I can add it as a custom search engine in Opera without adding a browser extension or leaving the page.
- aw3c2 14y agoopera has user style sheets built in. see the "author" and "user" mode. also check out the accessibility layout. no need to rely on a third party or even tell anyone what you are reading.
- a_bonobo 14y agoCan anyone explain the "Why is a process wedged?" part? I do understand that piping from /dev/random to /dev/null is going to run forever, but I do not understand the gdb-output, nor what that has to do with the rest of the text.
- jsnell 14y agoSo that whole section doesn't make a lot of sense to me. Running strace on an unkillable process tends to produce no output (if the process was actually making new system calls, it wouldn't be wedged). And worse, my experience is that attaching to a wedged process with ptrace() usually does nothing at all except also hang the attaching the process. This also applies to gdb. And finally, even if attaching worked, getting a disassembly of the current PC (in this case the syscall trampoline) would tell nothing useful about what's going on.
- agwa 14y agoI don't understand it either. The only syscalls going on there are reads from /dev/random and writes to /dev/null - both of those can be interrupted (in fact writes to /dev/null should be instantaneous). I think the author may be conflating applications blocked in system calls (since reads from /dev/random will block if the system lacks entropy) with applications blocked in uninterruptible system calls.
- syncsynchalt 14y agoThere are no reads from /dev/random successfully happening in the example (not after the first few blocks, anyway). /dev/random reads from an entropy pool that is quickly exhausted and slowly filled. The kernel will lock a process in D state while waiting for entropy. If you're expecting a stream of pseudo-random data then you can get that by directing /dev/urandom to /dev/null.
- agwa 14y ago> The kernel will lock a process in D state while waiting for entropy. No it doesn't, at least not in Linux 3.0.57 or any other kernel I can remember for the last many years. It blocks if it needs entropy, but it's interruptible meaning it's in the S state, not the D state, and can be killed.
- nonane 14y agoIt think one of the reasons the kernel can't kill a process is because one of the process's threads is blocked inside a kernel call (not completely sure about this). Hes using ps's 'wchan' option to get the address of the kernel function that the process is currently blocked or sleeping on. After he gets the wchan address, he uses gdb to map the address to function name. Taken from: http://unixhelp.ed.ac.uk/CGI/man-cgi?ps http://unixhelp.ed.ac.uk/CGI/man-cgi?ps nwchan WCHAN address of the kernel function where the process is sleeping (use wchan if you want the kernel function name). Running tasks will display a dash ('-') in this column.
- drivebyacct2 14y agosshfs used to have this problem and it was enough to bring nautilus and a lot of other applications to their knees as they tried to stat() my homedir and failed on the hung sshfs mount point.
- pyre 14y agoSame for the cifs.kext (or was is smbfs.kext) in early versions of OSX. Putting your laptop to sleep with a mounted Samba share was enough to slowly grind the system to a halt when you woke it up.
- agwa 14y agoDo not mount NFS with "soft" unless you really know what you're doing. NFS' behavior to hang is not "stupidity" - it's actually one of the best things about NFS. Applications do not deal well with failed reads/writes. If there's a brief network interruption or the NFS server goes down, it's WAY safer to cause applications to hang until the server comes back. Since NFS is a stateless protocol, when the server comes back, I/O resumes as if nothing ever happened. This helps make using NFS feel more like using a local filesystem. Otherwise it becomes a very leaky abstraction. Not being able to kill processes stuck on NFS I/O is annoying though, so you can mount with the "intr" option and that makes such processes killable. However, since Linux 2.6.25, you don't even need this and SIGKILL can always kill applications stuck in NFS I/O.
- NelsonMinar 14y agoThis would be a good time to re-read Waldo's "A Note on Distributed Computing", which points out how remote filesystems will never act like local filesystems. http://labs.oracle.com/techrep/1994/smli_tr-94-29.pdf http://labs.oracle.com/techrep/1994/smli_tr-94-29.pdf
- dredmorbius 14y agoExpecting NFS (or any other remote filesystem) to behave as if it were local is a fundamental error. Time, and speed of light, ultimately matter. If you need assurance, find a way of getting reliability in your system through redundancy and locality. Distinguish between "task has been delegated" and "task has been confirmed completed". Down any other path runs pain, and anyone who tells you otherwise is selling something. You're going to have to compromise: whole systems (or clusters) going titsup because your NFS heads had a fart, or lost commits. Neither is very attractive when shit's on the line. Your database relying on NFS is a fundamental error you'll have to design around.
- jrockway 14y agoNFS can be faster than local storage. Compare a very slow tape drive versus a very fast 10Gbps network connection to an NFS server with a huge in-memory cache. The problem is expecting any filesystem to be reliable.
- JimmaDaRustla 14y agoI always thought kill -9 won't always work because it currently has control of a system resource, like disk or something.
- dfox 14y agoThat is almost correct understanding. Processes that are waiting for things like disk I/O do not respond to any signals, not even KILL.
- ambrop7 14y agoI consider any uninterruptable sleep in the kernel a bug. There's no technical reason a process waiting for a resource (e.g. disk I/O) couldn't be killed on the spot, leaving the resource on its own. If it can't be, it just means it hasn't been implemented in the kernel.
- dfox 14y agoIt's bug motivated by compatibility. On original 70's implementations of Unix, file system I/O mostly led to busy wait in kernel and thus was not interruptible because it was simply not possible and there were applications that relied on this behavior. On UNIX, signal received during system call generally causes the kernel to abort whatever it was doing and requires application to deal with that situation and restart the operation, implementations of stdio in libc generally do the right thing, but most applications that do filesystem I/O directly do not (and surprisingly large number of commonly used network services behave erraticaly when network write(2) is interrupted by signal). And even applications that handle -EINTR from all I/O still have places where it is not handled (allowing interruptible disk I/O will cause things like stat(2) to return EINTR). Allowing SIGKILL to work and not any other signal is ugly special case, and while generally reasonable it is still special case that is relevant for things like NFS (with modern linux NFS client allowing you to disable this behavior) and broken hardware (and then trying to recover the situation with anything other than kernel-level debugger is mostly meaningless, with power cycling being the real solution when you can do that. Accidentally we currently have similar issue on one backend server where power-cycling is not an option).
- derleth 14y ago> [Fixing the bug] was simply not possible and there were applications that relied on this behavior [so now that it's technically possible to fix it, it would break compatibility so much we actually can't fix it anyway.] I'm beginning to wonder if it's possible to design a platform complex enough to be usable without running into this problem.
- asveikau 14y ago> implementations of stdio in libc generally do the right thing Tell me, what is the "right thing" for stdio to do when it sees EINTR? It strikes me that this can't really be solved at the library level. There are times when you'll want to retry and there are times when you'll want to drop your work and surface the error to the caller. Doesn't seem to me like a library can decide which is which. Which is probably why the I/O syscalls need to surface it in the first place. (I'd argue if a library like stdio, which does nothing but wrap syscalls and buffer stuff, can decide it, then there's no need for EINTR to exist at all because the syscall could theoretically make the same decisions.)
- nikster 14y agoMuch simpler answer: Bugs. If kill -9 does not work, its a bug. The kernel needs to be able to end processes no matter what the process is doing. By definition this should not be about how the misbehaving process was implemented. I imagine practical considerations are keeping these bugs in there, eg I can imagine the effort of making all processes killable would stand in no relation to the gains - its hard to do, and rare to occur,
- kqr2 14y agoSlightly off-topic, but kill -9 even has a rap song dedicated to it: http://www.youtube.com/watch?v=Fow7iUaKrq4 http://www.youtube.com/watch?v=Fow7iUaKrq4
- darwinGod 14y agoThe number of times I have spent 10 minutes staring at the output of 'pgrep processname' , when I had attached gdb to the process in another terminal session... Urgh!! :-/
- liotier 14y agoThat is not dead which can eternal lie And with strange aeons even death may die
- ucee054 14y agoph'nglui mglw'nafh Cthulhu R'lyeh wgah'nagl fhtagn
- kaeso 14y ago> ps Haxwwo pid,command | grep "rpciod" | grep -v grep pgrep(1) is there for a reason.