3 ms·
>It can take 10+, 30+ minutes for systems encountering this to resolve to some meaningful conclusion, and half the time, I'm desperately trying to ssh in so-as
by VectorLock 7y ago
>It can take 10+, 30+ minutes for systems encountering this to resolve to some meaningful conclusion, and half the time, I'm desperately trying to ssh in so-as to kill -9 the errant task anyway but ssh is paged out, and I wish the OOM killer would just do it for me instead of Linux trying to page everything through what feels like a single 4KiB page.
Sounds like you have a redundancy problem not a swap problem. If should just be able to kill a machine that gets into a bad way like that and move on. What if it wasn't swap but one of the million other things that could make your server crawl?
- deathanatos 7y agoGenerally speaking, swap/page thrashing is very easy to pick out from the metrics on a VM. (We use a system called Prometheus[1], which records and transmits metrics to a certralized service.) In particular, a machine that is swap/page thrashing will generally show as having no available RAM, a high amount (especially relative to baseline) of major (required disk reads) page faults, and often the CPU profile will be spending a lot of time in I/O wait, too, I think, though I usually just use out of RAM + page faulting. Also, the metrics service tends to go dark shortly afterwards — it's having the same issue as everything else on the VM at getting CPU time. Major page faults, perhaps the key stat for "this machine is page thrashing" since it directly corresponds to it, is found in /proc/vmstat and is called "pgmajfault". Though like I said, we generally had Prometheus and Grafana to turn these into pretty graphs, and to export them out of the VM itself, since when something is page thrashing, getting it to do anything is hard. CPU contention lacks the "out of RAM" part, and won't knock the metrics offline. Network contention can knock out the metrics, but often doesn't, and lacks the other signals: out of RAM/page faults. Disk I/O lacks the out of RAM & doesn't knock the metrics out since they don't require (beyond being paged in) disk I/O. (And those — CPU, RAM, network, and disk-ish — are about the only real resource dimensions on a VM.) Alternatively, if you let it play out and the VM eventually recovers, it might nonetheless decide to OOM kill a thing or to along the way, and those show up in dmesg / on the console, in you can get to those. > If should just be able to kill a machine that gets into a bad way like that and move on. I must admit that the production I love and cared for was not always perfect. Often, yes, I could, but there were a few spots where things weren't so rosy. Even when I could, I generally wanted to have some understanding as to why the VM went under, so as to not have the problem come back again, later, on a different VM. Page thrashing, in particular, is basically always symptomatic of a bug. And in distributed systems, the bugs are also distributed. The economics of getting devs enough time to develop to that quality vs. management wanting new features has been one of the hardest challenges of my career. Also for a while I lacked permission to actually kill a machine because ops was locking permissions down because "devs shouldn't have access to the actual machines"; like she says in one of the linked posts, fine, have my pager, you deal with the pages. [1]: https://prometheus.io/ https://prometheus.io/