5 ms·
Desktop machine? Swap. For some reason you have a single server and don't care about performance? Swap. Running an application that might have a working set
by VectorLock 7y ago
Desktop machine? Swap. For some reason you have a single server and don't care about performance? Swap. Running an application that might have a working set thats larger than RAM and the application doesn't understand how to do its own disk paging? Swaps good there!
Larger scale systems with redundancy? No swap.
Having swap in systems like this still doesn't make sense to me. It treads heavily on the "cattle not pets" philosophy. I shouldn't be ssh-ing into a machine thats swapping to see whats up. It should be killed. One server in the cluster starts swapping and falls out of step with its peers? It should be killed. When a machine starts swapping it falls into a while different performance regime than the rest of your systems, now you've got more variance in your response times. Not good when you care about your response times. Unless you have memory-pretending-to-be-disk for swap (in which case why isn't it just memory)
I've never seen a machine 'act funny' because it didn't have swap, its always the other way around. I don't think I've ever encountered a machine that used so much memory that the kernel didn't have buffers, but not so much that it invoked OOM killer. Unless there was a woefully misconfigured process running on the machine.
If a machine is well utilized CPU wise it is going to get absolutely crushed when it starts swapping.
Time and time again I see swap being an issue. The past year I've been in a Large Scale shop which for some ungodly reason it has swap (nowhere I've been in the past 10 years as swap as a general rule)
Don't even get me started with EBS IOPS exhaustion when you start swapping onto an EBS volume.
- Tuna-Fish 7y ago> I don't think I've ever encountered a machine that used so much memory that the kernel didn't have buffers, but not so much that it invoked OOM killer. Note that the reason this topic is currently in vogue is that this has become a lot easier recently. If you run the system on a modern low-latency SSD, the current OOM killer algorithm often fails to kill anything before the entire system is on it's knees with approximately 0 pages left for IO and non-anonymous memory, at which point the OOM killer will never run because the machine is so thoroughly locked. The proper way to fix this of course is to make the OOM killer hit earlier.
- VectorLock 7y agoI would like OOM killer to be smarter and maybe easier to configure. I'm glad you can instrument it better with BPF now, at least.
- kevin_nisbet 7y agoI largely agree. Although I do find the scenario of the kernel evicting mmapped pages causing performance degradation to be interesting and sounds plausible, but I haven't personally witnessed this behavior. Where I see swap tend to get especially detrimental is with GCed processes. I've spent significant effort tracking down long GC pauses to getting blocked on swapped pages (although the software was not optimized and responsible as well and this was spinning rust). But in line with the article and your comments, this depends on engineering the system to have headroom. IIRC processes that use more than a NUMA node worth of memory also run into some issues with the OOM killer with swap disabled, unless set to interleaved on the NUMA policy. So that's another thing to look out for when dropping swap, although I forget exactly why it happens.
- mankyd 7y ago> Larger scale systems with redundancy? No swap. Why not give them swap, set off pagers, and _maybe_ kill them? There could still be something worth investigating there, and having swap will make that easier You also don't want to have a cascading failure where a massive leak makes all your machines fill their ram, and start killing everything like crazy.
- kevin_nisbet 7y agoThis is a plausible route, but it still requires some engineering, specifically tweaking swapiness setting. Otherwise the swap will get used even with plenty of memory available, which in my experience can still cause havoc for GCed processes with high allocation rates on non-ssd disks.
- VectorLock 7y agoWhy let them live? Why wake myself up? Now your swapping systems are introducing a performance degradation.
- mankyd 7y agoQuote: "maybe kill some of them" Quote: "You also don't want to have a cascading failure where a massive leak makes all your machines fill their ram, and start killing everything like crazy." Cascading failures are a very real thing that have knocked whole systems offline. It sounds like the real solution is a balanced solution involving some engineering: kill them if you aren't killing _everything_. Page if the problem is ongoing, not if a couple of machines have a problem. Either way, you can add swap _and_ kill them. One does not preclude the other.