3 ms·
It's possible for the OS to recognize that several pages have the same content (ie: doing a hash) and then de-duplicate them. This can happen across multiple ap
by pradn 2y ago
It's possible for the OS to recognize that several pages have the same content (ie: doing a hash) and then de-duplicate them. This can happen across multiple applications. It's easiest for read-only pages, but you can swing it for mutable pages as well. You just have to copy the page on the first write (ie: copy-on-write).
I don't know which OSs do this, but I know hypervisors certainly do this across multiple VMs.
- deleted 2y ago[deleted]
- ChocolateGod 2y agoCorrect me if I'm wrong Linux only supports KSM (memory-deduping) between processes when doing it between VMs, as QEMU provides information to the kernel to perform it.
- yjftsjthsd-h 2y agohttps://www.kernel.org/doc/html/latest/admin-guide/mm/ksm.html https://www.kernel.org/doc/html/latest/admin-guide/mm/ksm.ht... > KSM was originally developed for use with KVM (where it was known as Kernel Shared Memory), to fit more virtual machines into physical memory, by sharing the data common between them. But it can be useful to any application which generates many instances of the same data Although... > KSM only operates on those areas of address space which an application has advised to be likely candidates for merging, by using the madvise(2) system call
- pradn 2y agoI wonder if you could just madvise the entire address space. My hunch is that it's for performance reasons only - fewer pages to scan, hash, and de-duplicate.
- MrDrMcCoy 2y agoYou can. There are custom kernel builds that do this, as well as shim loaders that do that madvise call. However, those are considerably less practical than the recent systemd feature that allows you to do it per unit. If applied to the system scope, it effectively covers the whole system. However, there are caveats: 1. You need a much better newer systemd than your distro likely packages. 2. KSM dedupe scans aren't free. Your system can spend its time doing that scan or it can spend its doing work with duplicate pages. Only relatively idle or highly homogenous systems would be free of the penalty. 3. For applications, especially statically linked ones, the duplicated code is not super likely to fall along page boundaries, thus the actual detectable duplication will be relatively low. That said, it's still great for densely deploying a high traffic microservice that's more memory than CPU bound.
- slabity 2y agoEven if the OS could perfectly deduplicate pages based on their contents, static linking doesn't guarantee identical pages across applications. Programs may include different subsets of library functions and the linker can throw out unused ones. Library code isn't necessarily aligned consistently across programs or the pages. And if you're doing any sort of LTO then that can change function behavior, inlining, and code layout. It's unlikely for the OS to effectively deduplicate memory pages from statically linked libraries across different applications.
- pradn 2y agoAh, good to know! Thank you for explaining. I guess much of this is why its hard to use shared libraries in the first place.