7 ms·
CVE-2026-31431: Copy Fail vs. rootless containers
- hackeman300 5mo agoIt's a shame, this seems like an interesting topic but I can't get past the blatant AI-isms littered throughout. >This is not raw shellcode — it is a fully formed ELF executable
- washbasin 5mo agoPlease post a tl;dr at the top or even in the subject. Many of us are scrambling to patch/reboot our **.
- nullsanity 5mo ago[dead]
- donaldjbiden 5mo agoThis isn't a new CVE. It's just documenting what happened when this person ran the exploit inside a certain type of container.
- isityettime 5mo agoIt already has a table of contents. The heading titled "why rootless containers stopped the escalation" is your tl;dr.
- PunchyHamster 5mo agotl;dr (not from article) echo -e 'install algif_aead /bin/false\n' > /etc/modprobe.d/disable-algif.conf that just prevents the faulty module from loading. So you have time to fix it properly (kernel upgrade) Technically there should be zero impact (the very very few tools that use it will fall back to userspace), I haven't even found that module loaded in infrastructure Then check if it is loaded, and if it is, unload/reboot
- chrisss395 5mo agoDumb question: is preventing the module from loading safe to blindly run on, e.g., Unraid, Proxmox, WSL2? Is it possible to break anything?
- cpach 5mo agoI would say any sanely written application would fall back to doing the requested operations in userspace if it cannot use the AF_ALG socket. It could fail though. But I have not yet heard of anyone noticing big problems due to disabling the problematic modules. And I have not noticed any such issues on our systems at ${DAYJOB}. IMHO, since these parts of the Linux kernel are so crappy I personally would say disabling them is a good default choice. YMMV. But if you encounter problems, then you can always re-enable the modules. (Preferably after upgrading your kernel, obviously.)
- PunchyHamster 5mo agocheck if module is loaded. if it isn't nothing is using it and you can safely add it. I'd also imagine most software doesn't fail but just use userspace lib
- mjmas 5mo agoThough this won't work for some kernels: If algif_aead was a builtin module, it needs to be disabled by adding initcall_blacklist=algif_aead_init to the boot cmdline. However initcall_blacklist requires the kernel to be built with CONFIG_KALLSYMS.
- atmosx 5mo agotl;dr: switch to podman :-) or (for docker, not mention in the post but...) just `allowPrivilegeEscalation=False` in the deployment's SCC and you'll be fine at the pod level. Most deployments don't need priv escalation anyway, the ones that do need to either limits perms through capabilities or make sure the node (meaning the kernel) is patched.
- cpach 5mo agoHow does allowPrivilegeEscalation=False help?
- atmosx 5mo agoHave you tested running the PoC in a pod with and without proviEsc set?
- cpach 5mo agoNo, I haven’t. My concern is to try to understand the mechanisms of the exploit. Copy Fail is not simply ”hey, kernel, give me root”. I would say it’s more general than that. It’s rather: ”Hey, kernel, when you present file /foo to a process, make the contents of that file appear according to my wishes”. Which can be used (in various ways) to advance the attacker’s position. That’s why I think it’s interesting to ponder if that power allows the attacker to simply sneak past security policies such as allowPrivilegeEscalation=false.
- eqvinox 5mo agoRunning sstrip on an ELF binary is called ELF "golfing"? TIL…
- Retr0id 5mo agoIt is, although real ELF golfers consider that a little naive.
- repelsteeltje 5mo agoSorry for posting a n00b question, but could you share etymology on this term golfing?
- Retr0id 5mo agoIn golf, lower scores are better.
- seanhunter 5mo agoWhich, if you think about it, means the objective of golf is to be the person who plays the least golf.
- mbreese 5mo agoIt’s manipulating the binary to make it as small as possible. In golf, the lowest score wins. So, in this context, the smallest binary that still works wins.
- jerf 5mo agohttps://www.perlmonks.org/index.pl?replies=1;node_id=437032;displaytype=print https://www.perlmonks.org/index.pl?replies=1;node_id=437032;... As befits a history of perl, it is full of random quotes and rambling discourses about history, but it has a lot of info in it.
- eqvinox 5mo agoIt does feel a little simplistic to get a special name. But lesser things have gotten fancier names...
- QuietLedge375 5mo ago[dead]
- 2bitencryption 5mo agotl;dr - within the container, the exploit works, and elevates to root (uid 0) within the container - BUT because that namespace actually maps to uid 1000 (the user) outside the container, the escalation does not flow up to the host. But… does this escape the container? If not (the author seems to indicate it does not) then does it matter if you are in Docker or rootless Podman, right, since the end result is always: you have elevated to root within the container. If the rest of the container filesystem isolation does its job, the end result is the same? Though I guess another chained exploit to escape the container would be worse in Docker? Do I have that right?
- firesteelrain 5mo agoThis is a problem and most people hadn’t considered it before because the caching is done to speed up build pipeline performance: “ While rootless containers prevent the attacker from escalating to host root, the page cache is still shared across the host. Containers that re-use the same base image layers share the same cached pages for those layers — if a malicious CI job corrupts a binary in the page cache, other containers launched from that same image could end up executing the poisoned version.”
- dwattttt 5mo agoI'm no expert, but the kernel is shared between all containers and the host. I don't believe the kernel maintains separate page caches for each container; a malicious CI job could corrupt a binary from any container, or the host.
- firesteelrain 5mo agoOnly if there is a shared inode between host and container.
- duped 5mo agoWhich is almost guaranteed if you're launching multiple containers with the same base image or shared layers.
- amluto 5mo agoSigh. 1. I would hope the default seccomp policy blocks AF_ALG in these containers. I bet it doesn’t. Oh well. 2. The write-to-RO-page-cache primitive STILL WORKED! It’s just that the particular exploit used had no meaningful effect in the already-root-in-a-container context. If you think you are safe, you’re probably wrong. All you need to make a new exploit is an fd representing something that you aren’t supposed to be able to write. This likely includes CoW things where you are supposed to be able to write after CoW but you aren’t supposed to be able to write to the source. So: - Are you using these containers with a common image or even a common layer in an image to isolate dangerous workloads from each other. Oops, they can modify the image layers and corrupt each other. There goes any sort of cross-tenant isolation. - What if you get an fd backed by the zero page and write to it? This can’t result in anything that the administrator would approve of. - What if you ro-bind-mount something in? It’s not ro any more.
- hlieberman 5mo agoIn fact, the authors specifically say on the very first line of their website that the copy/fail primitive can be used as a container escape. The entire premise of this article is flawed and irresponsible.
- deleted 5mo ago[deleted]
- eqvinox 5mo agoAIUI they haven't shown a container escape and are just claiming it so far. Or did I miss something?
- foreman_ 5mo ago[flagged]
- averi 5mo ago[flagged]
- deleted 5mo ago[deleted]
- M_bara 5mo ago> (like reading env vars and sending them to an external server) it'd not be able to send credentials or fetch a malware remotely at all due to the DNS queries being intercepted by eBPF and being sent to a CoreDNS proxy. Wouldn’t the exploit then just use ip addresses directly?
- averi 5mo agoYou can work with the idea of a DNS whitelist, as in you pass a list of allowed DNS entries via your .gitlab-ci.yml (or separate config) resolution happens and those entries (IPs) are stored in a list, any other IP not present in that list gets denied by eBPF (which can easily be used to rewrite the source and destination of a packet before the packet actually reaches the NIC for dispatch)
- Titan2189 5mo ago> [...] that root was just my unprivileged podman user on the host Couldn't you then simply re-run the exploit again as unprivileged podman user and gain root on the host?
- tuananh 5mo agodid anyone try it? it suppose to work right?
- kelnos 5mo agoNo, because you're still in the container, and there's no route to the host's root from there. If you can orchestrate a container escape from the container's "root", then you're on to something.
- wang_li 5mo agoThis pollutes the page cache, which affects the entire host. Getting "root" in a rootless container may mean nothing. But if it attacked the ls, ps, cat, grep, etc. commands and any process outside the container invokes that command it runs the payload of the attacker. What if the payload of the attack is just the same attack to escalate to root? So now you have escaped the container and gained root.
- averi 5mo ago[flagged]
- ezequiel-garzon 5mo agoPlease reply instead of (or in addition to) tagging the user you're replying to.
- pjmlp 5mo agoTagging isn't a feature in HN.
- ramon156 5mo agoThanks for the bikeshedding, they meant mentioning.
- pjmlp 5mo agoIt is also not supported, beyond people by sheer luck see their nick.
- anygivnthursday 5mo agoOr running their Claw scraping HN comments periodically for their mentions.
- zenoprax 5mo agoIf I see my points shoot up a bit I check my comment history to see what caused it.
- hlieberman 5mo agoThat's true... for the exploit demo that they released. The primitive that underlies the exploit, however -- a page cache write -- can easily bypass the container boundary. One only needs to hook an executable which is also present in the host.
- averi 5mo ago[flagged]
- walletdrainer 5mo agoThis feels LLM generated, lots of emdashes and even more text around a completely false premise.
- cpach 5mo agoWhat is the false premise in the article?
- Retr0id 5mo agoThat rootless containers mitigate kernel exploits.
- averi 5mo agoNowhere in the article is mentioned user namespaces completely mitigate the vulnerability, page cache corruption still happens but not being able to obtain root in the target host increases the attack vector to more than just a one liner into having to figure out whether specific shared base image layers are in use and by whom and by what binaries (think of a shared CI platform like the one we run for GNOME).
- Retr0id 5mo agoThe article does not prove that you can't get root on the host via page cache corruption, just that the specific exploit strategy they tried didn't work.
- averi 5mo agoThere's a specific reason why the exploit targets a setuid binary, if you poison it in memory it will be executed with the permissions of the user owning it, in this case root, meaning a setuid(0) + spawning a new shell will effectively give you root access on the host system, this for systems where uid=0 is equivalent inside and outside the container itself. The vulnerability is still there and is deadly serious, with rootless containers the attack vector just increases, the attacker will have to identify other factors (what containers are using a shared base image, what binaries are being called, what binaries should be overridden in memory etc). On top of this there's another thing worth mentioning, it's a common thing in Openshift (for non rootless podman) to allow CAP_SETGID/CAP_SETUID for being able to create a container within a container (this is called the allowPrivilegeEscalation in SCCs), that effectively grants you the ability to become uid=root in the container and in that scenario uid=0 matches the host uid=0. The important difference is that specific instance of the root user doesn't have CAP_SYS_ADMIN (or most of the other privileged kernel capabilities) meaning the actions the user can then perform are very limited.
- netheril96 5mo agoIf the goal is just preventing full root privileges, a CapabilityBoundingSet in a systemd unit will do. However copy fail can be used in many other ways not contained by containers or the above settings. For example it can modify the /etc/ssl/certs to prepare for MitM attacks. If you have multiple containers based on the same image then one compromised CA set affects another.
- est 5mo agoI added these AmbientCapabilities=CAP_NET_BIND_SERVICE CapabilityBoundingSet=CAP_NET_BIND_SERVICE NoNewPrivileges=yes to my .service. Is it good enough?
- cpach 5mo agoGood enough for what? I could be wrong, but I’m not sure those settings are enough to mitigate Copy Fail. If your distro offers a patched kernel, it’s best to upgrade to that one and reboot. You can also disable the vulnerable module (how to do it depends on what distro you’re using). But if you stay on an old unpatched kernel you might be exposed to other vulnerabilites.
- netheril96 5mo agoYou are misinterpreting my goal here. I have patched my kernel against copy fail but I am thinking of ways to harden my setup against future CVEs in the kernel. So the question is, before I learned about copy fail, what could I have done that would have limited the possible damage this vulnerability could do to me? CapabilityBoundingSet is one answer and rootless podman as mentioned in this article is another. They don’t prevent all but at least `su` is useless.
- cpach 5mo agoIf so, I would look into applying a decent seccomp profile. Other hardening solutions could be to run the workloads inside of a VM such as Firecracker, or gVisor. But that might be more work to implement compared to seccomp.
- CalmBirch127 5mo ago[dead]
- HollowRidge427 5mo ago[dead]
- BoldBrook418 5mo ago[dead]
- itvision 5mo agoExploit download/source: https://github.com/theori-io/copy-fail-CVE-2026-31431/blob/main/copy_fail_exp.py https://github.com/theori-io/copy-fail-CVE-2026-31431/blob/m... The dedicated website: https://copy.fail https://copy.fail
- kator 5mo ago[dead]
- codedokode 5mo agoI think it was a bad idea to put cryptographic APIs or VPN in the kernel. If userspace is too slow for this, you should either reduce context switch overhead, or create special kind of processes, which are isolated, but quick to switch into. They are repeating Windows mistakes.
- deleted 5mo ago[deleted]
- ohnei 5mo agoI don't think it was a bad idea, doing any idea requires an investment and a better investment would have been kernel layer, just ask the history of export control law what the US feared breaking more. Having security in userland means attacks in kernel or in userland are worthwhile against it. In the kernel it could have been secured better than OpenSSL was with less resources and could have had keys unavailable from userland. Instead it got basically no uptake as everyone hobbled along on slightly more resources spread even thinner on OpenSSL clones.
- cormorant 5mo agoIt's not faster than userspace, it's much slower normally. On special boards with crypto accelerators it can be faster, and there can be compliance reasons to want it. References: [1] https://www.chronox.de/libkcapi/html/ch01s02.html https://www.chronox.de/libkcapi/html/ch01s02.html [2] https://lwn.net/Articles/410763/ https://lwn.net/Articles/410763/ [3] https://trac.gateworks.com/wiki/linux/encryption#PerformaceComparisons https://trac.gateworks.com/wiki/linux/encryption#PerformaceC...
- cpach 5mo agoWell at least if it’s crufty stuff like AF_ALG that barely no-one is using and is kind of a forgotten place of the kernel. I don’t oppose reasonable crypto in the kernel, like WireGuard.
- cluckindan 5mo ago>barely no-one is using Except, you know, many things
- bawolff 5mo agoIt sounds like they are saying the exploit works but the proof-of-concept doesn't due to superficial reasons(?) That hardly seems like something to brag about.
- raddan 5mo agoIt’s not exactly superficial. It’s defense in depth: make sure that root inside a container is not root outside a container. There is also some good discussion about how the elevated user has access to page caches which can be dangerous when containers share pages (which is common). An attack “not working” for some seemingly trivial structural reason is a common trait of defense in depth. We would all love it if attacks like this were impossible, but absent some evidence of impossibility, why not hedge a little?
- angry_octet 5mo agoThey seem to be in a weird state of denial? Why don't they make it clear that it's just this POC that is blocked? It's like they don't understand.
- bawolff 5mo ago> make sure that root inside a container is not root outside a container. And its a great idea in general, it just doesn't stop this exploit. The proof of concept becomes root as a quick way to prove it has control of your computer. The system in the article isnt blocking the exploit its just blocking the mechanism to prove it worked. It still worked, just the test to verify is now giving a false negative. Good defense in depth disables neccesary steps that by themselves arent sufficient but are a neccesary condition. In the context of this exploit (but not in general) this mitigation is more like renaming the su command to mysu and hoping nobody notices.
- grimblee 5mo agoIf I understand correctly, rootfull podman with --userns=auto would also prevent the privilege escalation ?
- cpach 5mo agoHow?
- grimblee 5mo ago--userns=auto asign a different namespace for each container, so if you escape it you get a random uid far far away from root it also protects other containers from the compromise since they each have their own namespace and uid/gid range, the drawback though is that you can't mount shared volume unless you use a pod, since you would see files from outside your uid/gid range as owned by nobody and inaccessible.
- cpach 5mo agoThat might make Copy Fail harder to exploit, but I still wouldn’t bet money on CF being impossible to use in that scenario.
- grimblee 5mo agoSince in --userns=auto, root inside the container gets assigned to the first uid of the uid range assigned by podman, copyfail would succeed but you'd get uid 647831 and be able to do nothing with it
- angry_octet 5mo agoNo it wouldn't. The exploit is not impacted by namespaces.
- HollowRidge427 5mo ago[dead]
- QuietLedge375 5mo ago[dead]