4 ms·
How Cloudflare responded to the “Copy Fail” Linux vulnerability
- deleted 5mo ago[deleted]
- skinfaxi 5mo agoWould love to learn more about their internal behavioural detection program. > One of the first things our security team did was confirm that our existing endpoint detection would catch this exploit. Our servers run behavioral detection that continuously monitors process execution patterns. It doesn't rely on knowing about specific vulnerabilities; it watches for anomalous behavior across the fleet.
- CGamesPlay 5mo agoWould certainly be interesting to learn more about. A simple check: allowlist of known "processes that run as root". Any new process shows up, something happened.
- jeffbee 5mo agoBased on what? Proc title?
- CGamesPlay 5mo agoProc title is very easily forged (without root even). Obviously a real privileged process could modify the kernel and do whatever it wants, but if I were trying to detect this I would start with /proc/$id/exe.
- jeffbee 5mo agoMaybe, but there's a prctl to change that reference which a root process can use.
- Retr0id 5mo ago/proc/pid/exe is also easily forged, without root. For example you can do LD_PRELOAD=evil.so /bin/foo on any dynamic executable, or spawn /bin/foo unmodified and inject code via ptrace or /proc/pid/mem. I have a fileless, execless copyfail exploit that works by injecting shellcode directly into systemd's pid 1. (I should probably publish it at some point...)
- jeffbee 5mo agoYeah the whole system is based on the ability of one task to apparently become another task, that's how Unix works. So the indicators in /proc are just that: indicative at best. There's no reason the task should even be assumed to be executing code in a file. A process can map code into anonymous memory and continue executing there without even branching. Again this is considered a feature of the system rather than a flaw.
- parliament32 5mo agoIt's curious they're just "monitoring" rather than preventing. In a serious environment you'd run IPE with dm-verity/fs-verity to ensure binaries are whitelisted and integrity-checked at every execution.
- staticassertion 5mo agolol no one does that (edit: or, rather, that is extremely uncommon, even in "serious" environments, for a ton of reasons).
- parliament32 5mo agoLook at the FedRAMP requirements around integrity protection, then look at how massive the list of complaint products is. I promise, pretty much everyone in regulated environments is. It's so prevelant Azure is even pushing a turnkey solution for k8s https://learn.microsoft.com/en-us/azure/aks/use-azure-linux-os-guard https://learn.microsoft.com/en-us/azure/aks/use-azure-linux-...
- jeffbee 5mo agoIf you have much experience with fedramp, and it sounds like you do, perhaps you might agree that it is a huge list of things that superficially indicate doing something, without actually doing anything. As the documentation for IPE freely admits, it has no protective benefits because it is unaware of anonymous executable regions.
- parliament32 5mo agoIt sure has limitations, but "no protective benefits" is pretty wrong. In a real world example, if your containerized application has an RCE, you're preventing the attacker from executing binaries they tampered with or down/up-loaded. Combined with minimal distroless containers, it's a very effective attack surface reduction strategy, and works much better than the legacy scan-occasionally integrity-checking methods (rkhunter et al).
- dboreham 5mo agoThey might just compute a hash over the binary, or the code space in memory.
- mobeigi 5mo agoI'd very much like to learn more about this too, deserves its own blog post.
- staticassertion 5mo agoSyscalls and kernel module loading can both be logged, I assume that's sufficient here.
- skinfaxi 5mo agoYes but I am interested in hearing about cloudflare's implementation, how they scale it to their whole fleet, and what kinds of heuristics they are using to classifying behavior as anomalous.
- john_strinlai 5mo agothis is a techincal dive into how cloudflare responded, not a confirmation that they responded for whatever reason, unknown to me, hn automatically strips "how" from the start of titles. i cant remember ever seeing a title where this was an improvement.
- trollbridge 5mo agoStarting a title with “How” is standard clickbait.
- Goronmon 5mo agoIf we are taking that attitude why not go all the way? Titles are standard clickbait.
- miki123211 5mo agoWith LLMs, you could actually do anti-clickbait titles. Extract the article text with something like r.jina.ai, and ask an LLM to generate a ~80-character summary that explains the main point of the article for people too busy to read it. I do think this would genuinely be useful.
- john_strinlai 5mo agoback in my day, people just used the thing that rattles around inside their skull for such tasks
- senko 5mo agoTo do that, you need to read the article first, which is the point of click-bait titles. The point of the defense is to avoid exposing your neurons to that stuff.
- john_strinlai 5mo ago
- dboreham 5mo agoThe "Hunting for Exploitation" section is unclear to me: "The exploit leaves a distinctive trace in kernel logs when it runs." Hmm. Wouldn't a system with a compromised kernel also log exactly what the attacker wanted logged?
- QuantumNoodle 5mo agoAlso 48 hours prior the disclosure is a very narrow window? I wonder if their logs don't go back further or if there was another reason to look back only two days.
- cube00 5mo agoI guess the hope is the kernel has been able to successfully transmit that log message to the immutable central logging infra before it gets compromised. Although given the tendency for end point logging agents to run on buffers to reduce their network chattiness I do wonder if a fast acting exploit could dump that buffer before it manages to be transmitted. I don't think any of the agents are complex enough to immediately transmit permission elevation log messages over the regular background noise.
- rithdmc 5mo agoThe attack itself creates the logs, which - reading between the lines - are shipped to a central log server. A compromised server might not send any new indicators to the logs, but existing logs moved off device would still be available. I'd like to know what those distinctive traces are, which is also missing :(
- PunchyHamster 5mo agoYour exploit would have to get root and kill/exploit the logging daemon near instantly, else the log will already be sent to remote before you can change it locally
- srcreigh 5mo agoIt’s fascinating that already had a system which could identify the exploit at runtime. How can I learn more about that?
- sammy2255 5mo agoAny Cloudflare employees reading this, your network map has a few PoPs missing from it https://www.cloudflare.com/network/ https://www.cloudflare.com/network/ notably, Perth (PER) Australia. Hobart (HBA) Australia. Wellington (WLG), New Zealand. Christchurch (CHC), New Zealand. Nausori (SUV), Fiji.
- cube00 5mo ago> At the time of the "Copy Fail" disclosure, the majority of our infrastructure was running the 6.12 LTS version That could be as low as 50.1%, I wish they'd provide an actual percentage.
- jmclnx 5mo ago> Linux kernel build based on the community's Long-Term Support (LTS) CopyFail only highlights why Companies want LTS. If there was a supported kernel built prior to 2017, most large companies would still be on that version, avoiding this issue all-together. The corporate mindset is usually "never upgrade unless there is new hardware needed or critical software failure". All CopyFail did was reinforce that mindset. I wonder if CopyFail will cause enterprises put pressure on the Linux Foundation to maintain a "ultra LTS" were it is supported for 20 years ?
- PunchyHamster 5mo ago> CopyFail only highlights why Companies want LTS. If there was a supported kernel built prior to 2017, most large companies would still be on that version, avoiding this issue all-together. Sadly not really how it works for say Red Hat. They routinely backport features while keeping whatever "stable" number on kernel. We even had displeasure of them backporting a bug... same bug to 2 different RHEL versions
- tempest_ 5mo agoThe longer you wait the more painful the switch will eventually be.
- em-bee 5mo agofor the kernel? hardly. only if the kernel breaks userspace. which it shouldn't.
- PunchyHamster 5mo agofor us it was * Get list of modules from Puppet's facts, confirm module isn't used anywhere (it wasn't) * `install algif_aead /bin/false` in /etc/modprobe.d/disable-algif.conf * Run a check using exploit code to check it is no longer working I imagine CF runs more stuff that could use it I guess but apparently it's not often used API
- mkj 5mo agoIf they're already running a custom Linux kernel build, why did they have AF_ALG enabled? Seems the perfect situation to limit features to only those actually being used.
- computerfriend 5mo agoIn the article they explain that some of their services use it.
- mixdup 5mo agoAnd also as part of this, they have learned the lesson parent comment is trying to make: they called out that they are going to review their deployments and make sure there's no unused modules being deployed
- electra2012 5mo ago> Despite our practice of deploying Linux patch updates every two weeks, we remained vulnerable because a month-old mainline fix had yet to be backported to our primary kernel line. Hopefully a wake-up call to those who believe older distro LTS kernels are getting all the security fixes Canonical and Redhat would want you to believe.
- cluckindan 5mo agoHas anyone figured out whether this CVE was intentional?
- tptacek 5mo agoThis is an interesting post from Cloudflare, as usual, but it's not clear to me why they would have been vulnerable to CopyFail. Did I miss the point in this blog where that's addressed? What triggered the threat hunting and mitigation exploit? At what points in their architecture were they reliant on Linux user-based access control?
- robotbikes 5mo agoI would assume it was about protecting their servers from internal sources escalating privileges vs. them providing publicly accessible Linux shells.
- tptacek 5mo agoI mean, that's a real project, but Linux LPEs kind of grow on trees, so you can't literally rely on threat intelligence for this problem; presumably you handle it by drastically scoping down and surveilling what people do on prod hosts.
- aduwah 5mo agoThe whole IT industry is reliant on Linux user-based access controls, it is not a Cloudflare thing. Also leaving a massive gap like this behind would be a mistake on multiple levels. For example, it might get combined with another exploit that can achieve unprivileged access to some piece of metal, or you can have a disgruntled employee without admin access escalating their permissions on a box they aren't supposed to see all the secrets.
- TacticalCoder 5mo ago> For example, it might get combined with another exploit that can achieve unprivileged access ... Yeah. TFA mentions datacenters in 330 cities. That's a lot of Linux boxen. And many of those have, by definition, ports opened to the big bad Internet. These Linux servers are running services. They answer to ping, for a start. I even heard some are running DNS servers. Remote local exploits are a thing. What does CloudFlare prefer: that when the next remote local exploit surface all their fleet is one copy.fail away from privilege escalation to root or that they get the time (seen that they obviously have quite advanced detection measures in place) to detect the intruder before it gains root everywhere? It's Linux. It's datacenters in 330 cities. Linux powers the world and that's how things works. I, for one, I'm glad to own CloudFlare stocks since right after the 2022 crash and, for two, I'm happy they don't let their huge fleet of Linux servers with a non-patched exploit.