9 ms·
New Linux vulnerability affecting cgroups: can containers escape?
- frabbit 5y agoImportant note on this: "Fortunately, the default security hardenings in most container environments are enough to prevent container escape. Containers running with AppArmor or SELinux are protected. " So, all that hard work on SELinux continues to pay off.
- concerned_user 5y agoSadly, many answers to questions related to selinux issues, or howto's start with: Disable selinux.
- AaronFriel 5y agoI strongly believe that software that works for users is better than software that doesn't, and it's clear that for most lay folks, SELinux is software that doesn't work. SELinux remains inscrutable and unusuable to the lay person. Microsoft had the same problem with Windows XP and especially after its service pack 2 when the Windows Firewall was introduced, that it was difficult to debug and applications didn't prompt to open ports or have an API to do so. So many a lay person posted on forums "disable firewall". Users don't care why their tools don't work, they don't understand why or how to fix it. Technically complex SELinux audit tutorials are not helpful. There needs to be real, genuine attention to user experience an almost tutorial like CLI command. Something so simple anyone could safely make a program run. Whether that program is safe itself is another question, and users should be told that too.
- frabbit 5y ago> it's clear that for most lay folks, SELinux is software that doesn't work. I have always used selinux enabled systems. For the first few years it was a bit confusing and frustrating at times, but for the last (decade?) I have never had to butt heads with it. The default policies shipped by e.g. Fedora (a userland closest to the development of this work and therefore probably better maintained than some others) work out of the box without hassle. This very article refutes your assertion: here we see SELinux working for ordinary users without any additional fiddling. You, on the other hand, are probably exposed to this privilege escalation.
- AaronFriel 5y agoThat's not my assertion, my assertion is that SELinux doesn't work for a lot of people even if it works for you or I; and that's why you see the advice to disable it in forum posts. To be clear: SELinux is an important mitigation - just like the Windows Firewall - and one should not disable either.
- frabbit 5y agoI disagree. The advice to disable SELinux, like your assertion that it's too complicated for ordinary users, belongs to an older time. It's time to lay that myth to bed. Sure, if you're messing around with k8s and doing fun eBPF stuff you are going to need to be careful. But for just installing an OS, running it to do some web-browsing, gaming, image editing, wordprocessing? I would be highly surprised if the defaults do not work.
- deleted 5y ago[deleted]
- AaronFriel 5y ago> The advice to disable SELinux... time to lay that myth to bed. I think we agree, and Fedora / Red Hat have done great work setting up great defaults. But when a user encounters an issue with SELinux, the lack of feedback mechanisms to help them onto a better path results in them finding that advice.
- MonaroVXR 5y agoFedora literally gives you a notification and you can take action (Me a as novice Linux user)
- AaronFriel 5y agoThat's fantastic for Fedora desktop users. I don't expect you'd know, but is there a way to get the same quality of information via a CLI command?
- pjmlp 5y agoOne of the best things on Android is having SELinux and seccomp enabled.
- chunkyks 5y agoThings may have changed, but the last few times I looked, it was breathtakingly hard to a) identify if /when selinux is what's screwing you, then b) get selinux to stop it. I really wanted an audit mode that could also say "this command will unlock the specific thing I just blocked". That was a few years ago. Since then, I've turned off selinux whenever I'm getting screwed by some opaque process, stuff starts working, and closing it back down while leaving what I need open remains impossible black magic.
- chucky_z 5y agoIs audit2allow the thing you want?
- loudmax 5y agoProbably yes, but audit2allow is very hard to reason about. You can run it and hopefully it will enable you to allow the things you want to allow without also allowing things you didn't want. Red Hat doesn't seem to have any interest in making SELinux more accessible than programming in assembly. The UX for the tooling around SELinux is an absolute dumpster fire.
- chunkyks 5y agoI recall taking a stab at audit2allow a few years ago, and finding that it was incredibly opaque and felt like practising dark arts. At this point, it's probably true that I should get onboard the SELinux train and learn it properly, but it's just... ain't nobody got time for that.
- chucky_z 5y agoI believe this is considered one of the best videos: https://www.youtube.com/watch?v=_WOKRaM-HI4 https://www.youtube.com/watch?v=_WOKRaM-HI4
- mhitza 5y agoOn server environment that command is most of the time not installed by default. Quick! tell me which package I need to install to get audit2allow on a system; without using Google, dnf whatprovides, or repoquery --whatprovides. I'm still baffled why such an essential tool for quickly assessing violations and potential selinux booleans quick fixes is part of a obsfucated package name. I think some setroubleshoot family of tools might be installed by default on some systems, even if most answers will guide people to just use audit2allow.
- 2OEH8eoCRo0 5y agoWhy is the Linux community full of horrible advice?
- not2b 5y agoOften the issue is that what was decent advice five years ago can become horrible advice later, but the formerly decent advice is already moderated to the top on Stack Overflow and Reddit and is the first hit on a Google search, while the approach that people should be using now just isn't found. You'll often find horribly complex multi-step instructions for How to Do X ranked more highly than simple instructions about how to use a new interface to do everything in one step, because there was a window of time when the complex instructions were required, everyone was so grateful for them and upvoted them.
- silisili 5y agoI'd argue the vast majority of Linux desktop users (already a small group) don't use SELinux. So naturally when trying to help someone using something we don't have experience with and don't find necessary, that advice becomes more prevalent.
- yrro 5y agoWhich is an excellent indicator that the following advice is bad.
- kobalsky 5y agoyou are also safe if you are not running (EDIT: inside) the container as root, which is a common security practice for containers nowadays.
- suifbwish 5y agoCouldn’t you prevent against this sort of thing by using disposable VMs to host the containers? Sure it would be an extra layer of resources but it would double the complexity of the attack required to breach the physical node.
- yjftsjthsd-h 5y agoCorrect on both counts; you can, and it hurts performance / resource use. There's also intermediate options like gvisor. In practice, the performance issues mean that most people don't bother.
- leephillips 5y ago“Container” seems to be used throughout to mean “Docker container”. There are other types of containers.
- AaronFriel 5y agoI think container escape is well understood by most to mean (for Linux) cgroups and/or the stack most folks use (containerd, Docker). It's a generic term but useful term, like VM escape, even though there are many kinds of virtual machine managers and hypervisors.
- stormbrew 5y agothe other reply alluded to this, but to make it explicit: nothing about this CVE requires docker and it looks like you should be able to do it with a few syscalls in any process starting with a call to unshare(), unless something else (like selinux) is getting in your way.
- xxpor 5y agoBack in the day, people insisted that containers were not security boundaries and should not be treated as such. They're meant to contain things from going off the rails unintentionally, but an actual threat was another story. However, realistically, given the env that a container gives you, it certainly looks and feels like a security boundary. So are we just going to be stuck in this retroactive security cleanup mode forever? My point is that if it were designed from the ground up with the hard security boundary in mind, would we have ended up with containers in the first place? If not, is there any realistic way to go from where we are to where we should be? The only other design I'm familiar with that sort of comes close are MicroVMs. Those have the downside of actually needing to run a VM though, and most (all?) cloud providers don't allow nested virtualization so you're stuck running on an enormous bare metal box.
- stormbrew 5y ago> My point is that if it were designed from the ground up with the hard security boundary in mind, would we have ended up with containers in the first place? If not, is there any realistic way to go from where we are to where we should be? Yes, because systems that are designed with these kinds of security boundaries in mind already look like containers -- they're a natural match to actual capability-based systems like, for example, plan9's. The problem here stems entirely from trying to keep these globally-overriding capabilities like CAP_SYS_ADMIN and CAP_DAC_OVERRIDE while also allowing users to create their own namespaces. All these CVEs weren't things as long as only root could create new userns', and now that normal users can all these areas where things weren't checked are coming out of the woodwork. But a ground up capability-based system avoids this kind of problem by simply making it impossible to elevate to a privilege level like 'root' on POSIX systems, and so namespacing within those systems is incredibly natural to the point that it didn't really get a name (containers) until one was needed for linux' cognitive dissonance around the idea.
- tptacek 5y agoI'm not sure a year has gone by without a vulnerability that breaks shared-kernel isolation in reasonable configurations. Nobody was going to DAC or MAC out `waitid`, but `waitid` for a time take a kernel address for its siginfo_t parameter.
- jeffbee 5y agoThis style of writing sucks, and the abuse of the meaningless term "container" does nothing to clear it up. To reduce this CVE to one sentence: a process running in the top level control group, which has the ability to create user namespace, can take over the machine, because the kernel fails to check for CAP_SYS_ADMIN. See how easy that was?
- jackpirate 5y agoIsn't the whole purpose of this style of writing to define terms like "top level control group" and "CAP_SYS_ADMIN" for those people who don't already understand what they mean?
- jeffbee 5y agoThe article doesn't do that. It throws around jargon without defining it, or defining it vaguely or inaccurately.
- stormbrew 5y agoYou kind of missed the key to the whole thing here, though, which is that users are able to create userns' now by default. This is really important to understanding this and the last few container escape CVEs. The article doesn't do much better on that front, but it is in there at least.
- deleted 5y ago[deleted]
- encryptluks2 5y agoPodman and other container tools are now using user namespaces by default. I think it is clear there are some extra precautions needed, but ultimately the goal with running rootless containers is to improve security.
- nezirus 5y agoPodman also works fine rootless and with cgroups2, double win.
- MonaroVXR 5y agoDoes it support docker-compose?
- encryptluks2 5y agoYes, there is a podman-compose package as well.
- nezirus 5y agoActually, it supports docker-compose proper (v1 and v2, even if there are some bugs for v2)
- norenh 5y agoWow, all that long article and no mention of the versions affected. Fortunately it is mentioned in the redhat bugzilla for that CVE and there it states that it is fixed in stable kernel v5.16.6. I assume it is also fixed in stable stable kernels released at the same time: 5.15.20, 5.10.97, and 5.4.177 sources: https://bugzilla.redhat.com/show_bug.cgi?id=2051505 https://bugzilla.redhat.com/show_bug.cgi?id=2051505 https://lwn.net/Articles/883949/ https://lwn.net/Articles/883949/
- throwaway984393 5y agoI consider cgroups an administrative layer, not a security layer. They're for keeping apps from accidentally blowing up the host, not to prevent them hacking it. If you want security with containers, use Firecracker.
- fmajid 5y agoThe article is maddening in its generic “update to the latest kernel” advice and not listing the specific kernel versions fixed. They are: - Stable: 5.16.6 - LTS (for Alpine Linux): 5.15.20 Alpine 3.15 main is currently at 5.15.16 and thus vulnerable.