6 ms·
This exploit is interesting, but if you are doing container security correctly it’s actually not a big deal. In particular if you are setting per-container user
by hacknat 9y ago
This exploit is interesting, but if you are doing container security correctly it’s actually not a big deal. In particular if you are setting per-container usernamespaces, like you ought to be, then this exploit doesn’t do anything. In fact you can actively give a usernamespaced container any CAPs you want, because they are isolated to that container’s uid:gid offset.
Obviously, giving containers unecessary CAP privileges in unwise, but if you are practicing sound security best practices then there would be multiple layers of defense between you and this CVE. I think a strong AppArmor profile and SecComp profile would also make this CVE moot.
Edit:
Also, this exploit relies on you being able to fork up to a certain pid value. You can and should take advantage of Linux’s per cgroup ulimit functionality. No container needs more than 255 threads (even if they do you can make special exceptions for such applications).
Edit2:
Additionally this CVE relies on the getuid syscall being available, there is no reason to give a container this syscall, you should block it, ala this guide:
https://rhelblog.redhat.com/2016/10/17/secure-your-containers-with-this-one-weird-trick/ https://rhelblog.redhat.com/2016/10/17/secure-your-container...
I have to say I’m more than a little dissapointed in Twistlock for not pointing out what countermeasures you can employ against this and other CVEs.
- oblio 9y agoAs a somewhat of a container noob, could you expand on "per-container usernamespaces"?
- baq 9y agoFollow-up question: And why docker doesn't do that by default?
- hacknat 9y agoBecause the Docker project doesn’t make money off of security. It is actually quite infuriating, because they have become the de facto container image standard. Most of their security has actually come from Twistlock (I am not a Twistlock employee, FYI). My recommendation to most Admins or Devs that are serious about container security is to let your developers use docker, but run your images with CRI-O on your servers: http://cri-o.io/ http://cri-o.io/
- nvarsj 9y agoCRI-O is bleeding edge. I'm not sure it's ready for production usage. But it looks very promising. The sooner we can all dump docker in kubernetes the better.
- cpuguy83 9y agoThere are trade-offs to using userns and many ppl don't like the current set of trade-offs. In addition changing a default like this is a breaking change. Admins can enable userns by default in a daemon, but making it a hard-coded default is much more difficult. It's not just a matter of enabling user ns. There is no support at the vfs layer for uid/gid mapping. This means in order to use it, images must be chowned with the remapped ID's. Per-container mappings are not supported for this reason (it would require copying and chowning the entire image for each container mapping). Do you care to qualify your statement about CRI-O?
- ecnahc515 9y agoI recall seeing some patches submitted to make it possible to pass an uid/gid offset to the mount syscall at one point when people were implementing usernamespaces for container runtimes like docker. So is this fixable without having to make every file system implement this feature, or is there something else holding back better support for doing uid shifting for use with user namespaces?
- cpuguy83 9y agoThat has not been accepted into the kernel. It's called "shiftfs", which basically let's you perform the uid/gid shift on mount.
- eikenberry 9y agoViewing docker containers as anything more than a bundling and deployment system is a mistake. While they might help with security they will never be completely secure and you should architect your deployments with that in mind. Unless you are a giant enterprise shop with the resources to staff a decent sized K8s team, you should use the hosted solutions.
- andbberger 9y agoMaybe because it breaks things? I just enabled user namespaces after reading this post. Broke Jenkins and there doesn't appear to be an easy solution. I mount the docker socket in the Jenkins container, which is not an option with user namespaces as the user Jenkins now runs as does not have permission to access the socket. It seems to be possible to provide this user access to the socket through a socket proxy, but since all containers use the same user this seems to defeat the purpose of using namespaces in the first place. Cherry on top: although `docker run` supports running containers with custom userns settings, docker swarm, which I use to run Jenkins, does not. So as far as I can tell my only options are: 1. Go back to not using user namespaces 2. Make the docker daemon on the host available over HTTP, which is really something I was trying to avoid... Anyone have a more elegant solution?
- zenlikethat 9y agoMm, if you're bind mounting in the Docker socket, enabling user namespaces won't help much. You just have to deal with the fact that you have a privileged container (Docker API access == root, at least unless you're using authz). It'd be nice if we could see more RBAC around Docker API so you could do things like "grant only permission to run this one container".
- andbberger 9y agoTotally. But the vast majority of containers I use do not get a bind mount to the Docker socket... for which user namespaces would be a very nice feature.
- zenlikethat 9y agoYeah, definitely turn it on where possible, just important to realize that it's not a panacea (some people really hyped it up to be before the feature was released and criticized Docker for not having it at all). As always, gotta try and find the right sweet spot between convenience and attack surface.
- nassyweazy 9y ago
- fpoling 9y agoI tried it a couple of months ago. It immediately broke build of one of the images. It was a known bug. So I guess I just wait one more year to try. In the mean time, I make sure that all my containers runs as non-root with max security restrictions. The exception so far was sshd from OpenSSH and mostly due to incorrect porting from OpenBSD in portable ssh.
- LaGrange 9y agoAs far as I remember things, because it breaks overlay filesystems, which are a major space saver in Docker world. Something might have changed, but last time I checked, you couldn't "offset" uids/gids on a filesystem overlay, so every layer of the container would have to be copied and chowned (slowly). This would obviously only work for minimal containers (i.e. ones that don't contain a distribution), but software has to be pretty much built for such a case (e.g. statically linked, no dependencies on common tooling — popular with Go, but your Python application won't work edit: unless you copy all the layers, that is). You can read the docs here: https://docs.docker.com/engine/security/userns-remap/#prerequisites https://docs.docker.com/engine/security/userns-remap/#prereq..., and note that it stores image/container layers in subdirectories under /var/lib/docker. Tl;dr: user namespaces are inherently incompatible with many of the usability features Docker brings over other solutions, while they're not particularly useful for many popular use cases (no shared hosting, minor differences in consequence between escalating to the root of the container and its host - though that's an assumption frequently wrongly made).
- zenlikethat 9y agoAlso, people hold their bind mounts to the host near and dear, and user namespaces would break all kinds of things people expect to "just work" with bind mounts. Having user namespaces on by default would break tons of existing scripts, Compose/Kube files, etc. that do things like mount /var/lib/mysql into the container for persistence.
- hacknat 9y agoSure. User namespace-ing is a feature of container security that allows you to grant a process root access to a filesystem that itself is not root. To the running process it appears that it is or can run as root, but on the host it actually isn’t root, but some uid:gid offset. Here’s an article explaining more: http://man7.org/linux/man-pages/man7/user_namespaces.7.html http://man7.org/linux/man-pages/man7/user_namespaces.7.html The gist is that a container is further sandboxed by the kernel that is agnostic of the higher level security precautions. It’s not perfect by itself, but used in conjunction with other features like AppArmor or SELinux and SecComp it can make a container virtually sandboxed.
- borplk 9y ago> but if you are doing container security correctly ... The container that wasn't! (I get the gist of it, just tongue in cheek)
- jwilk 9y ago> Additionally this CVE relies on the getuid syscall being available, there is no reason to give a container this syscall, you should block it, Huh? Lots of legitimate things will break without working getuid(). > you should block it, ala this guide getuid() doesn't require any capabilities, so it can't be blocked by taking them away.
- hacknat 9y agoOops good call. I meant setuid, but either way I was wrong.
- eikenberry 9y ago> but if you are doing container security correctly Doing it correctly should be the case using the default settings. Defaulting to an insecure setup is a bug.
- dvdhnt 9y agoHmm. Perhaps this is a difference between dev and ops, but almost every tool we use comes out of the box with settings unfit for production. Instead, they're tuned for development, and in some cases, deploying to a staging environment. At least, this has been the case in my experience.
- sverhagen 9y agoAh, dev... ops... How about DevOps? As a (originally) dev I bring my app to production. How do I stand a fighting chance to reconfigure the defaults in the way you suggest, without suddenly gaining a whole new set of skills? Good defaults would be helpful, even if they're very conservative. I can break things open, but at least then I know what to read up on.
- kemitche 9y ago100% agree here. Ship with secure settings by default and have simple "developer guides" that show what to crack open for easier use in non-production environments.
- bacongobbler 9y agoI would argue the opposite. As a developer I want to have tools that make my life easier to - you guessed it - develop. Enabling unnecessary secure defaults that either hinder or don't apply to my use case is silly. There's a reason most users choose Ubuntu over OpenBSD as their workstation. I would put good money on the reason is because it's "secure enough" without getting too heavy handed on production use cases. However, I do agree that there has to be a balance. Most tooling I write tends to lean more towards the "good user experience" side first, and then document the production use case. Either that or release two separate (but similar) products; one for developers, one for operations teams. Docker's doing that with the Community Edition/Enterprise Edition, but I still think the Community Edition is far too heavy-handed when it comes to things like pulling images from "insecure" registries.
- bmitch3020 9y ago> In particular if you are setting per-container usernamespaces, like you ought to be, then this exploit doesn’t do anything. User namespacing in docker is enabled at the daemon level, not per container, so all containers share the same offset. This would ensure that a root user in the container would escape to a different uid on the host, but doesn't prevent someone from moving sideways through the containers on the same host. Note that enabling this will break the developer workflow of mounting files from the host into the container. I believe files will show up with the wrong ownership inside the container.
- hacknat 9y agoYou don’t need to use containerd. Other runtimes make it possible (per container offsets have been possible in runc for over year).
- wahern 9y ago1) User namespaces don't magically protect you from a vulnerability that allows writing to kernel memory. Neither would AppArmor. seccomp could theoretically, but waitid is a pretty fundamental Unix syscall and blocking it would break a lot of basic software. The author devised a particular exploit, but his example was hardly the only way to leverage the vulnerability. Being able to write to kernel memory is about as huge a vulnerability as you can get. Just because you can't think of a way to leverage a vulnerability doesn't mean an attacker can't; your failure of imagination is not evidence that it cannot be done. 2) Plenty of containers need more than 255 threads. Like, pretty much any Java server. In any event, this particular exploit doesn't necessarily require hundreds or thousands of simultaneous processes. 3) Blocking getuid is even worse than blocking waitid. Block getuid and you'll break glibc and god knows what. In any event, it would be futile as the real and effective UIDs are passed to the process through the auxiliary process vector when the kernel executes the process. 4) You're missing the forest for the trees. The real moral of the story is this: "In 2017 alone, 434 linux kernel exploits where found". Unless you're prepared to pour over every published exploit, 24/7, meticulously devise countermeasures, and be prepared to run effectively crippled software, you really shouldn't be relying on containers to isolate high-value assets. I wouldn't rely on VMs, either, as the driver infrastructure of hypervisors has also proven fertile ground for exploits.
- hacknat 9y agoIt’s not a full kernel memory CVE, you have +-255 bytes access to kernel memor from the cred pointer. I have no idea if that extends to userns or not. Also I think your confusing Java threads for system threads they are not the same. I think your being overly alarmist. You have to trust someone else’s code at some point, otherwise you’ll be paralyzed by non-productivity.
- geofft 9y ago> It’s not a full kernel memory CVE, you have +-255 bytes access to kernel memor from the cred pointer. I have no idea if that extends to userns or not. As I understand it, a kuid_t is the UID in the root namespace, so setting your cred->uid to 0 gets you considered as equivalent to root in the container host. Also, don't think that limited exposure to kernel memory saves you - take a look at the sudo "vudo" exploit from 2001, in which a single byte that was erroneously overwritten with 0, and then put back, turned out to be exploitable. http://phrack.org/issues/57/8.html http://phrack.org/issues/57/8.html (And in general, don't confuse the lack of public existence of an exploit with a proof that a thing isn't exploitable in a certain way.) > Also I think your confusing Java threads for system threads they are not the same. Current versions of the HotSpot JVM (where by "current" I mean "since about 1.1") create one OS thread per Java thread: http://openjdk.java.net/groups/hotspot/docs/RuntimeOverview.html http://openjdk.java.net/groups/hotspot/docs/RuntimeOverview.... "The basic threading model in Hotspot is a 1:1 mapping between Java threads (an instance of java.lang.Thread) and native operating system threads. The native thread is created when the Java thread is started, and is reclaimed once it terminates." Plus there are some other OS threads for the runtime itself. > I think your being overly alarmist. You have to trust someone else’s code at some point, otherwise you’ll be paralyzed by non-productivity. Sure, but you can choose which code to trust, and how to structure your systems to take advantage of the code you trust and not the code you don't. Putting mutually-distrusted things on physically separate Linux machines on the same network is a pretty good architecture: I trust that the Linux kernel is relatively low on CVEs that let TCP packets from a remote machine overwrite kernel memory.
- geofft 9y ago> Edit2: Additionally this CVE relies on the getuid syscall being available This exploit relies on it. The vulnerability does not. The exploit happens to use getuid() along the way to using heap spraying, but the writeup is pretty clear that neither getuid() nor heap spraying is required.
- hacknat 9y agoYeah and I’m wrong about that part anyways. You can’t cap out or block getuid without breaking glibc. I meant setuid, but that call isn’t used in this exploit. I got confused.
- quotemstr 9y ago> No container needs more than 255 threads > Additionally this CVE relies on the getuid syscall being available, there is no reason to give a container this syscall, The problem with MAC schemes is that, in practice, they lead to security people imposing random and arbitrary restrictions on general APIs in the name of the least privilege. In doing so, they break the orthogonality of general-purpose platform concepts and break the reductive mental model necessary to get anything done. It's a misunderstanding of what least privilege actually means. Security is better achieved by creating clear, principled security domains and boundaries, then controlling access to these domains in a general and transparent way. Saying "you, unix process, you can call system call X, but not system call Y, because in my opinion, Y is risky", when neither X nor Y breaks through a security domain, is bad practice. So is arbitrarily capping the number of threads in a container.