8 ms·
This assumes that Docker containers are being used like VMs. They're not designed to allow running isolated arbitrary code in some sort of multi-tenancy setup,
by notjack 11y ago
This assumes that Docker containers are being used like VMs. They're not designed to allow running isolated arbitrary code in some sort of multi-tenancy setup, they're designed to isolate dependencies and configuration between your own deployed services. This is a "security vulnerability" in the same way that putting up a white picket fence around a jail is - a gross mis-application of a tool built for a completely different purpose.
- insaneirish 11y ago> They're not designed to allow running isolated arbitrary code in some sort of multi-tenancy setup You mean Linux containers are not designed for this, nor are they designed to be secure. What a sad design failure. Some [1] container technology was developed with security as a first principle. [1]: http://us-east.manta.joyent.com/jmc/public/opensolaris/ARChive/PSARC/2002/174/zones-design.spec.opensolaris.pdf http://us-east.manta.joyent.com/jmc/public/opensolaris/ARChi...
- notjack 11y agoI should clarify I mean Docker containers specifically, good point.
- geofft 11y agoAnd you can run Docker containers on SmartOS, using "LX-branded zones", which do Linux syscall emulation: https://www.joyent.com/blog/triton-docker-and-the-best-of-all-worlds https://www.joyent.com/blog/triton-docker-and-the-best-of-al... I'd be curious if the Joyent folks believe Triton to have the security properties that everyone seems to want out of Docker. It doesn't involve the standard Docker daemon, so it might.
- lloydde 11y ago@geofft that is a great link, our CTO Bryan Cantrill continued to beat the drum at the Docker SF meetup in May 6, 2015. Video and transcript at https://www.joyent.com/developers/videos/bryan-cantrill-virtualizing-the-docker-remote-api-with-triton https://www.joyent.com/developers/videos/bryan-cantrill-virt... . My favorite line is, "And, so in terms of the security of containers, from our perspective they have to be secure because if you're able to break out of a container into the global zone of SmartOS, you could delete the company." Yesterday at http://containersummit.io/ http://containersummit.io/, Bryan's talk "Going Container Native" (video coming soon) was in the same vein, SmartOS (illumos) directly benefits from its Solaris lineage -- Solaris Zones were engineered for security. We, Joyent, absolutely do believe Triton has the security properties that everyone seems to want out of Docker, because Joyent has been running OS containers in multi-tenant production since ~2006. Open source SmartDataCenter [1], rebranded Triton, is the exact same code that is run in Joyent's Public Cloud. This truth in code, product and service contributed to me leaving an OpenStack company to join Joyent two years ago. 1. https://github.com/joyent/sdc https://github.com/joyent/sdc
- lloydde 11y agoBryan's talk "Going Container Native" from earlier this week is now online: http://containersummit.io/events/sf-2015/videos/going-container-native http://containersummit.io/events/sf-2015/videos/going-contai...
- JoshTriplett 11y agoAs far as I can tell, Solaris Zones are no more secure than Linux containers; both use kernel-level isolation mechanisms, while sharing the same kernel between zones/containers. And thus they both have the same security property: secure as long as you can't successfully exploit a syscall, but if you can, you get kernel-level access.
- _delirium 11y agoThe security property of "secure unless there's an exploit" applies to basically everything (maybe excluding formally verified software), including Xen and KVM, so I don't see how that's a specific knock against Zones. Yes, if there's an exploit that lets you do something you shouldn't be able to do, you can break out of the container/VM. But, like Xen and KVM, Zones are so far fairly successfully used as an isolation mechanism for the public cloud, for some years now. And all three make the claim that you ought to be able to rely on their isolation mechanisms, and any hole in them will be considered a serious bug. Whereas the way Docker is using Linux containers doesn't seem secure enough to make that claim or be used for that purpose, at least not yet.
- skarap 11y ago> Whereas the way Docker is using Linux containers doesn't seem secure enough... Which is totally expected, because docker tries to be useful to the largest user-base possible. It would be quite harder to use if didn't support directory mounting (via -v). And it would be a total nightmare for almost every user if you had to specify a list of allowed syscalls for every container. This reminds me the situation with SELinux a lot. It has improved a lot but I still see "disable SELinux" almost in every tutorial I read on CentOS, Fedora or RHEL.
- pjmlp 11y ago> This reminds me the situation with SELinux a lot. It has improved a lot but I still see "disable SELinux" almost in every tutorial I read on CentOS, Fedora or RHEL. Because security is hard. Just look at the Mac OS X users running as root and disabling Gatekeeper. Or the developers that stay away from the sandbox model. One of the nice things of mobile OSes is that there isn't a way around the container model. Although the history with permissions kind of messes it.
- kentonv 11y agoLinux containers are certainly intended to be secure. The problem is that the Linux kernel API is huge and privilege escalation bugs are found all the time. In order to make a container secure, you need to disable 95% of this API, e.g. using seccomp, not mounting /proc or /sys, drastically limiting /dev, etc. Sandstorm.io hasn't seen a single working breakout since we started keeping track over a year ago, despite numerous Linux kernel exploits going by in that time, because all the bugs have been in features that we turned off. https://docs.sandstorm.io/en/latest/developing/security-practices/#server-sandboxing https://docs.sandstorm.io/en/latest/developing/security-prac... That said, there is a cost in compatibility. Most apps can be made to run just fine in this constrained environment, but it does sometimes require tweaks. Docker is more interested in compatibility than in security, so naturally they don't do this kind of attack surface reduction by default. (You can configure it manually, but realistically if it isn't mandatory then few people will bother.)
- benmmurphy 11y agoIt is also important to note that Microsoft is not offering shared kernel on Azure using their new container technology because they believe it is too insecure. Microsoft might be a special case because their kernel has an especially large surface area. But I think the current cloud providers that are providing shared kernels are walking a very fine line when it comes to security. Seccomp obviously mitigates this problem and probably puts you in a better position than using a hypervisor if you lock things down very aggressively.
- acconsta 11y agoAre linux containers really not securable? That's how Google runs all their stuff in production. http://research.google.com/pubs/pub43438.html http://research.google.com/pubs/pub43438.html
- geofft 11y agoGoogle generally doesn't have to worry about mutually-untrusted containers. The meaning of "Linux containers are not secure" is that untrusted code should not be given root privileges within a container. Google is generally not doing that. They have e.g. trusted Gmail code running on the same machine as trusted YouTube code, handling untrusted emails and untrusted videos. But the Gmail team is not worried about the YouTube team hacking them, or vice versa. The security mechanisms just need to keep honest people honest. And when they do have untrusted, third-party code to run, for Google Cloud Platform, they use VMs or actual sandboxes: see section 6.1 of the paper you linked.
- acconsta 11y agoNo, but the Gmail team might be worried about some Russian guy finding a buffer overflow in YouTube's application. Then what? They can escalate privileges and read your email?
- KirinDave 11y agoGoogle's deployment of Docker is less susceptible to this. They do additional partitioning of applications.
- acconsta 11y agoYeah... so it seems like Linux containers can be sufficiently hardened?
- the_mitsuhiko 11y agoNot really. You physically seperate gmail and youtube.
- stephenr 11y agoIsn't this specifically a Docker failing? My understanding (from reading rather than use) is that LXC (the original Linux container project) provides a lot more in the way of security than Docker does?
- pjmlp 11y agoSpecially since HP-UX and Tru64 were doing it already in the late 90's.
- zurn 11y agoListening to what the Docker people write about security, they sure do sound like they are designed to be secure, even though they admit to a couple of potential shortcomings versus VMs. https://blog.docker.com/2013/08/containers-docker-how-secure-are-they/ https://blog.docker.com/2013/08/containers-docker-how-secure... concludes, "Docker containers are, by default, quite secure; especially if you take care of running your processes inside the containers as non-privileged users (i.e. non root)." They also recommend using SELinux with Docker to beef up security (see https://blog.docker.com/2014/07/new-dockercon-video-docker-security-renamed-from-docker-and-selinux/ https://blog.docker.com/2014/07/new-dockercon-video-docker-s...) As long as people are using containers as a security boundary it makes sense to pay attention to things like this one about the control socket.
- kentonv 11y agoThere is a section in that link titled "Specific Attack Surface of the Docker Daemon". Meanwhile they are ignoring the attack surface of the Linux kernel. There is no mention of seccomp in the whole post. People find privilege escalation bugs in the Linux kernel often enough that you can't really claim that running arbitrary native code is "secure" unless you're doing something to mitigate these. Things like: http://www.openwall.com/lists/oss-security/2015/07/22/7 http://www.openwall.com/lists/oss-security/2015/07/22/7 Note that SELinux won't do anything about this kind of bug. You have to use seccomp and similar to disable large swaths of the Linux kernel API, particularly exotic parts that aren't well-tested or reviewed.
- throwaway2048 11y agothere is no mention of seccomp because virtually no software can utilize it.
- lvh 11y agoAbsolutely; as I mention in the article. There's nothing new or exciting here; there's simply a big discrepancy between reality and how a lot of users understand it. This is partially true for the my-dev-user-is-part-of-the-docker-group case, but even more so for the container-with-docker-sock-access case.
- notjack 11y agoFirst let me state that this article is interesting and well written and taught me something new, thank you for that. However my main frustration is I'm not really sure what the article is advocating for here. Instead of not giving access to the docker daemon to containers (which is legitimately needed for complex deployments where one container needs to dynamically start up "sibling" containers, e.g. a CI service), wouldn't it make more sense to talk about not viewing Docker container security the same as VM security in the first place? Sure if you're going to do that anyway it makes sense to disable access to the socket, but then there's a million other things you'll have to do because docker containers are currently not primarily intended as a replacement for the isolation security of VMs. Their security is more like a useful extra layer, rather than a full blown replacement.
- nailer 11y agoRunning docker inside VMs loses all the IO benefits. That's not to say I use docker at all: I don't for this reason. Someday it will be solved and docker will actually be as production ready as its proponents think it is.
- nailer 11y agoSince folks seem to be new to docker and Xen: Hypervisors like Xen have very poor IO performance: one of the major benefits of docket is that it provides containerisation without the IO overhead of a hypervisor. Neither of those have been disputed by anyone, ever.
- nogox 11y agoCheck www.hyper.sh and https://github.com/hyperhq/runv https://github.com/hyperhq/runv They can boot a new VM with Docker images in 200ms, which is very close to LXC, and perfectly isolated by hypervisor. The problem of "Virtual Machine" is not "Virtual"/Virtualization, the problem is the full blown guest OS, aka "Machine".