8 ms·
Containers (meaning Docker) happened because CGroups and namespaces were arcane and required lots of specialized knowledge to create what most of us can intuiti
by alphazard 1y ago
Containers (meaning Docker) happened because CGroups and namespaces were arcane and required lots of specialized knowledge to create what most of us can intuitively understand as a "sandbox".
Cgroups and namespaces were added to Linux in an attempt to add security to a design (UNIX) which has a fundamentally poor approach to security (shared global namespace, users, etc.).
It's really not going all that well, and I hope something like SEL4 can replace Linux for cloud server workloads eventually.
Most applications use almost none of the Linux kernel's features. We could have very secure, high performance web servers, which get capabilities to the network stack as initial arguments, and don't have access to anything more.
Drivers for virtual devices are simple, we don't need Linux's vast driver support for cloud VMs. We essentially need a virtual ethernet device driver for SEL4, a network stack that runs on SEL4, and a simple init process that loads the network stack with capabilities for the network device, and loads the application with a capability to the network stack.
Make building an image for that as easy as compiling a binary, and you could eliminate maybe 10s of millions of lines of complexity from the deployment of most server applications. No Linux, no docker.
Because SEL4 is actually well designed, you can run a sub kernel as a process on SEL4 relatively easily. Tada, now you can get rid of K8s too.
- tptacek 1y agoThis makes sense if you look at containers as simply a means to an end of setting up a sandbox, but not really much sense at all if you think of containers as a way to make it easy to get an arbitrary application up and running on an arbitrary server without altering host system dependencies.
- ianburrell 1y agoI suspect that containers would have taken off even without isolation. I think the important innovation of Docker was the image. It let people deploy consistent version of their software or download outside software. All of the hassle of installing things was in the Dockerfile, and it was run in containers so more reliable.
- tptacek 1y agoI agree: I think the container image is what matters. As it turns out, getting more (or less) isolation given that image format is not a very hard problem.
- thaumasiotes 1y ago> I think the important innovation of Docker was the image. It let people deploy consistent version of their software or download outside software. What did it let people do that they couldn't already do with static linking?
- cschep 1y agoI can't tell if this is a genuine question or not but if it is.. deploying a Ruby on Rails app with a pile of gems that have c deps isn't fixed with static linking. This is true for python and node and probably other things I'm not thinking of.
- misnome 1y agoBecause it doesn’t need you to delve deep into the build system of every dependency and application you ever want to package?
- thaumasiotes 1y ago>> It let people deploy consistent version of their software
- zellyn 1y agoAgreed. There was a point where I thought AMIs would become the unit of open source deployment packaging, and I think docker filled that niche in a cloud-agnostic way
- Eikon 1y agomake tinyconfig can get you pretty lean already. https://archive.kernel.org/oldwiki/tiny.wiki.kernel.org/ https://archive.kernel.org/oldwiki/tiny.wiki.kernel.org/
- bombcar 1y agoIs that why containers started? I seem to recall them taking off because of dependency hell, back in the weird time when easy virtualization wasn't insanely available to everyone. Trying to get the versions of software you needed to use all running on the same server was an exercise in fiddling.
- alphazard 1y agoYes, totally agree that's a contributor too. I should expand that by namespaces I mean user, network, and mount table namespaces. The initial contents of those is something you would have to provide when creating the sandbox. Most of it is small enough to be shipped around in a JSON file, but the initial contents of a mount table require filesystem images to be useful.
- ctkhn 1y agoOn a personal level, that's why I started using them for self hosting. At work, I think the simplicity of scaling from a pool of resources is a huge improvement over having to provision a new device. Currently at an on-prem team and even moving to kubernetes without going to cloud would solve some of the more painful operational problems that send us pages or we have to meet with our prod support team about.
- mbreese 1y agoI think there were multiple reasons why containers started to gain traction. If you ask 3 people why they started using containers, you're likely to get 4 answers. For me, it was avoiding dependencies and making it easier to deploy programs (not services) to different servers w/o needing to install dependencies. I seem to remember a meetup in SF around 2013 where Docker (was it still dotCloud back then?) was describing a primary use-case was easier deployment of services. I'm sure for someone else, it was deployment/coordination of related services.
- dabockster 1y agoThe big selling points for me were what you said about simplifying deployments, but also the fact that a container uses significantly less resource overhead than a full blown virtual machine. Containers really only work if your code works in user space and doesn't need anything super low level (eg TCP network stack), but as long as you stay in user space it's amazing.
- tliltocatl 1y agoContainers and namespaces are not about security. They are about not having singleton objects at the OS level. Would have called it virtualization if the word wasn't so overloaded already. There is a big difference that somehow everyone misses. A bypassable security mechanism is worse than useless. A bypassable virtualization mechanism is useful. It is useful to be able to have a separate root filesystem just for this program - even if a malicious program is still able to detect it's not the true root. As about SEL4 - it is so elegant because it leaves all the difficult problems to the upper layer (coincidentally making them much more difficult).
- alphazard 1y ago> As about SEL4 - it is so elegant because it leaves all the difficult problems to the upper layer (coincidentally making them much more difficult). I completely buy this as an explanation for why SEL4 for user environments hasn't (and probably will never) take off. But there's just not that much to do to connect a server application to the network, where it can access all of its resources. I think a better explanation for the lack of server side adoption is poor marketing, lack of good documentation, and no company selling support for it as a best practice.
- frumplestlatz 1y agoThe lack of adoption is because it’s not a complete operating system. Using sel4 on a server requires complex software development to produce an operating environment in which you can actually do anything. I’m not speaking ill of sel4; I’m a huge fan, and things like it’s take-grant capability model are extremely interesting and valuable contributions. It’s just not a usable standalone operating system. It’s a tool kit for purpose-built appliances, or something that you could, with an enormous amount of effort, build a complete operating system on top of.
- josephg 1y agoYes. I really hope someone builds a nice, usable OS with SeL4 as a base. If SeL4 is like the linux kernel, we need a userland (GNU). And a distribution that's simple to install and make use of. I'd love to work on this. It'd be a fun problem!
- ants_everywhere 1y ago> Because SEL4 is actually well designed, you can run a sub kernel as a process on SEL4 relatively easily. Tada, now you can get rid of K8s too. k8s is about managing clusters of machines as if they were a single resource. Hence the name "borg" of its predecessor. AFAIK, this isn't a use case handled by SEL4?
- alphazard 1y agoThe K8s master is just a scheduling application. It can run anywhere, and doesn't depend on much (just etcd). The kublet (which runs on each node) is what manages the local resources. It has a plugin architecture, and when you include one of each necessary plugin, it gets very complicated. There are plugins for networking, containerization, storage. If you are already running SEL4 and you want to spawn an application that is totally isolated, or even an entire sub-kernel it's not different than spawning a process on UNIX. There is no need for the containerization plugins on SEL4. Additionally the isolation for the storage and networking plugins would be much better on SEL4, and wouldn't even really require additional specialized code. A reasonable init system would be all you need to wire up isolated components that provide storage and networking. Kubernetes is seen as this complicated and impressive piece of software, but it's only impressive given the complexity of the APIs it is built on. Providing K8s functionality on top of SEL4 would be trivial in comparison.
- ants_everywhere 1y agoI understand what you're saying, and I'm a fan of SEL4. But isolation isn't one of the primary points of k8s. Containerization is after all, as you mentioned, a plugin. As is network behavior. These are things that k8s doesn't have a strong opinion on beyond compliance with the required interface. You can switch container plugin and barely notice the difference. The job of k8s is to have control loops that manage fleets of resources. That's why containers are called "containers". They're for shipping services around like containers on boats. Isolation, especially security isolation, isn't (or at least wasn't originally) the main idea. You manage a fleet of machines and a fleet of apps. k8s is what orchestrates that. SEL4 is a microkernel -- it runs on a single machine. From the point of view of k8s, a single machine is disposable. From the point of view of SEL4, the machine is its whole world. So while I see your point that SEL4 could be used on k8s nodes, it performs a very different function than k8s.
- orbifold 1y agoIt would be great if we got "kernel independent" Nvidia drivers. I have some experience with bare-metal development and it really seems like most of what an operating system provides could be provided in a much better way as a set of libraries that make specific pieces of hardware work, plus a very good "build" system.
- stinkbeetle 1y agocgroups first came from resource management frameworks that IIRC came out of IBM and got into some distro kernels for a time but not upstream. Namespaces were not an attempt to add security, but just grew out of work to make interfaces more flexible, like bind mounts. And Unix security is fundamentally good, not having namespaces isn't much of a point against it in the first place, but now it does have them. And it's going pretty well indeed. All applications use many kernel features, and we do have very secure high performance web and other servers. L4 systems have been around for as long as Linux, and SEL4 in particular for 2 decades. They haven't moved the needle much so I'd say it's not really going all that well for them so far. SEL4 is a great project that has done some important things don't get me wrong, but it doesn't seem to be a unix replacement poised for a coup.
- onjectic 1y ago> Unix security is fundamentally good L. Ron Hubbard is fundamentally good! I kid, but seriously, good how? Because it ensures cybersecurity engineers will always have a job? seL4 is not the final answer, but something close to it absolutely will be. Capability-based security is an irreducible concept at a mathematical level, meaning you can’t do better than it, at best you can match it, and its certainly not matched by anything else we’ve discovered in this space.
- stinkbeetle 1y ago> good how? Good because it is simple both in terms of understanding it and implementing it, and sufficient in a lot of cases. > seL4 is not the final answer, but something close to it absolutely will be. Capability-based security is an irreducible concept at a mathematical level, meaning you can’t do better than it, at best you can match it, and its certainly not matched by anything else we’ve discovered in this space. Security is not pure math though, it's systems and people and systems of people.
- man8alexd 1y agocgroups are from Google. https://lwn.net/Articles/199643/ https://lwn.net/Articles/199643/
- noduerme 1y agoYou say applications and web servers kind of interchangeably. I don't know anything about SEL4. What if your application needs to spawn and manage executables as child processes? Is it Linux-like enough to run those and handle stuff like that so that those of us coding at the application layer don't need to worry about it too much?
- lmm 1y ago> Containers (meaning Docker) happened because CGroups and namespaces were arcane and required lots of specialized knowledge to create what most of us can intuitively understand as a "sandbox". That might be why Docker was originally implemented, but why it "happened" is because everyone wanted to deploy Python and pre-uv Python package management sucks so bad that Docker was the least bad way to do that. Even pre-kubernetes, most people using Docker weren't using it for sandboxing, they were using it as fat jars for Python.
- procaryote 1y agoNot only python, although python is particularly bad. Even java things wher fatjars exist you at some point end up with os level dependencies like "and this logging thing needs to be set up, and these dirs need these rights, and this user needs to be in place" etc. Nowadays you can shove that into a container
- themafia 1y ago> which has a fundamentally poor approach to security Unix was not designed to be convenient for VPS providers. It was designed to allow a single computer to serve an entire floor of a single company. The security approach is appropriate for the deployment strategy. As it did with all OSes, the Internet showed up, and promptly ruined everything.
- antod 1y agoI don't think Docker came about due to cgroups and namespaces being arcane, LXC was already abstracting that away. Docker's claim to fame was connecting that existing stuff with layered filesystem images and packaging based off that. Docker even started off using LXC to cover those container runtime parts.
- thaumasiotes 1y ago> namespaces were added to Linux in an attempt to add security to a design (UNIX) which has a fundamentally poor approach to security (shared global namespace, users, etc.) If the "fundamentally poor approach to security" is a shared global namespace, why are namespaces not just a fix that means the fundamental approach to security is no longer poor?
- tombert 1y ago> Drivers for virtual devices are simple, we don't need Linux's vast driver support for cloud VMs. We essentially need a virtual ethernet device driver for SEL4, a network stack that runs on SEL4, and a simple init process that loads the network stack with capabilities for the network device, and loads the application with a capability to the network stack. Make building an image for that as easy as compiling a binary, and you could eliminate maybe 10s of millions of lines of complexity from the deployment of most server applications. No Linux, no docker. Wasn't this what unikernels were attempting a decade ago? I always thought they were neat but they never really took off. I would totally be onboard with moving to seL4 for most cloud applications. I think Linux would be nearly impossible to get into a formally-verified state like seL4, and as you said most cloud stuff doesn't need most of the features of Linux. Also seL4 is just cool.
- m463 1y agoseems like all this was part of a long evolution. I think the whole thing has been levels of abstraction around a runtime environment. in the beginning we had the filesystem. We had /usr/bin, /usr/local/bin, etc. then chroot where we could run an environment then your chgroups/namespaces then docker build and docker run then swarm/k8s/etc I think there was a parallel evolution around administration, like configure/make, then apt/yum/pacman, then ansible/puppet/chef and then finally dockerfile/yaml
- man8alexd 1y agoThe irony is that dockerfile/yaml contains so much ugly bash code nowadays that it feels like we are back at configure/make stage.
- theamk 1y agoLuckily, no Dockerfile is ever as bad as old "configure" scripts were. As long I never have to worry about configure snippets that deal with Sun's CC compiler from 1990's, or with gcc-3, I will be happy.
- m463 1y agoIf you are talking about the RUN foo && \ bar && \ baz thing, I completely agree. I've always wondered if there could be something like: LAYER RUN foo RUN bar RUN baz LAYER to accomplish something similar, or maybe: RUN foo AND bar AND baz
- man8alexd 1y agoThere are also health/liveness checks, entry point code, sometimes embedded right in the Helm templates.
- otabdeveloper4 1y ago> which get capabilities to the network stack as initial arguments, and don't have access to anything more Systemd does this and it is widely used.
- zozbot234 1y ago> Cgroups and namespaces were added to Linux in an attempt to add security to a design (UNIX) which has a fundamentally poor approach to security (shared global namespace, users, etc.) Namespacing of all resources (no restriction to a shared global namespace) was actually taken directly from plan9. It does enable better security but it's about more than that; it also sets up a principled foundation for distributed compute. You can see this in how containerization enables the low-level layers of something like k8s - setting aside for the sake of argument the whole higher-level adaptive deployment and management that it's actually most well-known for.
- lproven 1y ago> I hope something like SEL4 can replace Linux for cloud server workloads eventually. Why not 9front and diskless Linux microVMs, Firecracker/Kata-containers style? Filesystem and process isolation in one, on an OS that's smaller than K8s? Keep it simple and Unixy. Keep the existing binaries. Keep plain-text config and repos and images. Just replace the bottom layer of the stack, and migrate stuff to the host OS as and when it's convenient.
- pjmlp 1y agoLinux is already being replaced by type 1 hypervisors on cloud server workloads. Anyone doing deployments in managed languages, regardless of AOT compiled, or using a JIT, the underlying operating system is mostly irrelevant, with exception of some corner cases regarding performance tweeks and such. Even if those type 1 hypervisors happen to depend on Linux kernel for their implementation, it is pretty much transparent when using something like Vercel, or Lambda.
- deleted 1y ago[deleted]
- nisegami 1y agoWhy go a step further and deploy all cloud workloads using webassembly?
- pianopatrick 1y agoThe story I heard was that containers let you use less memory and better share the kernel and CPU compared to Virtual Machines, such that you could run more applications on the same servers. This translates into direct cost savings, which is why large companies with large server farms were willing to pay their engineers to develop the technology and transition to the technology. In terms of security, I think even more secure than SEL4 or containers or VMs would be having a separate physical server for each application and not sharing CPUs or memory at all. Then you have a security boundary between applications that is based in physics. Of course, that is too expensive for most business use cases, which is why people do not use it. I think using SEL4 will run into the same problem - you will get worse utilization out of the server compared to containers, so it is more expensive for business use cases and not attractive. If we want something to replace containers that thing would have to be both cheaper and more secure. And I'm not sure what that would be