6 ms·
Learning containers from the bottom up
- ashater 5y agoGood article, steps one level below container managers like Docker or k8s. Obviously not the indepth of how Linux kernel manages container processes but a good write-up.
- kodah 5y agoThis is a great article. I disagree with this: > Now, when you have a decent understanding of containers - from both the implementation and usage standpoints - it's time to tell you the truth. Containers aren't Linux processes! This is a bit of wordplay, I'm assuming, in absence of a word that defines the operating system features that power the concept of containers. To Linux, there is no (to my knowledge) concept of a "container". The container runtime runs your process(es) as the parent and uses the operating systems features to isolate it and restrict it/them. A virtual machine would just be a full emulated version of this, rather than using the operating system to virtualize the network stack. The author is right in that there is no such thing as a container, but only as much as containing is a thing you do, imo. What users think of containers are still just processes though, and I don't think that's an entirely useless abstraction to be cognizant of.
- spenrose 5y ago> The author is right in that there is no such thing as a container, but only as much as containing is a thing you do, imo. What users think of containers are still just processes though, and I don't think that's an entirely useless abstraction to be cognizant of. Fantastic distillation. Thank you!
- jjtheblunt 5y agowhy not think of them as process (group) spawned with particular parent process setup, in particular the cgroups etc configuration effecting isolation.
- musicale 5y agoIn the bad old days before setns() it was more of a pain to add processes to a container since they had to be children of an existing process in that container.
- otterley 5y agoI would go even further - containers are process trees. They just happen to be process trees with the following attributes: (a) they (usually) have separate namespaces (network/pid/uts/cgroups/mount); (b) they (usually) have dropped capabilities; and (c) they (usually) are in cgroups that have resource reservations and/or limits. Under the hood, that's all containers are!
- pm90 5y agoThis really got me at first. Since I had seen Windows virtual vm on Linux (and vice versa), my mental model was still “full virtualization”. But the processes running in a Linux container are still _linux_ processes, they’re just isolated (fairly) well.
- jandrewrogers 5y agoThe container runtime intercepts some syscalls, altering the observable behaviors of the kernel in ways that can adversely impact software that is otherwise perfectly designed to operate outside the container runtime. Normal processes don’t have their syscalls intercepted and this is material difference to the extent it is not transparent. If running the same properly designed software exhibits material differences in behavior between a bare metal process and a containerized process, then they aren’t the same as a matter of practical semantics. Ironically, virtualized processes have much closer equivalence to a bare metal process than containerized processes in practice. Saying a container is “just a process” is like saying a virtual machine is “just a process”, both are true in some sense depending on how you define “process”. But as a matter of practical engineering, they are different kinds of things.
- kodah 5y agoAre you talking about SecComp and namespacing?
- jandrewrogers 5y agoThe root cause is likely SecComp. The notoriously poor I/O performance of containerized code, regardless of configuration, is largely a side effect of syscall interception. In particular it breaks software that does I/O scheduling in user space, which is idiomatic and explicitly supported by the Linux kernel, even on virtual machines, but this use case conflicts with the container abstraction so runtimes offer an ersatz version that allows the code to run albeit poorly.
- yencabulator 5y agoMore likely you're running your containers using overlay or some FUSE thing, and that's causing the I/O slowdown. What syscalls do you think are intercepted, how? Speaking as someone who can write kernel code, I'm not aware of any such thing specific to containers. (As far as the linux kernel is concerned, there's no such thing as a container.) If you're talking about BPF, that can be used outside of containers, e.g. systemd can limit any unit, and using it is not part of a definition of what a container is.
- stevebmark 5y agoI think it’s important to understand that containers aren’t Linux processes. Containers can run more than one process. Containers can be stopped and restarted even though the initial process is gone forever. And containers have their own isolated writable layer.
- kuizu 5y agoA nice blog series explaining in detail each Linux kernel mechanism making up containers: https://www.schutzwerk.com/en/43/posts/linux_container_intro/ https://www.schutzwerk.com/en/43/posts/linux_container_intro...
- otterley 5y agoAgreed - this is a far more comprehensive, logical, and technically correct explanation of how containers work under the hood.
- musicale 5y agoDocker and Kubernetes embody a number of design decisions that might be a good fit for some users (and for Google) but add more complexity and overhead than I usually need or want for my typical use case of basic isolation and resource limits. Fortunately the container architecture is flexible so that you can use as much or as little of it as you like. I also tend to think that if you want stronger isolation for security purposes then you will want a lightweight VM rather than a container (and if you are worried about side channels, probably hardware partitioning - good luck.)
- porker 5y agoFor a quick overview of containers I found https://wizardzines.com/zines/containers/ https://wizardzines.com/zines/containers/ super helpful.
- yencabulator 5y ago> ... but containers are needed to build images Incorrect. The images are mere files(/subtrees), and you can write one however you wish.