5 ms·
Mount Mayhem at Netflix: Scaling Containers on Modern CPUs
- vivzkestrel 7mo ago- can someone kindly explain why there are 2 websites that all claim to be netflix tech blog? - website 1 https://netflixtechblog.medium.com/ https://netflixtechblog.medium.com/ - website 2 https://netflixtechblog.com/ https://netflixtechblog.com/
- geodel 7mo agoI mean Netflix is dealing with big, important things like container scaling, creating a million micro services talking to each other and so on. Having multiple tech blogging platform on Medium is not something they have a spare moment to think about.
- hhh 7mo agoThe second one is a hosted custom domain for the medium blog iirc
- owenthejumper 7mo agoWhy is this so badly AI written? Netflix can surely pay for writers. At this point I refuse to read any content in the AI format of: - The problem - The solution - Why it matters
- deleted 7mo ago[deleted]
- ViktorRay 7mo agoArticles like this are pretty cool. It’s so interesting to see the behind the scenes that happens whenever we watch a Netflix movie.
- haneul 7mo agoInteresting, another case of removing HT improving performance. Reminds me of doing that on Intel CPUs of a few gens ago.
- ahoka 7mo agoIt's quite logical that by saturating your CPU it can only decrease performance.
- spockz 7mo agoIn this case the CPU wasn’t really saturated with work but with contention on global locks. The contention is lessened by removing the amount of concurrent mounts that are being done. I wonder if simply setting a maximum number of concurrent mounts in the code or by letting containerd think there are only half the amount of cores, would have solved the contention to the same amount.
- parliament32 7mo agoIt took them this long to move from docker to containerd?
- rixed 7mo agoI am not familiar with the nitty gritty of container instance building process, so maybe I'm just not the intended audience, but this is particularly unclear to me: > To avoid the costly process of untarring and shifting UIDs for every container, the new runtime uses the kernel’s idmap feature. This allows efficient UID mapping per container without copying or changing file ownership, which is why containerd performs many mounts Why does using idmap require to perform more mount?
- martijnvds 7mo agoThis kind of id mapping works as a mount option (it can also be used on bind mounts). You give it a mapping of "id in filesystem on disk" to "id to return to filesystem APIs" and it's all translated on the fly.
- rixed 7mo agoThank you! Going to ask an LLM to lecture me on this when I have some time; good to see that humans are still the best at giving just the right amount of explanation :)
- nineteen999 7mo agoThe costly process probably explains why they just started injecting ads in my plan where there previously weren't any. And also explains why rather than be leveraged into a more expensive plan to help them pay for their containers, I cancelled my subscription. Not like there's more than 1% content there worth paying for these days anyway.
- yjftsjthsd-h 7mo agoOkay, I'll ask the dumb question: Couldn't you also reduce the number of layers per container? Sure, if you can reuse layers you should, but unless you've done something very clever like 1 package per layer I struggle to think that 50 is really useful?
- gucci-on-fleek 7mo ago> unless you've done something very clever like 1 package per layer I struggle to think that 50 is really useful? 1 package per layer can actually be quite nice, since it means that any package updates will only affect that layer, meaning that downloading container updates will use much less network bandwidth. This is nice for things like bootc [0] that are deployed on the "edge", but less useful for things deployed in a well-connected server farm. [0]: https://bootc-dev.github.io/bootc/ https://bootc-dev.github.io/bootc/
- seabrookmx 7mo agoIt doesn't work this way really? It's called a layer because each layer on top depends on the layers below. If you change the package defined in the bottom most layer, all 49 above it are invalid and need re-pulled or re-built.
- gucci-on-fleek 7mo ago> If you change the package defined in the bottom most layer, all 49 above it are invalid and need re-pulled or re-built. I also initially thought that that was the case, but some tools are able to work around that [0] [1] [2]. I have no idea how it works, but it works pretty well in my experience. [0]: https://github.com/hhd-dev/rechunk/ https://github.com/hhd-dev/rechunk/ [1]: https://coreos.github.io/rpm-ostree/container/#creating-chunked-images https://coreos.github.io/rpm-ostree/container/#creating-chun... [2]: https://coreos.github.io/rpm-ostree/build-chunked-oci/ https://coreos.github.io/rpm-ostree/build-chunked-oci/
- minitech 7mo agoThat’s mostly a Dockerism (and even Docker has `COPY --link` these days). The underlying tech supports independent layers.
- DeathArrow 7mo agoSo using the "old" container architecture could have been better than wasting time implementing the new architecture, dealing with the performance issues and wasting more time fixing the issues?
- seabrookmx 7mo agoThis completely ignores all their reasons to move to the new architecture in the first place. My understanding is that they had a mostly in-house architecture (that predated Kubernetes' rise) and by moving to this new platform, they are now much more closely aligned with standard Kubernetes. They can now utilize EKS for their control plane, and leverage the many community provided features previously unavailable to them.
- spockz 7mo agoInterestingly. This enhancement has been proposed in July 2025, accepted and merged in August 2025, and released in November 2025. The blog post is also from November. And now it shows here.
- s_ting765 7mo agoInteresting blog post. For what it's worth, I count 7 em-dashes used.