3 ms·
RE small images: advice like this always seems to ignore that Docker images are cached incrementally. If you build, say, a Go binary and then put it in a popula
by returningfory2 5y ago
RE small images: advice like this always seems to ignore that Docker images are cached incrementally. If you build, say, a Go binary and then put it in a popular stock Debian or Alpine Docker image, pushing the final image to a registry will only transfer the diff over the network. The size of the diff is basically the compressed Go binary and doesn’t depend on the base image you’re using!
Then when you pull the Docker image onto a machine it’s likely the base image is already in cache so again it’s only the diff that is transferred.
It’s sometimes said this Alpine stuff is a premature optimization, but in most cases it’s not even an optimization.
- varikin 5y ago> It’s sometimes said this Alpine stuff is a premature optimization, but in most cases it’s not even an optimization. Not always. If you have thousands of containers running, the network cost* for changing the base image is massive because each instance needs possibly all new layers, not just the last layer or 2. Having a smaller base image is very practical at this point. * And for network cost, there is both the literal cost with egress from dockerhub, or another registry, but also the load on your network infra.
- p_l 5y agoActually using that caching in a way that benefits you is, well, very hard. Especially since a lot of the time you have a lot of stuff included in order to start the final program, and every step above the base image will result in different layer that will be separate and probably not cacheable Also, Alpine images are used for being, well, small in general - there's not much to exploit on such image, nor is there going to be a lot of unplanned-for data, and especially it will speed up deployment when you upload a new image. It takes special, dedicated CI/CD work to have common base images to exploit sharing of layers properly.
- justin_oaks 5y agoWhile that's true, you or your company have to standardize on a base image to get benefit of shared image layers. In my job, I rarely have an exact match on the base layer because I'll use images from Docker Hub. Although many of those images use the same small set of distros (Alpine, Debian, or Ubuntu), any given image is likely from a different base version of that distro, thus there is no common base image. There may be differences in the base image even if you use the same tag since the same image may change over time. Thus if you pull debian:bullseye today, and then pull again tomorrow, you may have two different images. Generally I don't bother trying to get a common base image and I'm happy enough just attempt to lower the overall size of my images.
- nickjj 5y ago> It’s sometimes said this Alpine stuff is a premature optimization, but in most cases it’s not even an optimization. I do use Debian Slim over Alpine without thinking but image size does matter in a lot of cases once you leave your dev box or start to use Kubernetes. For example if you're using AWS Fargate images aren't cached[0]. That means every time you deploy your app a new pod is going to get created and it has to wait for your image to download. Fargate already has a 30-40ish second spin up time inherit to the platform. Having to wait X seconds or even minutes more to download your app's image delays things quite a lot, especially if you have something like a pre-sync hook running to do something like running a database migration. Now every deploy pays this penalty twice. [0]: https://github.com/aws/containers-roadmap/issues/696 https://github.com/aws/containers-roadmap/issues/696