3 ms·
I'm working at a shop where we have to use extremely large images (10 - 40 GiB) on a platform of 60+ medium sized blades (4 - 12 cores, 32 - 256 GiB RAM, 200 -
by El_RIDO 5y ago
I'm working at a shop where we have to use extremely large images (10 - 40 GiB) on a platform of 60+ medium sized blades (4 - 12 cores, 32 - 256 GiB RAM, 200 - 1000 GiB SSDs). Startup time for running such huge images often is 5 - 10 minutes. They contain prepackaged test data so that they can be run in parallel against newly built versions of our software - there are about 30ish of these images.
You not only store these images once in the registry, but on all container hosts that they might run on. And you need to transmit them, if they aren't cached there, yet. And you may want to redeploy new images several times a day, as they get updated, new ones added, etc.
To be fair, the above is not the use case containers were originally intended for. And if your 10 GiB container image only runs on 3 or 5 nodes and only gets updated once a week, you obviously don't have to worry about the overhead too much.
But at scale (beyond 10 nodes) and/or on very fast development cycles (=more then one deployment per day), size IMHO starts to matter.
- hardwaresofton 5y agoI can totally see this starting to matter, but I do want to point out that it hasn't ground your business to a halt just yet and is very very unlikely to. From what I've seen most ops pain comes from a lack of automation when deploying -- as in if the process is 5-10 or even 15 minutes, no one really cares (outside of an emergency) if it's completely automated. Do you find that is true with your work as well? I can certainly see that it can be a painful problem (10-40GiB is huge), but I expect most smaller shops that aren't running on their own servers/colo as such these image sizes were never even an option (I can't imagine trying to upload 40GB across the public internet to launch one instance!).
- ethbr0 5y agoI've worked in the general automation space for awhile, and agree that time-to-result is rarely the limiting factor. It's also not usually total-person-clock-time. It's number-of-distinct-manual-steps. (5 minutes manual) + (15 minutes automation) + (5 minutes manual) + (30 minutes automation) = everyone hates doing it (10 minutes manual) + (45 minutes automation) = everyone's happy and can use their time productively
- js8 5y agoBecause every manual step means things can go wrong. You might also need to look it up if you don't do it often. All these increase the overhead of manual steps compared to clock time.
- ethbr0 5y agoPartly. I think it's also because manual steps require synchronization. Not only do I have to do something, I also have to set an alarm / watch the clock to know when to do something.
- tiagod 5y agoAnd the CI with many-GB Docker images is very painful... Turning on layer caching usually makes the process even slower as it needs to pull and unpack the previous image before starting, and if you turn it off you're downloading a ton of deps on every build. If you separate the heavy stuff into a base-image you still have to load it on the beginning of CI, which without beefy machines with local SSD caching can take a loooong time.
- daniellarusso 5y agoSo, is there a point where it makes sense to use virtual machines instead of containers?
- El_RIDO 5y agoFor this particular use case: We previously used VMs and snapshots for that workload. The problems we encountered were: - snapshots aren't really intended to be portable (we used ESXi and also KVM on LVM backed volumes), hence we had to write and maintain tooling to have a "snapshot repository", versioning for these and distributing them to the target nodes - that did all work, but was even slower (startup times, which can include shipping and setting up the snapshot, was 15 - 60 minutes) - a VM will duplicate all of the services and the kernel, so all of that has to be started as well, where as the containers only start the services under test. - Using VMs makes it much more difficult for our developers and QA to retreive a particular image and replicate a failing test on a given version of the tested software locally, especially when considering a wide range of client OSs we see in our not that large group (MacOS, various Linux distros, even Windows) A general observation on trade-offs: With containers layering and COW you get fast development, but have to pay the performance bill when you happen to download and apply an image for the first time. Similarly, taking a snapshot of a LV under LVM or on a ESXi VM is fast, but applying the snapshot is slow. We had therefore at one point considered using Ceph RDB and their COW clone snapshots[0]. It would let us do "cheap" restores of snapshots. Our initial tests showed that the network bandwidth requirements[1] would have needed some serious infrastructure re-engineering in order to keep up with our I/O expectations. And again, the containers slot in nicely with commonly available local resources and allow working offline, to a degree. [0] https://docs.ceph.com/en/latest/rbd/rbd-snapshot/ https://docs.ceph.com/en/latest/rbd/rbd-snapshot/ [1] the environment in question uses 10 - 40 Gb/s network links, 40 on the gitlab git and container registry side, 10 on the blade side.
- daniellarusso 5y agoI appreciate the response and you answered my next question about the networking specs!
- 5y ago