9 ms·
Cache is King: A guide for Docker layer caching in GitHub Actions
- notnmeyer 2y agothis is pretty neat—it’s been a while since i’ve tried caching layers with gha. it used to be quite frustrating. my previous experience was that in nearly all situations the time spent sending and retrieving cache layers over the network wound up making a shorter build step moot. ultimately we said “fuck it” and focused on making builds faster without (docker layer) caching.
- adityamaru 2y agoYeah that still holds true to some extent today with the GHA cache. Blacksmith colocates its cache with our CI runners, and ensures that they're in the same local network allowing us to saturate the NIC and provide much faster cache reads/writes. We're also thinking of clever ways to avoid downloading from a cache entirely and instead bind mount cache volumes over the network into the CI runner. Still early days, but stay tuned!
- notnmeyer 2y agothere’s probably a cool consistent hashing solution where jobs are routed to a host that that is likely to have the cache stored locally already and can be mounted into the containers.
- kylegalbraith 2y agoYup! We observed the same thing back before we built Depot. The act of saving/loading cache over a GHA network pretty much negated any performance gain from layer caching. So, we created a solution to persist cache to NVMe disks and orchestrate that across builds so it's immediately available on the next build. All the performance of layer caching without any network transfer. The registry cache idea is a neat idea, but in practice suffers the same problem.
- notnmeyer 2y agototally, your approach is the right one and anything reasonable is going to focus on collocating the cache as close as possible to where the build runs.
- remdoWater 2y agotitle suggestion: Cache Rules Everything Around Me (C.R.E.A.M.)
- maxmcd 2y agoI weep for this period of time where we don't have sticky disks readily available for builds. Uploading the layer cache each time is such a coarse and time-consuming way to cache things. Maybe building from scratch all the time is a good correctness decision? Maybe stale values in disks is a tricky enough issue to want to avoid entirely? If you keep a stack of disks around and grab a free one when the job starts you'd end up with good speedup a lot of the time. If cost is an issue you can expire them quickly. I regularly see CI jobs spending >50% of their time downloading the same things, or compiling the same things, over and over. How many times have I triggered an action that compiled the exact same sqlite source code? Tens of thousands? Maybe this is fine, I dunno.
- parentheses 2y agoI agree. The notion that everything must be docker is nice in principle but requires a lot of performance optimization work early on. Earlier than one would need with "sticky disks" as you called them.
- adityamaru 2y agoThis is exactly the sort of insight that led us to work on Blacksmith. Since we own the hardware we run CI jobs on there are some exciting things we can do to make these "sticky disks" work the way you describe it. Stay tuned!
- deleted 2y ago[deleted]
- withinboredom 2y agoInteresting. I remember working on a project where the first clean build would always fail, and only incremental builds could succeed. I was a junior at the time, so this was 15-20 years ago. I remember spending some time trying to get it to succeed from a clean build and my lead pulling me aside: he said it was an easy fix, but if we fixed it, the ops guys would insist on building from scratch for every build. So please, stop. Personally, unless you have an exotic build env, it’s usually faster and easier to simply build in the runner. If you need a docker image at the end, build a dockerfile that simply copies the artifacts from disk.
- cpfohl 2y agoThis is wild. I've spent the last three weeks working on this stuff for two separate clients. Important note if you're taking advice: cache-from and cache-to both accept multiple values. Cache to just ouputs the cache data to all the ones specified. cache-from looks for cache hits in the sources in-order. You can do some clever stuff to maximize cache hits with the least amount of downloading using the right combination.
- adityamaru 2y agooh TIL, that is interesting
- Arbortheus 2y agoThat’s a great idea.
- user- 2y agoI feel like ive seen so many new companies just providing cheaper github actions.
- clintonb 2y agoThey provide Actions _runners_ because GitHub runners are quite expensive (per CPU and GB of memory) compared to the underlying cost of a Kubernetes node on most cloud providers or bare metal. Of course, that assumes you’ve already paid the cost to setup a cluster, which is not free.
- aayushshah15 2y ago(blacksmith co-founder here) it's unfortunate the amount of expertise / tinkering required to get "incrementalism" in docker builds in github actions. we're hoping to solve this with some of the stuff we have in the pipeline in the near future.
- damianh 2y agoThe fact that GitHub don't provide a better solution here has to be actually costing them money with the network usage and extra agent time consumed. Right?
- aayushshah15 2y agoGitHub has perverse incentives to not fix this problem because they charge customers based on usage (by the minute), so they make more money by providing slower builds to end-users.
- krainboltgreene 2y agoThey've also just completely refocused to AI in the last two years thanks to the microsoft/ChatGPT situation.
- boronine 2y agoI've spent days trying all of these solution at my company. All of these solutions suck, they are slow and only successful builds get their layers cached. This is a dead end. The only workable solution is to have a self-hosted runner with a big disk.
- aayushshah15 2y agohow do you ensure isolation between runs on a self hosted runner that way?
- boronine 2y agoWhat kind of isolation do you need? We are building our own code so I don't see the need for isolation beyond a clear directory.
- airspeedjangle 2y agoShared runner infrastructure in a big company. It's pretty common to treat these situations as multi-tenant low trust environments.
- damianh 2y agoPlenty of marketplace actions will install things and/or mutate the runner. It's a matter of time before someone does something or there's a build that doesn't cleannup after itself (e.g. leaving test processes running) that ruins the day for everyone else.
- damianh 2y agoSelh-hosted runners can be ephemeral too. With such either mount the cache as a disk or bake docker layers/images into the runner image.
- remdoWater 2y agoThis requires a lot of work from a dev inf team, though. Not as straightforward for an average team.
- solatic 2y agoAs someone who spent way too much time chasing this rabbit, the real answer is Just Don't. GitHub Actions is a CI system that makes it easy to get started with simple CI needs but runs into hard problems as soon as you have more advanced needs. Docker caching is one of those advanced needs. If you have non-trivial Docker builds then you simply need on-disk local caching, period. Either use Depot or switch to self-hosted runners with large disks.
- adityamaru 2y agototally agree, github actions has done an excellent job at this lowest layer of the build pipeline today but is woefully inadequate the minute your org hits north of 50 engineers
- aayushshah15 2y agoDid you consider using a local (in the same VPC) docker registry mirror perhaps? https://docs.docker.com/docker-hub/mirror/ https://docs.docker.com/docker-hub/mirror/
- solatic 2y agoIt's not the pulls that are the problem, it's caching intermediate layers from the build that is the problem. As soon as you introduce a networked registry, the time it takes to pull layers from the registry cache and push them back to the registry cache are frequently not much better than simply rebuilding the layers, not to mention the additional compute/storage cost of running the registry cache itself. It's just a problem that requires big, local disks to solve.
- DanielHB 2y agoYeah I had the exact same problem and came to the same conclusion.
- bushbaba 2y agoCan’t you use s3 + mountpoint for most distributed CI cache needs?
- manx 2y agoEarthly solves this really well: https://earthly.dev https://earthly.dev They rethink Dockerfiles with really good caching support.
- oftenwrong 2y agoThe caching support is mostly the same. Both Earthly and Dockerfile are BuildKit frontends. BuildKit provides the caching mechanisms. A possible exception is the "auto skip" feature for Earthly Cloud, since I do not know how that is implemented.
- adamgordonbell 2y agoAlso CACHE keyword, for cache mounts. Makes incremental tools like compilers work well in the context of dockerfiles and layer caches. That can extend beyond just producing docker iamges as well. Under the covers the CACHE keyword is how lib/rust in Earthly makes building Rust artifacts in CI faster. https://github.com/earthly/earthly/issues/1399 https://github.com/earthly/earthly/issues/1399
- vladaionescu 2y agoI would add that 1. Earthly is meant for full CI/CD use-cases, not just for image building. We've forked buildkit to make that possible. And 2. remote caching is pretty slow overall because of the limited amount of data you can push/pull before it becomes performance-prohibitive. We have a comparison in our docs between remote runner (e.g. Earthly Satellites) vs remote cache [1]. [1]: https://docs.earthly.dev/docs/caching#sharing-cache https://docs.earthly.dev/docs/caching#sharing-cache
- ValtteriL 2y agoDocker layer caching is one of the reasons I moved to Jenkins 2 years ago and have been very happy with it for the most part. I only need to install utils once and all build time goes to building my software. It even integrates nicely with Github. Result: 50% faster feedback. However, it needs a bit initial housekeeping and discipline to use correctly. For example using Jenkinsfiles is a must and using containers as agents is desirable.
- adityamaru 2y agodo you self host your jenkins deployment in your AWS account?
- ValtteriL 2y agoSelf host
- remdoWater 2y agowhat do you mean by discipline here?
- ValtteriL 2y agoBasically using exclusively declarative pipelines with Jenkinsfiles in SCM, avoiding cluttering Jenkins with tools aside from docker, keeping Jenkins up to date and protected with proper auth. Jenkins is the most flexible automation platform and its easy to do things in suboptimal ways (eg. Configuring jobs using the GUI). There's also a way to configure Jenkins the IaC way and I am hoping to dig into that at some point. The old way requires manual work that instictly feels wrong when automating everything else.
- dboreham 2y agoNo! So much time spent debunking such broken "caching" solutions. Computers are very fast now. Use proper package/versioning systems (part of the problem here is that those are often also broken/badly designed).
- aayushshah15 2y agoThis is simply false. For starters, GitHub actions by default run on Intel Haswell chips from 2014 (in some cases). Secondly, hardware being faster doesn't obviate the need for caching, especially for docker builds where your layer pulls are purely network bound.
- kbolino 2y ago"Computers are very fast now" is largely because of caching. The CPU has a cache, the disk drive has a cache, the OS has a cache, the HTTP client has a cache, the CDN serving the content has a cache, etc. There may be better ways to cache than at the level of Docker image layers, but no caching is the same as a cache miss on every request, which can be dozens, hundreds, or even thousands of times slower than a cache hit.
- tanepiper 2y agoI have this set up in our pipeline, we also build the image early and use assets to move it between jobs. We've also just switched to self-hosted runners, so might look into shared disk. But in the long run, as annoying as it is out build pipelines reduced but quite a few minutes per build.
- adityajp 2y ago(Co-founder of Blacksmith here) Glad it worked really well for you. What made you switch to self-hosted runners?
- jpgvm 2y agoThe trick to Docker (well OCI) images is never under any circumstance use `docker build` or anything based on it. Dockerfile is your enemy. Use tools like Bazel + rules_oci or Gradle + jib and never spend time thinking about image builds taking time at all.
- dindresto 2y ago+1 to this, migrating our build setup to Nix + nix2container decreased our pipeline duration for incremental changes by a lot, thanks to Nix's granular caching abilities.
- jpgvm 2y agoYeah I really need to actually sit down and learn Nix, seems like it can solve this in a more general way for cases where the thing you want to run is packaged for Nix already.
- Arbortheus 2y agoPlease no! Do not use Bazel unless you have a platform team with multiple people who know how to use it - e.g. large Google-like teams. We had “the Bazel guy” in our mid-sized company that Bazelified so many build processes, then left. It has been an absolute nightmare to maintain because no normal person has any experience with this tooling. It’s very esoteric. People in our company have reluctantly had to pick up Bazel tech debt tasks, like how the rules_docker package got randomly deprecated and replaced with rules_oci with a different API, which meant we could no longer update our Golang services to new versions of Go. In the process we’ve broken CI, builds on Mac, had production outages, and all kinds of peculiarities and rollbacks needed that have been introduced because of an over-engineered esoteric build system that no one really cares about or wanted.
- jpgvm 2y agoBazel isn't for everyone which is why I suggested using any similar tool, jib, Nix, etc. Just not Dockerfile (or if you are going to use Dockerfile only use ADD). Also just because you don't have experience with something doesn't make it a bad choice. I would recommending understanding it first, why your coworker chose it and how other tools would actually do in the same role, grass is often greener on the other side until you get there. Personally I went through a bit of an adventure with Bazel. My first exposure to it was similar to yours, was used in a place I didn't understand for reasons I didn't understand, broke in ways I didn't understand and (regretfully) didn't want to spend time understanding. The reality was once I sat down to use it properly and understood the concepts a lot of things made sense and a whole bunch of very very difficult things became tractable. That last bit is super important. Bazel raises the baseline effort to do something with the build system, which annoys people that don't want to invest time in understanding a build system. However it drastically reduces the complexity of extremely difficult things like fully byte for byte reproducible builds, extremely fast incremental builds and massive build step parallelization through remote build execution.
- glenjamin 2y agoDocker layer caching is complicated! CircleCI has an implementation that used to use a detachable disk, but that had issues with concurrency It’s since been replaced with an approach that uses a docker plugin under the hood to store layers in object storage https://circleci.com/docs/docker-layer-caching/ https://circleci.com/docs/docker-layer-caching/
- remram 2y agoIs it any better than buildx cache then, that also stores caches in object storage (via OCI registry)?
- mshekow 2y agoI took a detailed look at Docker's caching mechanism (actually: BuildKit) in this article https://www.augmentedmind.de/2023/11/19/advanced-buildkit-caching/ https://www.augmentedmind.de/2023/11/19/advanced-buildkit-ca... There I also explain that IF you use a registry cache import/export, you should use the same registry to which you are also pushing your actual image, and use the "image-manifest=true" option (especially if you are targeting GHCR - on DockerHub "image-manifest=true" would not be necessary).
- remram 2y agoThanks, this is a very thorough explanation. Is there really no way to cache the 'cachemount' directories?
- mshekow 2y agoThe only option I know is to use network shares/disks, but you need to make sure that each share/disk is only used by one BuildKit process at a time.
- daulis 2y agoAfter years of lurking, I made an account to reply to this "image-manifest=true" was the magic parameter that I needed to make this work with a non-DockerHub registry (Artifactory). I spent a lot of time fighting this, and non-obvious error messages. Thank you!! We use a multi-stage build for a DevContainer environment, and the final image is quite large (for various reasons), so a better caching strategy really helps in our use case (smaller incremental image updates, smaller downloads for developers, less storage in the repository, etc)
- tkiolp4 2y agoDocker has been among us for years. Why isn’t efficient caching already implemented out of the box? It’s a leaking abstraction that users have to deal with. Annoying at best.
- omeid2 2y agoEfficient caching exists when caching makes sense, layers are meaningfully cached. What most people need but don't use is base layers that are upstream of their code repo and released regularly, not at each commit. Containerisation has made reproducible environments so easy that people want to reproduce it at each CI run, a bit too much.
- joe0 2y agoThey actually have recently, but it’s a separate (payed) product offering: https://www.docker.com/products/build-cloud/ https://www.docker.com/products/build-cloud/
- ikekkdcjkfke 2y agoAnybody have a build system that builds as fast or faster than locally?
- bagels 2y agoI think that bitbucket offers this out of the box on their CI product (pipelines)
- Cloudef 2y ago- uses: DeterminateSystems/nix-installer-action@main - uses: DeterminateSystems/magic-nix-cache-action@main
- spurin 2y agoIt would be possible to offload the caching to Docker Build Cloud transparently, it’s part of the Docker subscription service, every account gets free minutes - 50 free minutes a month so depending on usage, you may be able to get this at zero cost. With this approach, you’d use buildx and remotely, they would manage and maintain cache amongst other benefits. It does require a credit card signup (which takes $0 to mitigate fraud). Full transparency, I’m a Docker Captain and helped test this whilst it was called Hydrobuild.
- SJC_Hacker 2y agoSmaller images are also another way to go. At my last company, image sizes were like 2-3Gb. I was able to prune that down to ~1.5 GB. Boost and a custom clang/llvm build were particular major offenders here. There's quite a bit of cruft that can be pruned.