9 ms·
Faster Gitlab CI/CD pipelines
- benatkin 5y agoI see this part: key: yarn.lock Is this going to make it start without a cache every time that yarn.lock changes? Isn't that a bit overkill? Normally `yarn install` only downloads updated packages.
- dnsmichi 5y agoCache keys are unique identifier names for a cache, as otherwise a global cache is used by all jobs, 'key: default'. https://docs.gitlab.com/ee/ci/yaml/index.html#cache https://docs.gitlab.com/ee/ci/yaml/index.html#cache https://docs.gitlab.com/ee/ci/yaml/index.html#cachekey https://docs.gitlab.com/ee/ci/yaml/index.html#cachekey The key identifier can also be the job name, the branch as commit ref, or something else unique for this job, or project pipeline. The example in the blog post could also use key: yarn-cache-$CI_COMMIT_REF_SLUG to better reflect its purpose. GitLab 13.11 added support for multiple cache keys per job: https://about.gitlab.com/releases/2021/04/22/gitlab-13-11-released/#use-multiple-caches-in-the-same-job https://about.gitlab.com/releases/2021/04/22/gitlab-13-11-re... If you are using a monorepo, or work with submodules and different package systems, you may have Python, Ruby, NodeJS in the same CI job. By default, a defined cache needed all 'path' entries as a list, using the same global cache. Specifying multiple keys with different path locations allows to keep caches separated, and as such, better performance for each specific job. Some jobs may not need the NodeJS, and can only specify to use the Ruby cache key for example. In case you like to invalidate the cache every time a specific file (yarn.lock, go.sum, etc) changes, you can explicitly configure this behavior using cache:keys:files https://docs.gitlab.com/ee/ci/yaml/index.html#cachekeyfiles https://docs.gitlab.com/ee/ci/yaml/index.html#cachekeyfiles This can help prevent corrupted caches, e.g. having downloaded older packages which are stale and not used by current dependency trees. Your code still optionally imports them, and jobs fail because of the old dependency. You cannot reproduce that problem in your dev environment though, starting with a fresh container and no caches. I have been debugging these things before, it takes a while to identify local job caches as the culprit. That said, suggesting to go with a little less performance gain and invalidate caches when dependencies change - if it makes sense for the package manager with often changing recursive dependencies. I've seen it with Python. Tip for failing jobs - by default, the caches are not saved, meaning to say, a large pip install command remains volatile, even if only the user defined unit test command failed afterwards. To avoid a slow down in the pipeline, you can use cache:when:always to always save the cache. https://docs.gitlab.com/ee/ci/yaml/index.html#cachewhen https://docs.gitlab.com/ee/ci/yaml/index.html#cachewhen Exercises to learn with Python are at slide 109 https://docs.google.com/presentation/d/12ifd_w7G492FHRaS9CXAXOGky20pEQuV-Qox8V4Rq8s/edit#slide=id.gf2ada56f71_0_303 https://docs.google.com/presentation/d/12ifd_w7G492FHRaS9CXA...
- benatkin 5y agoGood information! > In case you like to invalidate the cache every time a specific file (yarn.lock, go.sum, etc) changes, you can explicitly configure this behavior using cache:keys:files https://docs.gitlab.com/ee/ci/yaml/index.html#cachekeyfiles https://docs.gitlab.com/ee/ci/yaml/index.html#cachekeyfiles They're using cache:keys:files so it will be installing all of them each time yarn.lock changes. When a build is triggered where yarn.lock hasn't changed, it does a build without downloading all the packages. Come to think of it, builds don't always run in chronological order, so it could wind up with extra packages. Yarn has autoclean, but it says to avoid using it. NPM seems to be quite OK with it, though: https://docs.npmjs.com/cli/v7/commands/npm-prune https://docs.npmjs.com/cli/v7/commands/npm-prune I think caching two folders - one that contains the downloads and one that contains the installed packages - could be the way to go. Yarn and npm have caches to prevent downloading files. And maybe only cache the downloads on the main branch.
- dnsmichi 5y agoThank you for the great thoughts :) > And maybe only cache the downloads on the main branch. $CI_COMMIT_REF_SLUG resolves into the branch when executed in a pipeline. Using it as value for the cache key, Git branches (and related MRs) use different caches. It can be one way to avoid collision but requires more storage with multiple caches. https://docs.gitlab.com/ee/ci/variables/predefined_variables.html https://docs.gitlab.com/ee/ci/variables/predefined_variables... In general, I agree, the more caches and parallel execution you add, the more complex and error prone it can get. Simulating a pipeline with runtime requirements like network & caches needs its own "staging" env for developing pipelines. That's a scenario not many have, or might be willing to assign resources onto. Static simulation where you predict the building blocks from the yaml config, is something GitLab's pipeline authoring team is working on in https://gitlab.com/groups/gitlab-org/-/epics/6498 https://gitlab.com/groups/gitlab-org/-/epics/6498 And it is also a matter of insights and observability - the critical path in the pipeline has a long max duration, where do you start analysing and how do you prevent this scenario from happening again. Monitoring with the GitLb CI Pipeline Exporter for Prometheus is great, another way of looking into CI/CD pipelines can be tracing. CI/CD Tracing with OpenTelemetry is discussed in https://gitlab.com/gitlab-org/gitlab/-/issues/338943 https://gitlab.com/gitlab-org/gitlab/-/issues/338943 to learn about user experiences, and define the next steps. Imho a very hot topic, seeing more awareness for metrics and traces from everyone. Like, seeing the full trace for pipeline from start to end with different spans inside, and learning that the container image pull takes a long time. That can be the entry point into deeper analysis. Another idea is to make app instrumentation easier for developers, providing tips for e.g. adding /metrics as an http endpoint using Prometheus and OpenTelemetry client libraries. That way you not only see the CI/CD infrastructure & pipelines, but also user side application performance monitoring and beyond in distributed environments. I'm collecting ideas for blog posts in https://gitlab.com/gitlab-com/marketing/corporate_marketing/corporate-marketing/-/issues/5677 https://gitlab.com/gitlab-com/marketing/corporate_marketing/... For someone starting with pipeline efficiency tasks, I'd recommend setting a goal - like shown in the blog post X minutes down to Y - and then start with analysing to get an idea about the blocking parts. Evaluate and test solutions for each part, e.g. a terraform apply might depend on AWS APIs, whereas a Docker pull could be switched to use the Dependency proxy in GitLab for caching. Each environment has different requirements - collect helpful resources from howtos, blog posts, docs, HN threads, etc. and also ask the community about their experience. https://forum.gitlab.com/ https://forum.gitlab.com/ is a good spot too. Recommend to create an example project highlighting the pipeline, and allowing everyone to fork, analyse, add suggestions.
- aetherspawn 5y agoIt’s hard to know whether to cache CI or not. On one hand, without the cache builds can be very slow. But on the other hand, you’ll see in a lot of projects random commits like “blow away corrupted cache”, which makes you wonder whether building the cache from scratch is an important step of reproducible builds. I personally rather let the builds run longer and be absolutely certain. Maybe there’s a good middle ground for dev commits vs final merge commits, but unfortunately there’s no machinery in ie GitHub to specify a commit as final before merge.
- whazor 5y agoI remember there is an empty cache button in the UI of Gitlab.
- dnsmichi 5y agoYep, in the pipeline view at the top right. https://docs.gitlab.com/ee/ci/caching/#clearing-the-cache https://docs.gitlab.com/ee/ci/caching/#clearing-the-cache
- exdsq 5y agoIt's not perfect but you could have a word in the commit message that the pipeline looks for and acts upon, so default using cache but let you not use it with "NO-CACHE: <message>"
- boleary-gl 5y agoGitLab team member here - thanks for this write up! Great to see your thought process throughout.
- john_cogs 5y ago+1 - thanks for sharing!
- stabbles 5y agoMy impression is that Github actions are more convenient, as jobs are split into steps, and steps share the filesystem state
- iechoz6H 5y agoNot so convenient if your code is on GitLab.
- dnsmichi 5y agoGreat post, thanks for sharing. We should link that in the Pipeline Efficiency docs: https://docs.gitlab.com/ee/ci/pipelines/pipeline_efficiency.html https://docs.gitlab.com/ee/ci/pipelines/pipeline_efficiency.... I've given a talk about similar ideas for efficient pipelines at Continuous Lifecycle, the slides have many URLs inside to learn async: https://docs.google.com/presentation/d/1nq7Q4WMv6rQc6WFJCRqjtgtgQfgtKTtwSFw4IXXuRsA/edit https://docs.google.com/presentation/d/1nq7Q4WMv6rQc6WFJCRqj... And if you want dive deeper, a free full day workshop with exercises to practice config, resource, caches, container images and more. I've created it for the Open Source Automation Days in early October. Slides with exercises: https://docs.google.com/presentation/d/12ifd_w7G492FHRaS9CXAXOGky20pEQuV-Qox8V4Rq8s/edit https://docs.google.com/presentation/d/12ifd_w7G492FHRaS9CXA... Exercises+solutions: https://gitlab.com/gitlab-de/workshops/ci-cd-pipeline-efficiency-workshop/pipeline-efficiency-workshop-exercises https://gitlab.com/gitlab-de/workshops/ci-cd-pipeline-effici... I did not have time yet to write a blog post sharing more insights on the exercises, but they should be self-explaining on the slides, with solutions in the repository. Let me know how it goes, feel free to repurpose for your own blog posts, and send documentation updates please :)
- dnsmichi 5y ago> We should link that in the Pipeline Efficiency docs I've shared the resources and this HN topic with GitLab's technical writing team to make tutorials more visible on docs.gitlab.com https://gitlab.com/gitlab-org/technical-writing/-/issues/511#note_763331836 https://gitlab.com/gitlab-org/technical-writing/-/issues/511...
- iduoad 5y agoSo much to learn here, thank you for sharing!
- dnsmichi 5y agoI went ahead and blogged about the workshop, thanks y'all for the inspiration :) https://dev.to/dnsmichi/efficient-devsecops-pipelines-in-a-cloud-native-world-free-workshop-hmk https://dev.to/dnsmichi/efficient-devsecops-pipelines-in-a-c... If you learn a new trick or gem, please blog and share with our community :-) Overview of the topics inside the workshop: - Introduction: CI/CD meets Dev, Sec and Ops - CI/CD: Terminology and first steps - Analyse & Identify - Learn using the GitLab CI Pipeline Exporter to monitor the exercise project throughout the workshop. - Efficiency actions - Config Efficiency: CI/CD Variables in variables, job templates (YAML anchors, extends), includes (local, remote), rules and conditions (if, dynamic variables, conditional includes), !reference tags (script, rules), maintain own CI/CD templates (include templates, override config values), parent-child pipelines, multi project pipelines, better error messages to fix failures fast - Resource Use Efficiency: Identification, max pipeline duration analysis, fail fast with stages grouping, fail fast with async needs, analyse blocking stages pipeline (solution with needs), matrix builds for parallel execution (pratice: combine matrix and extends, combine matrix and !reference), extends merge strategies (with and without !reference) - CI/CD Infrastructure Efficiency: Optimization ideas, custom build images, optimize builds with C++ as example, GitLab runner resource analysis (sharing, tags, external dependencies, Kubernetes), local runner exercise, resource groups, storage usage analysis, caching (Python dependency exercise, including when:always on failed jobs) - Auto-scaling: Overview, AWS auto-scaling with GitLab Runner with Terraform, insights into Spot Runners on AWS Graviton - Group discussion - Deployment Strategies: IaC, GitOps, Terraform, Kubernetes, registries - Security: Secrets in CI/CD variables, Hashicorp Vault, secrets scanning, vulnerability scanning - Observability: CI/CD Runner monitoring, SLOs, quality gates, CI/CD Tracing - More efficiency ideas: Auto DevOps, Fast vs Resources, Conclusion and tips
- nhoughto 5y agoOne improvement that would be nice in gitlab is an extension / similar feature to the FF_USE_FASTZIP flag. The default way to cache things is zip, which no matter the compression isn't a great fit, you really want a format that is stream-able like tar. When its stream-able you can download the tar and unpack it at the same time, not waiting for the full download before beginning to unpack. Obviously you would compress the tar (normally) with lz4 or zstd or similar, using this approach you should generally see a reduction in total time to get and unpack the cache. node_modules even more so since it generally has zillions of small files, so the unpack time can be quite high. Be nice if gitlab supported this approach OOTB. Edit: An even better way if your environment supports it is to use a mountable image (esp for node_modules) for your caches, this basically removes the entire unpack phase of making the cache available, instead of unpacking you just mount the image and off you go. In macOS this looks like a sparseimage via attached hdituil, or in linux a squashfs image with a writable overlay mount (you need a privileged container for this if in a container). Since the OOTB gitlab.com runners are a root-user linux VM (i think?) this approach should work quite well. SquashFS images are a great fit for node_modules especially as it moves the creation of the zillions of tiny files to cache-creation time rather than cache-use time. If you share the cache images via a hostPath mount or similar for existing images you can effectively make caches available in 0 seconds (just mount an already downloaded image and done).
- john_cogs 5y agoThanks for the feedback. I passed it along to the pipeline authoring team.
- n0w 5y agoThe later should already be possible. I believe you can specify a docker image to be used for running a job so you can bake things like node_modules into the image.
- nhoughto 5y agoYep true, generally baking caches into build images is a bit fragile tho, the more caches and different types of jobs you have the larger the set of combinations of images you need to cover it. Often ends up easier to decouple build image from caches, and ideally calculate a cache key based on a yarn.lock or similar (not supported by gitlab). Also image pull performance is generally pretty poor from most registries, you will get much higher speeds out of a straight s3 download then pulling an image via docker or containerd, even tho ecr for ex is backed by s3 . I haven’t actually used ootb gitlab caching in quite a long time because of these limitations, but would be nice for it to work great ootb!
- diftraku 5y agoOne point worth of note when it comes to Docker caching, more specifically pulling images, is the rate-limiting on Docker Hub. While hosted GitLab might make use of a transparent pull-through cache (as I've gathered from glancing at relevant parts of the docs), you can benefit a lot by using one with your own local GitLab instance (assuming it does not already provide it via container registry). We ended up switching to Harbor[1] from the vanilla registry and almost by chance stumbled on the fact that it supported a pull-through cache from various other sources (including Docker Hub). This was especially useful after we hit the rate-limit after one of our pipelines got out of hand and decided to rebuild every locally hosted Docker image (both internal and external). [1]: https://goharbor.io/ https://goharbor.io/
- jayd16 5y agoI work in games, mostly Unity but also backend services and such. I haven't had a lot of success with gitlab caches. In my experiments, it's usually faster and less error prone to simply not use the gitlab shared cache. Most caching benefits seem to come from proper git ignore usage and to forgo the network and file IO cost of a gitlab's cache system. What kinds of things should be cached? What kind of success are others seeing?
- elurg 5y agoIs there any good way to cache a built container and use it as the image for later steps? I think that caching files is much less effective than caching a built container which already has dependencies installed. For tools like apt and pip the time it takes to "install" can be longer that the time it takes to "download".
- mshekow 5y agoJust push it to some image registry (e.g. the one that comes with GitLab), and then use it as image on the next job. You can also use tags to enforce that it is the same runner who runs the two jobs, so that pulling the image becomes instant
- pid-1 5y agoGreat article, wish I had something like that 3 years ago. Adding my personal tips: - Do not use GitLab specific caching features, unless you love vendor lock in. Instead, use multi stage Docker builds. This way you can also run your pipeline locally and all your GitLab jobs will consist of "docker build ..." - Upvote https://gitlab.com/gitlab-org/gitlab-runner/-/issues/2797 https://gitlab.com/gitlab-org/gitlab-runner/-/issues/2797 . Testing GitLab pipelines should not be such a PIA.
- raffraffraff 5y agoIn a previous life, I set up CI runner images (Amazon AMIs) that had all of our docker base images pre-cached, and ran custom docker cleanup script that excluded images with certain tags. This meant that a new runner would be relatively quick off the blocks, and get faster as it built/pulled more images. You can get better cache hit from tagging your gitlab runners and pinning projects to certain tags. Also this: https://medium.com/titansoft-engineering/docker-build-cache-sharing-on-multi-hosts-with-buildkit-and-buildx-eb8f7005918e https://medium.com/titansoft-engineering/docker-build-cache-... ... Sharing cache on multiple hosts using buildkit + buildx
- pid-1 5y agoI use GitLab's private registry + scheduled pipelines to prebuild our base images, but that's definitely some extra spice. Thanks for sharing!
- dnsmichi 5y agoGreat tips, thank you! > Instead, use multi stage Docker builds. This way you can also run your pipeline locally and all your GitLab jobs will consist of "docker build ..." There's a section in the pipeline efficiency docs with more tips and tricks for optimizing Docker images: https://docs.gitlab.com/ee/ci/pipelines/pipeline_efficiency.html#docker-images https://docs.gitlab.com/ee/ci/pipelines/pipeline_efficiency....
- kzrdude 5y agoSerious question: what's the energy/resource usage tradeoff for CI? When are we burning too much resources on pointless testing and hearing data centers? I'm not saying CI is bad, but where is the threshold where it becomes wasteful, how big should the tested configuration matrix be?
- jftuga 5y agoThis is an excellent article. How does GitHub CI/CD compare to this?
- iduoad 5y agoAs Far As I know, You can do pretty much the same thing with Github CI/CD (Github Actions).
- e_proxus 5y agoHere’s an example of how to enable caching in GitHub Actions: https://gist.github.com/eproxus/d74315864fd6897bb47741e8de5bc980#cache https://gist.github.com/eproxus/d74315864fd6897bb47741e8de5b... The example is a bit Erlang specific but the cache action it contains is quite generic.
- dnsmichi 5y agoShared this thread into a new community forum topic for valuable resources: https://forum.gitlab.com/t/ci-cd-pipeline-efficiency-resources/ https://forum.gitlab.com/t/ci-cd-pipeline-efficiency-resourc...