9 ms·
We cut our CI pipeline execution time in half
- imiric 4y ago> We noticed a strong correlation between crazy utilization spikes and CI failure rates. This is interesting, and is something I've also suspected on many CI systems that offer free public runners (CircleCI, GitHub Actions, etc.). For seemingly no reason at all, tests were very flaky and unstable in CI, which couldn't be reproduced on local machines. I tried everything from resource-limited containers, to identically spec'd VMs, and never was able to reproduce certain failures. This made issues very hard to troubleshoot and fix. Of course, you might say that this unstable environment surfaced race conditions in our tests or product, and that's true, but it's incredibly frustrating to have random failures that are impossible to reproduce locally, and having to wait for the long experiment-push-wait for CI development loop. I suspect this is caused by over provisioning of the underlying hardware, where many VMs are competing for the same resources. This seems quite frequent on Azure (GH Actions). In the article's case they patched it by making their environment more stable, which is a solution we can't do on public runners, but I'd caution them that they're only patching the issue, and not really fixing the root cause. The flakiness still exists in their code, and is just not visible when the system is not under stress, but will surface again when you least want it to, possibly in production.
- alrocar 4y agoYep, default runners in most CI platforms shared resources so they are prone to produce flakiness (depending on your set up). That was one of the reasons we ended up setting up our own runners. Didn't mention in the post but we use spot VM instances.
- CottonMcKnight 4y agoTL;DR: how a data company uses their own product.
- dijit 4y agoI have a somewhat related question. I'm using gitlab-ci with it's docker executor, and overall I'm very happy with it. I use it on some rather beefy machines, but most of the CI time is not spent compiling, it is spent instead on setting up the environment. Are there any tips/tricks to speed up this startup time? I know stuff like ensuring that artifacts are not passed in if not needed can help a lot, but it seems that most of the execution time is simply spent waiting for docker to spin up a container.
- kyriakos 4y agoI noticed that many times using cache makes gitlab ci take longer than just fetching node dependencies again via npm install.
- buttersbrian 4y agoHow do you setup or provision your environment? And what does this environment look like?
- dijit 4y agoDocker on Debian 11 bare metal with gitlab-ci installed the "blessed" way (by adding gitlabs apt repos). No optimisation to the baseOS other than mounting the /var/lib/docker on a RAID0 array with noatime on the volume and CPU mitigations disabled on the host Compilation is mostly go binaries (with the normal stuff like go vet/go test). Rarely it will do other things like commit-lint (javascript) or KICS/SNYK scanning. the machines themselves are Dual EPYC 7313 w/ 256G DDR4.
- hdjjhhvvhga 4y agoWhere do you keep your bare metal machines if I my ask? I wanted to do a similar setup a while ago (building/testing on Hetzner bare metal, deployments and the rest on AWS) but due to Amazon's pricing policy the cost of traffic would be enormous.
- 4y ago
- berkle4455 4y agoCI has been such a productivity killer. You don’t need it. Stick with CD only and you can ship.
- alrocar 4y agoI'm really interested in different points of view. I guess you mean the kind of trunk based development? But still some sort of CI happens, maybe locally. Never worked in a different way than using a local / remote CI pipeline, that's why I'm curious.
- jdkoeck 4y agoCI is a prerequisite for CD.
- berkle4455 4y agoIt’s literally not.
- nerdponx 4y agoIt might be required in that it's impossible to deliver continuously if changes are not integrated continuously, for some definition of "integrated".
- teach 4y agoI'm tempted to just downvote you and move on with my life but I'm genuinely curious. Given that it's meaningless to Deploy something without Integrating the changes, what do you _actually_ mean by "You don’t need [CI]. Stick with CD only." Are you just talking about testing the changes? Help us out here.
- xboxnolifes 4y agoI guess you could just have an unchanging project redeploy itself every hour or so.
- 4y ago
- hankstenberg 4y ago[flagged]
- teach 4y agoAm I a curmudgeon? Not to take away from this cool writeup, but I'm familiar with a few CI/CD tools, particularly QuickBuild, Jenkins and Spinnaker. So this jumped out at me: > Our CI process was pretty standard: Every commit in an MR triggered a GitLab Pipeline, which consisted of several jobs. me: nodding silently > Those jobs would run in an auto-scaling Kubernetes cluster with up to 21 nodes me: what the actual deuce? Is this really "pretty standard"?
- alrocar 4y agoThanks for this comment. I guess there's sometimes we (developers) take things for granted when they are not, and that puts a lot of pressure on us instead of celebrating our wins. I would change now "pretty standard" by "we don't invented the wheel" xD :pray:, in the end I wanted to mean we use existing tools and "just" put them together
- brightball 4y agoIf you have it, it’s awesome. You can get parallel execution of so much, spin up environments for each branch for QA and dynamic scans. IMO it’s the optimal use case for K8s
- hedora 4y agoYou have to be at a certain scale for k8s to make sense in a CI environment. In particular, it needs to be economical to spend 10-50% of a full time employee to maintain the Kubernetes cluster (even if it is some managed thing like EKS). Also, the duty cycle on the 21 nodes needs to be low enough to justify the complexity over just buying 21 computers (or getting annual pricing on 21 VMs). You could use spot instances for the EKS nodes, but then PRs will randomly fail because their instances disappear. That wastes developer salary money and productivity. Assuming you have a ventilated room you don't care about, you could run 21 desktop towers off of ~ two-four 120V circuits. (Or buy a rack and pay ~ 2x as much for the hardware.) 21 build hosts would cost ~$21-42K. Power is probably averaging 50W per machine (they are probably mostly idle even when running tests, since they have to download stuff.) That's about 720KWh per month. US average electrical pricing is $0.20 / kWh; punitive California rates are about $0.40. So, in the punitive case, that's $288 / month. Running 21 machines probably requires as much annoying maintenance work as EKS, though the maintenance includes swapping bad hardware, fiddling with ethernet cables, and wearing ear protection (if a rack is involved) instead of debugging piles of yaml and AWS roles, optimizing to stay in budget, etc, etc.
- Scubabear68 4y agoMaybe I am just an old fuddy duddy conservative, but this struck me from the post: “In the grand scheme of things, one week isn’t that long. But to us, it felt like forever. We are constantly iterating and release multiple changes every day”. I assume they mean multiple production releases? Is this because the product lacks maturity or stability, or is it just your culture? I am asking because I am trying to imagine the impact of this on existing customers. It sounds like an awful lot of churn. This obviously happens a lot in the “you are the product” space like Facebook, Google, etc. But this looks to be a data analytics product with paid tiers. Curious what tooling and processes you have to support this, and how you keep customers happy with this model.
- mariosisters 4y agoI think it’s that you are an old fuddy duddy :P Actually, if you work with SMBs/enterprises, I agree with you on customer facing changes. In my past life we would ship very frequently (often more than once a day) but always had to feature flag changes that large clients might see or be affected by. Even something as simple as tweaking the layout of a core flow could cause support headaches and angry customers — customers worth 10s of thousands of dollars per month. Is it worth losing a customer to CD a new button placement?
- chrisandchris 4y agoI can only image how clean code looks & works that is full of feature flags. Glad that I don't need to do that to often :)
- tempest_ 4y agoThey keep adding them until they need an internal library to manage their collection of feature flags. Then after a while they graduate to feature flags as a service (of which there are a bunch of cloud services trying to have a go at)
- mariosisters 4y agoThe right approach is to immediately remove the flags after rollout… The actual approach is to maintain a million fucking feature flags, ensuring that almost all possible combinations are essentially untested… better hope you did a good job separating concerns!