18 ms·
Keeping master green at scale
- 7e 7y agoIs this novel? Other companies have had this for ages.
- ricardobeat 7y agoNo, they haven't. This is a system to queue commits, not a simple CI setup. This problem only comes up when you start having contention due to commit volume in a monorepo (think thousands commits/day). This is only the 3rd one I've heard about. > This paper introduces a change management system called SubmitQueue that is responsible for continuous integration of changes into the mainline at scale while always keeping the mainline green. Based on all possible outcomes of pending changes, SubmitQueue constructs, and continuously updates a speculation graph that uses a probabilistic model, powered by logistic regression. The speculation graph allows SubmitQueue to select builds that are most likely to succeed, and speculatively execute them in parallel. Our system also uses a scalable conflict analyzer that constructs a conflict graph among pending changes. The conflict graph is then used to (1) trim the speculation space to further improve the likelihood of using remaining speculations, and (2) determine independent changes that can commit in parallel
- zyang 7y agoI don't quite understand the problem they are trying to solve. Is there so many change sets that they couldn't provision enough ci servers, hence the "speculation graph with probabilistic model"?
- yzmtf2008 7y agoSay you change test A, and I change test B. Before, they both expect a value VAL to be 1, but now test A expects it to be 2 and test B expects it to be 3. We both submit a change, and both changes passed CI on that respective branch. You merged your change into master since it looks OK, and I do too. Now master is broken. Womp womp.
- joshuamorton 7y agoSort of, though not really. Imagine I have three changes, C1 modifies F1, C2 modifies F2, and C3 modifies F1. There's no relation between F1 and F2. At low-ish rate of submission, you test and commit C1, then test and commit C2, then when you try and test and commit C3, you rebase, and re-test and commit. (the merge doesn't conflict so can be automatically fixed) Now assume all three changes are submitted by 3 different engineers in the span of a minute and engineers don't want to manually rebase. The rebase/build/submit time is less than the time between changes! So you have a tool that queues up the changes, and at each change you 1. Rebase onto current head 2. build with the new changes 3. Submit But that's still really slow. Since everything is sequential. If my change takes ~30m to test, it blocks everyone else who depends on my change. So OK, do things in parallel: Build and test C1, C1 + C2, and C1+C2+C3. Then, as soon as C1 is finished testing, you can submit all 3. There's still 2 problems though: C2 is unreasonably delayed, and "what if C1 is broken". So, if C2 and C1 don't conflict, you can actually just submit C2 before C1 even though the request to submit was made after. But when there really is a dependency, like C3 and C1, the question is, do I build and test {C1, C1+C3}, {C3, C3+C1}, or something else. SubmitQueue appears to try and address that question. "Given potentially conflicting changes (not at a source level but at a transitive closure level), how do I order them so that the most changes succeed the fastest, assuming some changes can fail, and I have enough processing power to run some, but not all, permutations of changes in parallel"?
- jrochkind1 7y agoawesome explanation
- roshanj 7y agoThe problem also occurs if your CI Build + Test steps take a while to run, even on a small team pushing dozens of commits per day. Two code-conflict-free changes may pass a pre-merge build+test cycle independently but may logically break one another if both changes are merged into master. Using a submit/merge queue guarantees that each change has passed tests with the exact ordering of commits it would be merged onto. The example described here is a better explanation: https://github.com/bors-ng/bors-ng#but-dont-githubs-protected-branches-already-do-this https://github.com/bors-ng/bors-ng#but-dont-githubs-protecte...
- UweSchmidt 7y agoIn companies like that, is there any consideration given to minimizing conflict-prone actions, like, say renaming functions, an activity that could conflict with any commit that uses the old function name, but which in itself is unlikely to break anything? Maybe certain commits could be scheduled over the weekend? I guess I just have a hard time imagining how many buys developers really commit important work all at once on large projects...
- foobiekr 7y agoI think they are more common than you are thinking. I am familiar with several, even going back to the svn and cvs era, all of which predated the whole formalized-and-named CI/CD thing. In my experience, we called this model submit-to-commit, and depending on the specific manifestation, worked with diffs or branches. I'm talking 1990s. The fancy bits in this implementation from the paper are interesting but the model itself is not that unusual.
- paxys 7y agoWhich companies?
- bostik 7y agoUs, for instance.[0,2] But sure enough, we definitely weren't the first to go down this path. Facebook was using (or developing the tech for) server-side rebasing in 2015.[1] Gitlab provides native server-side rebase functionality, likely inspired by various parties already having developed tools to do the same. These aren't new ideas. But handling them at the scale where you land hundreds or even thousands of commits a day to a repo and require the ability to deploy at will, that's where engineering comes into play. 0: https://smarketshq.com/marge-bot-for-gitlab-keeps-master-always-green-6070e9d248df https://smarketshq.com/marge-bot-for-gitlab-keeps-master-alw... 1: https://softwareengineering.stackexchange.com/questions/278713/why-server-side-repository-merge-is-a-terrible-idea-in-git https://softwareengineering.stackexchange.com/questions/2787... 2: https://github.com/smarkets/marge-bot https://github.com/smarkets/marge-bot
- lozenge 7y agoYours seems identical to Bors, just for Gitlab instead of GitHub? That isn't really what's described in the OP.
- bostik 7y agoYup, pretty much. I was mostly answering the parent, who in turn was questioning the lack of novelty. The concept of an evergreen master with testing done in branches, followed by automated merges/rebases is not special. Quite a few companies have been doing it for years, it's the off-the-shelf tooling and subsequent publicity that haven't necessarily been around as long. As for OP's material? The automated conflict resolution via reordering to optimise parallelism - that certainly feels novel.
- ungzd 7y agoOther companies often just wait for tests to finish, while at the time of running tests proposed changes (branch/PR) might be not based on current version of master. Then they just rebase/merge after tests pass, without running tests again. For smaller projects, this rarely breaks. For monorepo with lots of committers rate of breakage becomes too large. Next step is to serialize all proposed changes, so they are rebased one on top of other before running tests. This eliminates breakage due to merging, but does not scale: > The simplest solution to keep the mainline green is to enqueue every change that gets submitted to the system. A change at the head of the queue gets committed into the mainline if its build steps succeed. > > This approach does not scale as the number of changes grows. For instance, with a thousand changes per day, where each change takes 30 minutes to pass all build steps, the turnaround time of the last enqueued change will be over 20 days. This paper is about scaling a variant of such queue.
- chairleader 7y agoQuite a premise: "Giant monolithic source-code repositories are one of the fundamental pillars of the back end infrastructure in large and fast-paced software companies."
- draw_down 7y agoI'm willing to agree with that premise.
- googlemike 7y agoI like it at Google!
- Aqua_Geek 7y agoGoogle has the team + tooling to properly support it. The same cannot be said for many other orgs.
- nevir 7y agoMany teams
- glandium 7y agoGoogle has more people working on the problem than many other companies have employees.
- Aqua_Geek 7y agoI don’t doubt it. They also do more traffic through their VCS than most companies do through their main product.
- erik_seaberg 7y agoDo they? They didn't used to. In 2015 we were routinely dead in the water, unable to test and deploy anything from our google3 projects because some random submitted a CL for a project we didn't even care about. Teams would appoint "build cops" whose job is to complain as quickly as possible because that's all we could do about it. Every problem you could have with bad dependencies is entirely self-inflicted. The Right Thing™ is to choose a known-good version, and update when you have the bandwidth to pay down the tech debt.
- huac 7y ago"Based on all possible outcomes of pending changes, SubmitQueue constructs, and continuously updates a speculation graph that uses a probabilistic model, powered by logistic regression. The speculation graph allows SubmitQueue to select builds that are most likely to succeed, and speculatively execute them in parallel" This is either brilliant or just something built for a promotion packet
- mcqueenjordan 7y agoPromotion-oriented design, no doubt.
- burakcosk 7y agono it is not, this whole thing being used in production and really reduces the time changes submitted and they are merged into master. and the master is almost always green, meaning developers can build and test any piece of code without problem. They have designed this as a result of a need, not just a fancy project.
- draw_down 7y ago"I don't experience this problem, therefore this problem is not real."
- sundargates 7y agoI can guarantee you that none of the ideas on the paper were born out of a desire to get promoted. They were invented because ML models helped figure out which set of builds we need to run more accurately at scale.
- xiphias2 7y agoI still hope you got promoted though, you deserve it :)
- pastor_elm 7y agoSounds so much simpler outside the context of a 'research' paper: >When an engineer attempts to land their commit, it gets enqueued on the Submit Queue. This system takes one commit at a time, rebases it against master, builds the code and runs the unit tests. If nothing breaks, it then gets merged into master. With Submit Queue in place, our master success rate jumped to 99%. https://eng.uber.com/ios-monorepo/ https://eng.uber.com/ios-monorepo/
- revskill 7y agoWhat's exactly a monothlic ? Is it only related to codebase (monothlic vs monorepo) ? Or it's about runtime like microservices vs monothlic.
- jade12 7y agoFrom the first sentence of the abstract: > monolithic source-code repositories A monorepo is a monolithic repository
- ricardobeat 7y agoTo answer the parent, it doesn’t imply a monolith application, but deployment to multiple server roles and apps will happen using the same source repository.
- underrun 7y agoAdrian Colyer dug into this a little further on the morning paper: https://blog.acolyer.org/2019/04/18/keeping-master-green-at-scale/ https://blog.acolyer.org/2019/04/18/keeping-master-green-at-... His analysis indicates that what uber does as part of its build pipeline is to break up the monorepo into "targets" and for each target create something like a merkle tree (which is basically what git uses to represent commits) and use that information to detect potential conflicts (for multiple commits that would change the same target). what it sounds like to me is that they end up simulating multirepo to enable tests to run on a batch of most likely independent commits in their build system. For multirepo users this is explicit in that this comes for free :-) which is super interesting to me as it seems to indicate that an optimizing CI/CD systems requires dealing with all the same issues whether it's mono- or multi- repo, and problems solved by your layout result in a different set of problems that need to be resolved in your build system.
- ori_b 7y ago> For multirepo users this is explicit in that this comes for free :-) Only if you spend the time to build tools to detect commits in your dependencies, as well as your dependent repositories, and figure out how to update and check them out on the appropriate builds. So, no, it doesn't come for free.
- msangi 7y agoPackage managers solve it quite well. Just depend on the latest version of your dependencies and tag a new version whenever they change.
- ricardobeat 7y agoThis doesn’t work when an underlying system changes, and upgrading is mandatory for all clients or package dependants (happens often at scale for a multitude of reasons).
- ljm 7y ago
- richardwhiuk 7y agoAnyone fancy comparing this to bors?
- drodgers 7y agoThe main difference is in the conflict-detection system. Whereas bors only has a single queue, this new system can have one queue for each set of changes which doesn't interact with any other set. Eg. if you've got an ios app, a webapp, and a bunch of documentation all in the same repo, then this system will automatically work out that changes to each of those independent projects can be tested and merged in parallel, because they can't possibly conflict. It relies on understanding the inputs and outputs for all CI build steps to work out how changes to particular files might conflict. Also, it has a much more sophisticated understanding of how likely a change is to be the source of failure, which it updates in response to repeated test runs. It can then prioritise the changes which are most likely to succeed.
- richardwhiuk 7y agoIs the logic of which queue what files trigger automatically or manually determined?
- sundargates 7y agoActually we have compared it in our paper. Bors builds one change at a time. On the other hand, Submit Queue speculatively builds several changes at a time based on the outcomes of other pending changes in the system. Apart from that, Submit Queue uses a conflict analyzer to find independent changes in order to commit changes in parallel as well as trim the speculation graph. We have also evaluated the performance of Single-Queue (idea of Bors) on our workloads. In fact, as described in the paper, the performance of this technique at scale was so high (~132x slower) that we omitted its results. Submit Queue on the other hand operates at 1-3x region compared to an optimal solution. I recommend you to read the paper here for further details. https://dl.acm.org/citation.cfm?id=3303970 https://dl.acm.org/citation.cfm?id=3303970
- 7y ago
- ratRaces 7y agoThis isn't really news. Managing a technical code base is mostly about excluding total fucking dipshits. Excluding total fucking dipshits is often about hiring smart people, and not exterminating their enthusiasm and optimism, so that their willingness to perform doesn't wither and die in a smothering rat race environment.
- techmortal 7y agoHow common is this in the industry? Do multirepos run on a batch?
- jonthepirate 7y agoHaving been at both Lyft and DoorDash where I've been an engineer responsible for unit test health, I decided to do a side project called Flaptastic (https://www.flaptastic.com/ https://www.flaptastic.com/), a flaky unit test resolution system. Flaptastic will make your CI/CD pipelines reliable by identifying which tests fail due to flaps (aka flakes) and then give you a "Disable" button to instantly skip any test which is immediately effective across all feature branches, pull requests, and deploy pipelines. An on-premise version is in the works to allow you to run it onsite for the enterprise.
- roskilli 7y agoI don't want to come across as negative, but just an observation and to play devil's advocate - wouldn't it be better to fix the flaky test or delete it entirely instead of build a feature to disable it during a test run in an automated fashion? Whenever our team has a significant number of flakey tests (more than 1-2) we usually schedule a bug squash session to fix them and amortize the cost over the whole team.
- viklove 7y agoBest practice is actually just to disable all tests that are failing. Can't hold up our sprint deadlines!
- tomschlick 7y agoFailing != flaking. If your tests interact with any level of randomness (seed data, time based constraints, etc) you're going to find the occasional test that doesn't work and subsequently works on the rebuild. If something is consistently failing I would assume this tool does not disable it.
- jonthepirate 7y agoWhat you really want to do is first disable a test you know is unhealthy to unblock everybody. Then, you fix it. After you've reintroduced it healthy, you can turn it back on.
- cjfd 7y agoA possible complication would occur if there are tests that occasionally fail.
- Scaevolus 7y agoThere's a nice middle ground between this and a one-at-a-time submit queue: have a speculative batch running on the side. This gives nice speedups (approaching N times more commits, where N is the batch size) with minimal complexity. One useful metric is the ratio between test time and the number of commits per day. If your tests run in a minute, you can test submissions one at a time and still have a thousand successful commits each day. If your tests take an hour, you can have at most 24 changes per day under a one-at-a-time scheme. I worked on Kubernetes, where test runs can take more than an hour-- spinning up VMs to test things is expensive! The submit queue tests both the top of the queue and a batch of a few (up to 5) changes that can be merged without a git merge conflict. If either one passes, the changes are merged. Batch tests aren't cancelled if the top of the queue passes, so sometimes you'll merge both the top of the queue AND the batch, since they're compatible. Here's some recent batches: https://prow.k8s.io/?repo=kubernetes%2Fkubernetes&type=batch https://prow.k8s.io/?repo=kubernetes%2Fkubernetes&type=batch And the code to pick batches: https://github.com/kubernetes/test-infra/blob/0d66b18ea7e8d3f216287ad06b11042c12bc6e48/prow/tide/tide.go#L759 https://github.com/kubernetes/test-infra/blob/0d66b18ea7e8d3... Merges to the main repo peak at about 45 per day, largely depending on the volume of changes. The important thing is that the queue size remains small: http://velodrome.k8s.io/dashboard/db/monitoring?orgId=1&panelId=10&fullscreen&from=now-7d&to=now http://velodrome.k8s.io/dashboard/db/monitoring?orgId=1&pane...
- viraptor 7y agoThe same thing was (is?) done in openstack with zuul, I believe. When you going to merge something, your branch goes on top of things already going through the CI.
- Scaevolus 7y agoWe talked to the Zuul team, they use more parallelism but it's similar: https://zuul-ci.org/docs/zuul/user/gating.html https://zuul-ci.org/docs/zuul/user/gating.html Most of the complexity and suffering of a submit queue evolves from the interactions between your VCS and CI systems. Keeping things simple is great! Kubernetes' CI system is Prow, which runs the tests as pods in a Kubernetes cluster. Dogfooding like this is great, since the team you're providing CI for can also help fix bugs that arise.
- antimora 7y agoI am still trying to wrap my head around a giant monolithic repo model instead of breaking codes into multiple repos. At Amazon, for example, they have multi repos setup. A single repo represents one package which has major version.The Amazon's build system builds packages and pulls dependencies from the artifact repository when needed. The build system is responsible for "what" to build vs "how" to build, which is left to the package setup (e.g. maven/ant). I am currently trying to find a similar setup. I have looked as nix, bazel, buck and pants. Nix seems to offer something close. I am still trying to figure how to vendor npm packages and which artifact store is appropriate. And also if it is possible to have the nix builder to pull artifacts from a remote store. Any pointer from the HN community is appreciated. Here is what I would like to achieve: 1. Vendor all dependencies (npm packages, pip packages, etc) with ease. 2. Be able to pull artifact from a remote store (e.g. artifactory). 3. Be able to override package locally for my build purposes. For example, if I am working on a package A which depends on B, I should be able to build A from source and if needed to build B which A can later use for its own build. 4. Support multiple languages (TypeScript, JavaScript, Java, C, rust, and go). 5. Have each package own repository.
- dilyevsky 7y agoMonorepos are really nice if you want to enforce consistent and sane engineering practices and not waste time managing all the repos individually by teams. Bazel has target caching including remote caching which can be shared across multiple engineers/execution environments. The tricky part would be ensuring your builds are hermetic and reproducible (which is also easier to achieve in monorepo setup).
- PKop 7y ago> At Amazon, for example, they have multi repos setup. And didn't you find that this created massive headaches trying to build many disparate and inconsistent dependencies across repos? I think the benefits touted from mono-repos are exactly illustrated by the pain points working with Amazon's multi repo setup, in my opinion. https://danluu.com/monorepo/ https://danluu.com/monorepo/ "Refactoring an API that's used across tens of active internal projects will probably a good chunk of a day." This was my experience.
- jl-gitlab 7y agoWe're building some similar tech at GitLab, though without the dependency analysis yet. Merge Requests now combine the source and target branches before building, as an optimization: https://docs.gitlab.com/ee/ci/merge_request_pipelines/#combined-ref-pipelines-premium https://docs.gitlab.com/ee/ci/merge_request_pipelines/#combi... Next step is to add queueing (https://gitlab.com/gitlab-org/gitlab-ee/issues/9186 https://gitlab.com/gitlab-org/gitlab-ee/issues/9186), then we're going to optimistically (and in parallel) run the subsequent pipelines in the queue: https://gitlab.com/gitlab-org/gitlab-ee/issues/11222 https://gitlab.com/gitlab-org/gitlab-ee/issues/11222. At this point it may make sense to look at dependency analysis and more intelligent ordering, though we're seeing nice improvements based on tests so far, and there's something to be said for simplicity if it works.
- shimont 7y agoI think that what works for companies like Uber/Google/Facebook is not applicable to the rest of fortune 500 or all of the rest of the companies. disclaimer: I am one of Datree.io founders. We provide a visibility and governance solution to R&D organizations on top of GitHub. Here are some rules and enforcement around Security and Compliance which most of our companies use for multi-repo GitHub orgs. 1. Prevent users from adding outside collaborators to GitHub repos. 2. Enforce branch protection on all current repos and future created ones - prevent master branch deletion and force push. 3. Enforce pull request flow on default branch for all repos (including future created) - prevent direct commits to master without pull-request and checks. 4. Enforce Jira ticket integration - mention ticket number in pull request name / commit message. 5. Enforce proper Git user configuration. 6. Detect and prevent merging of secrets.