28 ms·
Google stores billions of lines of code in a single repository (2016) [pdf]
- myhf 4y ago(published July 2016)
- thunderbong 4y agoAlso, it's a PDF link
- dang 4y agoAdded. Thanks!
- chrisa 4y agoHere's a talk version given by Rachel (one of the authors) about the same topic: https://www.youtube.com/watch?v=W71BTkUbdqE https://www.youtube.com/watch?v=W71BTkUbdqE
- sabujp 4y agoand my previous director is scaling github now :)
- GreedClarifies 4y agoThis is from the golden age of Google. Of particular note is that they published this many years after it had been shipped to their internal customers. This was not some position paper about "why we focus on ai" after not shipping any of their "breakthroughs".
- rvcdbn 4y agoI really wish they would make this tech available via gcloud. Seems like it would be very popular and a great way to attract other gcloud business away from MS/GitHub which scales horribly.
- grahar64 4y agoThey tried that by making a bit available with a remote cloud builder for Bazel. It failed for some reason and they pulled it. I think building something that scales for one big repo is just a completely different problem than making it scale for a lot of small repos.
- seedless-sensat 4y agoBazel is not failing in the open source world though
- ameliaquining 4y agoI think that maintaining a hosted service has significantly higher fixed costs than maintaining an open source project whose users are responsible for deploying it themselves. So a higher degree of adoption would be necessary to justify it.
- grahar64 4y agoDont get me wrong, I love Bazel. But tooling for it to build multi-language multi-billion line monorepos doesn’t exist outside big companies.
- deleted 4y ago[deleted]
- blindriver 4y agoLong term projects like this don't get any attention because the chance of getting a promotion from it are almost nil. And after the layoffs, it's pretty clear that no matter how hard you work, you can get fired so what's the point in dedicating your career to something like this?
- deanCommie 4y agoNo wonder noone at Google can't ship everything if they constantly have to stop development of their feature so they can do mandatory upgrades of their dependencies...
- ameliaquining 4y agoMost of that work is done by the owners of the dependencies, rather than the dependents. This is sometimes a problem for open source dependencies, though, as there isn't always anyone whose job it is to keep them up to date. Some amount of NIH syndrome is because reinventing the wheel can be less work than integrating an existing wheel that was designed for a different vehicle with different specs.
- zdw 4y agoMonorepos are great... but only if you can invest in the tooling scale to handle them, and most companies can't invest in that like Google can. Hyrum Wright class tooling experts don't grow on trees. A good article to reference when this topic gets raised: http://yosefk.com/blog/dont-ask-if-a-monorepo-is-good-for-you-ask-if-youre-good-enough-for-a-monorepo.html http://yosefk.com/blog/dont-ask-if-a-monorepo-is-good-for-yo...
- no_wizard 4y agoYou can get better tools now though, like Turbo Repo or NX. They don’t require the same level of investment as Bazel but they don’t always have the same hermetic build guarantees, though for most it’s “good enough”.
- lallysingh 4y agoBuild in docker.
- patrick451 4y agoYou don't need google scale tooling to work with a mono repo until you are actually at google scale. Gluing together a bunch of separate repos isn't exactly free either. See, for example, the complicated disaster Amazon has with brazil. In the limit, there are only two options: 1. All code lives one repo 2. Every function/class/entity lives in its own repo with a third state in between 3. You accept code duplication This compromise state where some code duplication is (maybe implicitly) acceptable is what most people have in mind with a poly-repo. The problem though is that (3) is not a stable equilibrium. Most engineers have such a kneejerk reaction against code duplication that (3) is practically untenable. Even if your engineers are more reasonable, (3) style compromise means they constantly have to decide "should this code from package A be duplicated in package B, or split off into a new smaller package C, which A and B depend on". People will never agree on the right answer, which generates discussion and wastes engineering time. In my experience, the trend is almost never to combine repos, but always to generate more and more repos. The limiting case of a mono repo (which is basically it's natural state) is far more palatable than the limiting case of poly-repo.
- dang 4y agoRelated: Why Google Stores Billions of Lines of Code in a Single Repository (2016) - https://news.ycombinator.com/item?id=22019827 https://news.ycombinator.com/item?id=22019827 - Jan 2020 (121 comments) Why Google Stores Billions of Lines of Code in a Single Repository (2016) - https://news.ycombinator.com/item?id=17605371 https://news.ycombinator.com/item?id=17605371 - July 2018 (281 comments) Why Google stores billions of lines of code in a single repository (2016) - https://news.ycombinator.com/item?id=15889148 https://news.ycombinator.com/item?id=15889148 - Dec 2017 (298 comments) Why Google Stores Billions of Lines of Code in a Single Repository - https://news.ycombinator.com/item?id=11991479 https://news.ycombinator.com/item?id=11991479 - June 2016 (218 comments)
- deleted 4y ago[deleted]
- randyrand 4y agoiOS and Windows are “monorepos” too. The software is built daily, and everyone must be on the same version of every library. Under the hood there are a bunch of repos, and there are exceptions, but largely operates as a monorepo.
- jbm 4y agoIs this still the case for Windows? I remember hearing something like this when I was getting my BCompSci, but I assumed it must have changed since then.
- KolmogorovComp 4y ago> Google’s codebase is shared by more [...] than 25,000 Google software develop- ers from dozens of offices in countries around the world. > Access to the whole codebase encourages extensive code sharing and reuse [...] Doesn't this strategy result in a great risk of massive code leaks from rogue employees? Even if read access are logged and the culprit found, it's too late once it's been published.
- ameliaquining 4y agoMost source code just isn't that interesting or sensitive.
- scarface74 4y agoIf you had every line of code that Google wrote, what would you do with it? But I found this discussion on HN. https://news.ycombinator.com/item?id=11790438 https://news.ycombinator.com/item?id=11790438
- ameliaquining 4y agoWell, if you had the search ranking algorithms or the bot-detection algorithms or anything inherently adversarial like that, then you could do all kinds of nefarious things. But that stuff's locked down more tightly. Likewise with a few ultra-hard-tech things where the implementation's a major competitive edge.
- forgotusername6 4y agoI imagine looking for vulnerable areas of the code might be something people would be interested in doing. Maybe start with login or billing or something. You could also look at recent activity to spot new, unannounced projects. You could use blame to find who wrote what and target them for anything from job offers to social engineering attacks.
- ameliaquining 4y agoMost of that information is readily available on the corporate intranet without having to dig through source code. Security-by-obscurity isn't something to rely on (again, except in the case of things like abuse detection where there's no alternative).
- yazaddaruvala 4y agoHaving worked at Google and Amazon. Honestly their systems are almost identical. Amazon just creates a monotonically increasing watermark outside the “repo”. Google uses “the repo” to create the monotonically increasing watermark. Otherwise, Google calls it “merge into g3” Amazon calls it “merge into live”. Amazon has the extra vocabulary of VersionSets/Packages/Build files. Google has all the same concepts, but just calls them Dependencies/Folders/Build files. Amazon’s workflows are “git-like”, Google is migrating to “git-like” workflows (but has a lot of unnecessary vocabulary around getting there - Piper/Fig/Workspace/etc). I really can’t tell if the specific difference between “mono-repo” or “multi-repo” makes much practical difference to the devs working on either system.
- faizshah 4y agoI haven’t worked at google but I think there is one other difference. At amazon teams “merge from live” and have control of their own service’s CD pipeline. They might manually release the merged changes or have full integ test coverage. The Amazon workflow offers more flexibility to teams (whether or not that might be desirable). Not sure how deployments and CD work at google but I think the picture is different at google for unit tests, integ tests etc. Amazon teams have more control over their own codebase and development practices whereas, based on what I know, google has standardized many parts of their development process.
- dmoy 4y agoDid you work on a team at Google that uses branches? Most teams do not, so there is no "merge into g3".
- yazaddaruvala 4y agoEvery single Fig/Piper workspace is a “branch” in a git-like workflow. It’s then “merged into g3” from that workspace.
- safog 4y agoThere are no presubmits that prevent breaking changes from "going into live". If some shared infra updates are released, the merge from live breaks for multiple individual teams rather than preventing the code from getting submitted in the first place.
- sn_master 4y agoBecause Google does something, doesn't mean it's a good thing to do for anyone else. This kind of infrastructure is very expensive to maintain, and suffers from many flaws like -almost- everyone being stuck using SDKs that are several versions behind the latest production one even for the internal GCP ones.
- lopkeny12ko 4y agoThere's a lot of love for monorepos nowadays, but after more than a decade of writing software, I still strongly believe it is an antipattern. 1. The single version dependencies are asinine. We are migrating to a monorepo at work, and someone bumped the version of an open source JS package that introduced a regression. The next deploy took our service down. Monorepos mean loss of isolation of dependencies between services, which is absolutely necessary for the stability of mission-critical business services. 2. It encourages poor API contracts because it lets anyone import any code in any service arbitrarily. Shared functionality should be exposed as a standalone library with a clear, well-defined interface boundary. There are entire packaging ecosystems like npmjs and pypi for exactly this purpose. 3. It encourages a ton of code churn with very low signal. I see at least one PR every week to code owned by my team that changes some trivial configuration, library call, or build directive, simply because some shared config or code changed in another part of the repo and now the entire repo needs to be migrated in lockstep for things to compile. I've read this paper, as well as watched the talk on this topic, and am absolutely stunned that these problems are not magnified by 100x at Google scale. Perhaps it's simply organizational inertia that prevents them from trying a more reasonable solution.
- zhengyi13 4y ago> It encourages poor API contracts because it lets anyone import any code in any service arbitrarily. Perhaps that might be the default case, but the build system has a visibility system[1] that means that you can carefully control who depends on what parts of your code. Separately, while some might build against your code directly, a lot of code just gets built into services, and then folk write their code against your published API, i.e. your protobuf specification. [1]: https://bazel.build/concepts/visibility https://bazel.build/concepts/visibility
- klodolph 4y ago> The next deploy took our service down. How would multi-repo change this? A dependency updated, and code broke, and the new version was broken—but you update dependencies in multi-repo anyway, and deployments can be broken anyway. I don’t see how multi-repo mitigates this. > It encourages poor API contracts because it lets anyone import any code in any service arbitrarily. This has nothing at all to do with monorepos. Google’s own software is built with a tool called Bazel, and Meta has something similar called Buck. These tools let you build the same kind of fine-grained boundaries that you would expect from packaged libraries. In fact, I’d say that the boundaries and API contracts are better when you use tools like Bazel or Buck—instead of just being stuck with something like a private/public distinction, you basically have the freedom to define ACLs on your packages. This is often way too much power for common use cases but it is nice to have it around when you need it, and it’s very easy to work with. A common way to use this—suppose you have a service. The service code is private, you can’t depend on it. The client library is public, you can import it. The client library may have some internal code which has an ACL so it can only be imported from the client library front-end. Here’s how we updated services—first add new functionality to the service. Then make the corresponding changes to the client. Finally, push any changes downstream. The service may have to work with multiple versions of the client library at any time, so you have to test with old client libraries. But we also have a “build horizon”—binaries older than some threshold, like 90 days or 180 days or something, are not permitted in production. Because of the build horizon, we know that we only have to support versions of the client library made within the last 90 or 180 days or whatever. This is for services with “thick clients”—you could cut out the client library and just make RPCs directly, if that was appropriate for your service. > It encourages a ton of code churn with very low signal. The places I worked at that had monorepos, you might filter out the automated code changes there to do automated migrations to new APIs. One PR per week sounds pretty manageable, when spread across a team. Then again, I’ve also worked at places where I had a high meeting load, and barely enough time to get my work done, so maybe one PR per week is burdensome if your are scheduled to death in meetings.
- bandika 4y ago[flagged]
- Scubabear68 4y agoI’d really love to know what the breakdown of those 2 billion lines of code is by product. What a huge number.
- marcrosoft 4y agoI love monorepos. I feel like they are even more helpful for small teams and smaller scale. The productivity of being able to add libraries by creating a new folder or refactor across services is unbeatable.
- gardenhedge 4y agoI've never experienced a monorepo like Googles. How does it work? Are Chrome and Gmail in the same repo? I assume they're built separately and pushing code to one doesn't affect the other.
- charcircuit 4y agoNo, Chrome and Gmail are in different monorepos. >How does it work? Different projects are in different folders instead of different repos. >I assume they're built separately and pushing code to one doesn't affect the other. Yes, building or testing something only builds its dependencies.
- tfsh 4y agoFor GP, note Chrome is a special case because it's an open source-first project so it is not in the same repo as Gmail. However products are in the same repo such a gmail, youtube, search (frontend, mobile, server, infra, etc), photos, maps, play, translate and literally thousands of other internal and external products and projects.
- Karellen 4y ago> The Google codebase includes approximately one billion files and has a history of approximately 35 million commits spanning Google’s entire 18-year existence. Wait, that's an average of nearly 30 new files per commit. Not 30 files changed per commit, but whatever changes are happening to existing files, plus 30 brand new files. For every single commit. Although... > The total number of files also includes source files copied into release branches, files that are deleted at the latest revision, [...] I'm not quite sure what this is saying. Is it saying that if `main` contains 1,000 files, and then someone creates a branch called `release`, then the repo now contains 2,000 files? And if someone then deletes 500 files from `main` in the next commit, the repo still contains 2,000 files, not 1,500? If that's the case, why not just call every different version of every file in the repo a different file? If I have a new repo and in the first commit I create a single 100-line file called `foo.c`, and then I change one line of `foo.c` for the second commit, do I now have a repo with two files? I mean, if you look at the plumbing for e.g. `git`, yes, the repo is storing two file objects for the repo history. But I don't think I've ever seen someone discuss the Linux git repo and talk about the total number of file objects in the repo object store. And when the linked paper itself mentions Linux, it says "The Linux kernel is a prominent example of a large open source software repository containing approximately 15 million lines of code in 40,000 files" - and in that case it's definitely not talking about the total number of file objects in the store. I don't think it's entirely clear what the paper even means when it talk about "a file" in a source code repository, or if it even means the same thing consistently. I'm not sure it's using the most obvious interpretation, but I can't understand why it would pick a non-obvious interpretation. Especially if it's not going to explain what it means, let alone explain why it chose one meaning over another.
- bananapub 4y agoyou're misunderstanding a bunch of things. > The total number of files also includes source files copied into release branches I guess you haven't used Perforce or similar. a branch is a sparse copy of just the changed files/directories. they are not used very much. > files that are deleted at the latest revision so it means "one billion files have existed in the history repo, some are currently deleted". > I don't think it's entirely clear what the paper even means when it talk about "a file" in a source code repository, seems pretty clear - a source code repo has lots of files. at the most recent revision, some exist, some were deleted in some past revision. more will be added (and deleted) in later revisions. it's very much not the same model as git. hope that clears things up.
- quantum_state 4y agosomething is seriously wrong if Google needs 2B loc to do its things …
- 0x6c6f6c 4y agoHow do you propose you provide the number of services Google has without lots of code? For context, the entirety of the Google suite is in there, and a lot more. I'm even somewhat surprised it's that little with their scale.
- deleted 4y ago[deleted]
- thwoeriuowie 4y agoGoogle's code may be a monorepo, but back when I was there you only ever 'checked' out particular projects for editing etc. It's a bit silly to talk about some aspects of Google separated from the whole dev env in there.
- denvercoder904 4y agoIs the code for the Search project in the mono repo as well? How does Google handle access control for their mono repos? Where's the secret sauce stored?
- charcircuit 4y agoThere is directory / file level ACL. Due to AI the secret sauce isn't as important as all of the data. Recommendation algorithms don't need to be super confidential since it ultimately turns into "make content that people will want recommended to them."
- gorgoiler 4y agoImagine you have two teams in one monorepo and requirements.txt has pinned numpy at 1.22. One team wants to upgrade to 1.24 but the upgrade breaks the other team’s code as it was dependent on an emergent property* in the older version of numpy. How would you handle this situation as an IC? As a manager of one of the teams? As a skip-level manager of both teams? As a budding IC on the team that wants the upgrade, you may want to go fix up the other team’s code for them so you can bring them along with the upgrade. Realistically, the further you get from Google’s level of engineering discipline and skill the more likely you are to encounter the following in the needs-1.22 codebase: - horrible code that is hard to understand and therefore hard to refactor - code with no tests, making it risky to refactor - the team that wrote it have all left or been fired and no one is available to help understand it - they are a remote team with no social relationship to you who interact entirely online, in writing, in the style of an aggressive subreddit mod - deeply entrenched factions mean that even if you offer them a patch they will default refuse it because who are you to work on their codebase and they don’t need the upgraded numpy so why should they waste resources on reviewing something they don’t want - misguided adherence to status enhancing terms like “audit” and “compliance” mean jobsworth ICs refuse to even look at your patch because someone somewhere once heard a friend of a friend whose company failed SOC2 because engineer from floor X made a change to code owned by floor Y and it went against policy All of these social problems are real ones I have encountered and if you have solved these then you’re probably already happily in a monorepo already. If instead you work in an org full of teams pointing guns at each other in a fight to the death to stop any kind of cross org collaboration from sullying the purity of the tribal system then know this: it gets better, and if you build the right social connections then the technical efficiency of having your monobusiness executing its monomission inside a monorepo is within reach! *bug
- fouronnes3 4y agoWhile I found your comment insightful and sadly very accurate, it's fundamentally a human problem, not a technical problem. So I don't think the solution to it should be technical like "don't use a monorepo and those problems will go away!", but rather organisational in nature.
- 4y ago
- teleforce 4y agoPrevious discussions on HN (2020): https://news.ycombinator.com/item?id=22019827 https://news.ycombinator.com/item?id=22019827
- dgnemo 4y agoBig fan of monorepo approach here. Still, I have recently hit a major issue with the fact that GIT (and other common version control sw) don't have per-directory ACL. Has anyone dealt with this issue? Which VCS / configuration have you adopted?