13 ms·
Why Google Stores Billions of Lines of Code in a Single Repository (2016)
- dfjpitcher 8y agoBecause its founders Brin and Page have inborn genetic belief in a single Jewish God. It's the same force that led Einstein to search for Grand Unified Field Theory.
- gervase 8y agoShould probably have a [2016] tag.
- guessmyname 8y agoIndeed, but to be fair, the information in the article is based on several research papers from 2011 [1]. And I am 100% sure the idea of having a monolithic project is several years older than that. I am grateful that the article is re-posted in multiple websites, because just the other day I was in an interview and, while doing my coding challenge, overheard the conversation of a young computer science graduate and another interviewer. The interviewer asked him to explain what was a monolithic repository and the benefits. This guy had no idea what the interviewer was talking about and right there I realized that what many of us take for granted terminology-wise in the IT world, will certainly be a foreign language to young students who are just entering the work force. [1] http://info.perforce.com/rs/perforce/images/GoogleWhitePaper-StillAllonOneServer-PerforceatScale.pdf http://info.perforce.com/rs/perforce/images/GoogleWhitePaper...
- paulie_a 8y agoIm sure properly organized it's okay, but from what I've seen it's mediocre at best, especially with legacy/technical debt it's a huge mistake. Start breaking that repo apart, because it probably isn't very/hopefully depending on the debt that exists.
- oxguy3 8y agoOne of the big advantages of the monorepo is actually that it prevents technical debt from accumulating. If a change somewhere else breaks your code, you can't put off dealing with it -- you are forced to fix the issue immediately.
- paulie_a 8y agoThat makes a lot of sense and I definitely like the idea of that. Unfortunately unless you either spend a tremendous amount of effort in a legacy system to make that reality, or start with a new green field, it's not realistic day to day.
- erik_seaberg 8y agoTech debt is a useful tool and we shouldn't have zero tolerance. I can see wanting to deprecate old versions promptly, but I can't see instantly deprecating every old version with no workaround for mitigating emergencies.
- techbio 8y agoPrevious thread: https://news.ycombinator.com/item?id=11991479 https://news.ycombinator.com/item?id=11991479
- deleted 8y ago[deleted]
- curtis 8y agoI think monorepos make a lot of sense when you're talking about millions of lines of code. I'm not at all sure they make sense when you're talking about billions.
- gravypod 8y agoI don't think the number of linea matters. I think the interconnection of your code matters. If you have 2 sets of services that are completely uncoupled the having two monorepos for those two deployments make sense. If you can guarantee atomic changes across all services that interconnect you have the benefits monorepos give you.
- ebikelaw 8y agoEven at google this is true. There are naturally multiple monorepos :) For example the Linux kernel devs have their own. This makes sense since the kernel-user interface is strongly defined.
- mason55 8y agoIsn’t this only true if you’re doing full CI? Otherwise I could update my service and you can update yours to work with mine but unless we coordinate deployments you still have to worry about interface mismatches. I guess the alternative is you can just never (for a loose definition of never) make breaking changes to an interface. You can only enhance or create a new version.
- jldugger 8y agoWell, this particular monorepo has two billion LoC. But it's not a git monorepo, which matters significantly.
- ChrisCinelli 8y agoManaging dependencies and versions across repos is a pain. Refactoring across repos is quite hard when your code spreads across repos considering the tree of dependencies. Unfortunately Git checkout all the code, including history, at once and it does not scale to big codebases. The approach that Facebook chose with Mercurial seems a good compromise ( https://code.fb.com/core-data/scaling-mercurial-at-facebook/ https://code.fb.com/core-data/scaling-mercurial-at-facebook/ )
- antt 8y agoGit works very well when the code is distributed. Which funnily enough is in the name. That we are using git as a centralized repository is a case of "Why do I need a screwdriver when I have a hammer?".
- EpicEng 8y ago> Managing dependencies and versions across repos is a pain
- shub 8y agoThere's nothing about git that requires you to use it like the kernel does. Centralized version control is just a special case of decentralized, if you're using git. You still get the benefits of your repo being a peer of the master repo, like local branches.
- csdreamer7 8y agoDoesn't the Git Virtual File system that Microsoft is contributing to Git take care of this? https://blogs.msdn.microsoft.com/devops/2017/02/03/announcing-gvfs-git-virtual-file-system/ https://blogs.msdn.microsoft.com/devops/2017/02/03/announcin... Edit: don't just down vote. If you have a problem with my comment, tell me why.
- justincormack 8y agoThis currently only works on Windows, although they are planning OSX and Linux ports.
- deleted 8y ago[deleted]
- ridiculous_fish 8y ago> Google's monolithic software repository, which is used by 95% of its software developers worldwide, meets the definition of an ultra-large-scale4 system, providing evidence the single-source repository model can be scaled successfully This 95% number is the most surprising part of the article. That implies that the sum of engineers working on Android + Chrome + ChromeOS + all the Google X stuff + long tail of smaller non-google3 projects (Chromecast, etc) constitute only 5% of their engineers. Is e.g. Android really that small?
- dlubarov 8y agoThey must have meant that 95% of Google engineers use the monorepo in some capacity, even if the majority of their work is done in a different repo.
- dlp211 8y agoI think you're interpretation is incorrect. A better way to think of this is that those 5% of people work exclusively on those projects. I'd be very surprised to learn that only 5% of Google engineers work on those projects.
- hyperpape 8y agoI don’t know how to parse the number, but 5% of a billion still leaves 50 million lines of code, or three Linux kernels worth.
- Karishma1234 8y agoThe 95% number probably does not mean what you are saying but what it means is 95% of developers are using it for some reason with say a non-zero commits over lifetime of a developer OR simply checking it out and using it for dependancies.
- Too 8y agoThat 95% is most likely more figurative than fact.
- jbergknoff 8y agoHow does CI work with a monorepo? Do you always have to run all the tests and build all the artifacts? Or are there nice ways to say "just build this part of the repo"?
- FartyMcFarter 8y agoFor safe-looking changes, it's OK to only run a subset of the tests (usually including the tests that directly test the changed library). For changes that are more likely to break distant code, you can run all tests (perhaps bundling together several changes in order not to overload the system). Alternatively you can take the risk of breaking tests post-submit... this is not very good citizenship, but in some cases it might be reasonable (when the risk is small).
- ebikelaw 8y agoDependencies are explicit so the build tool (Bazel, to an approximation) compute the transitive closure of requirements of the desired target. There are more details about testing at [1] 1: https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/45861.pdf https://static.googleusercontent.com/media/research.google.c...
- dekhn 8y agoYou specify targets. Just like using bazel: bazel build //tensorflow/blah/.... I maintain a small part of the monorepo, and it's really nice to be able say "Run every test that transitively depends on numpy with my uncommitted changes", so you can know if your changes break anybody who uses numpy when you update the version. Personally I think it would be neat if there was an external "virtual monorepo" that integrated as-close-to-head of all software projects (starting at the root, that's things like absl and icu, with the tails being complex projects like tensorflow), and constantly ran CI to update the base versions of things. Every time I move to the open source world, I basically have to recompile the world from scratch and it's a ton of work.
- dlubarov 8y agoIt's flexible; presubmit tests can be configured per-directory. There's also an option to run all tests of packages that could be affected by a change based on the Blaze dependency graph. If you're making changes to a package with tons of dependencies such as Guava, for a risky change you might want to run all affected tests, but for a minor change you might want to run just the standard unit tests. As a compromise, there's also an option to run a random sample of affected tests.
- throwaway021916 8y agoBecause its founders Brin and Page have inborn genetic belief in a single Jewish God. It's the same force that led Einstein to search for Grand Unified Field Theory.
- hobls 8y agoI feel terrible for anyone who sees this and thinks, “ah! I should move to a monorepo!” I’ve seen it several times, and the thing they all seem to overlook is that Google has THOUSANDS of hours of effort put into the tooling for their monorepo. Slapping lots of projects into a single git repo without investing in tooling will not be a pleasant experience.
- nine_k 8y agoMost folks who consider a monorepo don't have billions of lines of code, and often not even millions. Linux kernel is a monorepo.
- threeseed 8y agoLinux kernel is one functional piece of work though. Imagine if we combined KDE, Gnome, Linux Kernel, ZFS etc all in the one monorepo.
- truthof080616 8y agoBecause its founders Brin and Page have inborn genetic belief in a single Jewish God. It's the same force that led Einstein to search for Grand Unified Field Theory.
- jorblumesea 8y agoIs this really relevant for anyone except for "google scale" companies? For most teams, managing 30-40 services backed by git repos isn't a huge task and doesn't cause many problems. Is there mature tooling that helps teams manage this, or is this proprietary google magic tooling?
- fastball 8y agoMost teams can probably get by with much fewer than 30-40 services. Unless you have 30-40 groups within your team.
- jorblumesea 8y agoEven if they had that, managing the contract between and in a small group isn't super difficult.
- vbezhenar 8y agoI've used monorepo for few small related projects and it worked just fine for me. Much easier to make related changes across several projects.
- whack 8y agoMaybe I'm not cool enough to understand this, but I don't see the draw for monorepos. Imagine if you're a tool owner, and you want to make a change that presents significant improvements for 99.9% of people, but causes significant problems for 0.1% of your users. In a versioned world, you can release your change as a new version, and allow your users to self-select if/when/how they want to migrate to the new version. But in a monorepo, you have to either trample over the 0.1%, or let the 0.1% hold everyone else hostage. Conversely, imagine if you're using some tools developed by a far off team within the company. Every time the tooling team decides to make a change, it will immediately and irrevocably propagate into your stack, whether you like it or not. If you were at a startup and had a production critical project, would you hardcode specific versions for all your dependencies, and carefully test everything before moving to newer versions? Or would you just set everything to LATEST and hope that none of your dependencies decide to break you the next day? Working with a monorepo is essentially like the latter.
- perfunctory 8y ago> Working with a monorepo is essentially like the latter. Not really. In the dependencies analogy the author of the dependency has no way to test the dependee(s). While with monorepo this is exactly what you do, "the tooling team" will "carefully test everything" before "propagate into your stack" (and it doesn't have to be irrevocable).
- whack 8y agoIn practice, at any medium/large organization, the tooling team doesn't know your system, and its nuances, nearly well enough to "carefully test everything". Having a solid automated test suite does help. But I personally would like to be in control of when my project updates its dependencies, instead of being forced to always pull everything from LATEST.
- perfunctory 8y ago> the tooling team doesn't know your system they don't know it because they don't use monorepo. monorepo makes "solid automated tests" easier since basically there is only one version to test. The instinct against pulling everything from LATEST, developed in the traditional world, is perfectly understandable. However in monorepo "your" project is also tooling-team's project. "being forced" becomes "being helped". It's shared responsibility.
- a-dub 8y agoIt should be noted that the monolithic model is somewhat encouraged by the client mapping system in Perforce, which was Google's first version control system so it is unclear to me if this was deliberate or just a side effect of the best VCS of the time. I also still have doubts around the value of a monorepo, in the article they claim it's valuable because you get: Unified versioning, one source of truth; Extensive code sharing and reuse; Simplified dependency management; Atomic changes; Large-scale refactoring; Collaboration across teams; Flexible team boundaries and code ownership; and Code visibility and clear tree structure providing implicit team namespacing. With the exception of the niceness of atomic changes for large scale refactoring, I don't really see how the rest are better supported by throwing everything into one, rather than having a bunch of little repos and a little custom tooling to keep them in sync.
- malkia 8y agoIncrementally monolithic CL number is also useful. You can mark quite a lot of things with it - not only binary releases, but other developments too (configuration files, etc.). At the end your binary "version" comprises of main base CL + cherrypicked individual CL's - rather than branch with these fixes - I guess one can encode this too with git/hg - by using sha hashes - but this becomes much bigger in terms of information, and human handling it. I guess not very strong point, but using CL numbers (I'm working with perforce mostly these days) makes things easier. And having one CL monothonically increasing all over all source code you have even better - you can even reference things easier - just type cl/123456 - and your browser can turn it into a link. Among many other not so obious benefits...
- lpghatguy 8y agoMost popular Git frontends (GitHub and GitLab too, I believe) let you link to commits with just the first 5-6 characters of the hash. I don't think that's much different to remember than a Perfore CL number.
- malkia 8y agoTo me the issue is when mentally trying to work with these numbers, P4 & G4's numbers increment, so I can tell which one came before the other - I can't do this with hashes. I'm sure I can get used to the other way, but this cannot easily be ignored.
- prepend 8y agoI love these articles. Is there a wiki or collection of detailed descriptions of large company tech practices that isn’t marketing blargh. I read years ago about Google data ingest, locator process but neglected to bookmark so now can’t find the reference.
- mlinksva 8y agoMe too. I don't know of a collection, but others can be found at https://ai.google/research/pubs/ https://ai.google/research/pubs/ https://research.fb.com/publications/ https://research.fb.com/publications/ https://www.microsoft.com/en-us/research/search/?q&content-type=publications https://www.microsoft.com/en-us/research/search/?q&content-t... and similar (though only a small fraction give hints about at scale practices, and those would be neat to collect in one place). Closely related to this post: just noticed a 2018 case study on Advantages and Disadvantages of a Monolithic Repository https://ai.google/research/pubs/pub47040 https://ai.google/research/pubs/pub47040
- emmelaich 8y ago(2016)
- senozhatsky 8y agoWell, it's not so uncommon. For instance, OpenBSD, NetBSD repos are sort of monolithic. And, believe it or not, there are some advantages. For instance, let's take a look at OpenBSD 5.5 [0] release notes: > OpenBSD is year 2038 ready and will run well > beyond Tue Jan 19 03:14:07 2038 UTC OpenBSD 5.5 was released on May 1, 2014. While Linux is still "not quite there yet" y2038-wise. y2038 is a very complex issue, while it may look simple - time_t and clock_t should be 64-bit. This requires changes both on the kernel -- new sys-calls interfaces [stat()], new structures layouts [struct stat], new sizeof()-s, etc. -- and the user space sides. This, basically, means ABI breakage: newer kernels will not be able to run older user space binaries. So how did OpenBSD handle that? The reason why y2038 problem looked so simple to OpenBSD was a "monolithic repository". It's a self-contained system, with the kernel and user space built together out of a single repository. OpenBSD folks changed both user space and kernel space in "one shot". IOW, a monolithic repository makes some things easier: a) make a dramatic change to A b) rebuild the world c) see what's broken, patch it d) while there are regressions or build breakages, goto (b) e) commit everything [0] http://www.openbsd.org/55.html?hn http://www.openbsd.org/55.html?hn [UPDATE: fixed spelling errors... umm, some of them] -ss
- glandium 8y agoThe reason why y2038 problem looked so simple to OpenBSD has little to do with "monolithic repository" and everything to do with "happy to break kernel ABI compatibility". You're saying as much yourself. Monolithic repository might have been a tool that helped enforce it, but that's not what made it happen. It's the decision that ABI could be broken that did. And that's also why it hasn't happened in Linux yet. Even if there was a monorepo containing all the open source and free software in the world (or at least, say, that you can find in common distros), the fact that there's a contract to never break the ABI makes it simply hard to do.
- dguest 8y agoThis is a very good point. I work in an organization that just switched to monolithic and it's been going very well, with hundreds of active developers and millions of lines of code. But our developers are students or academics. As many as half don't understand the concept of an ABI. So the monorepo works quite well for us because rebuilding from scratch is something we do multiple times a day.
- haglin 8y agoGoogle's handling of their source code makes me wanna work there. I don't like distributed version control systems with hundreds of repositories spread out. It makes management more complicated. I understand this is a minority view, but that is my experience. It was easier to work in a single Perforce repository than hundreds of Git or Mercurial repos.
- djur 8y agoDistributed vs. centralized VCS has very little directly to do with many vs. monolithic repos. After all, git was originally developed for a project with a large monolithic repo. Distributed VCS and many small repos got popular around the same time, but that's partly coincidental (microservice architectures getting popular, npm community preferring extremely small libraries) and partly because of GitHub making it very cheap in money/time to have many git repos.
- nicodjimenez 8y agoI have slight experience with both monorepos and smaller repos and I think they can both work. The advantage of smaller repos is that it forces different components to expose well designed API's. Bigger repos make sense for products and embedded software, smaller repos make sense for platforms build up of small services communicating on the internet.
- djur 8y agoSmaller repos force different components to expose APIs, but I don't think it forces or even encourages the APIs to be well designed. In some cases, having work spread across multiple repos can impede iterative development, meaning that you risk half-assed or, uh, two-and-a-half-assed implementations. Also, when someone's asking for review for a change that encompasses, say, a change to a service, a change to a client library for that service, and a change to 2-3 other services that use that client library, I know that I cringe a little when suggesting a change, knowing that to implement it is going to require a commit on all of these different repos, waiting for CI to run on each one, etc. I try to only use that impulse to counter the urge to bikeshed, but the temptation is there.
- mlthoughts2018 8y agoOne of my former managers had worked a long time at Google and was present for the advent of Google’s in-house tooling developed around their monorepo. His account was that it was basically accidental, at first resulting from short term fire drills, and then creating a snowball effect where the momentum of keeping things in the Perforce monorepo and building tooling around it just happened to be the local optimum, and nobody was interested in slowing down or assessing a better way. He personally thought working with the monorepo was horrible, and in the company where I worked with him, we had dozens of isolated project repos in Git, and used packaging to deploy dependencies. His view, at least, was that the development experience and reliability of this approach was vastly better than Google’s approach, which practically required hiring amazing candidates just to have a hope of a smooth development experience for everyone else. I laugh cynically to myself about this any time I ever hear anyone comment as if Google’s monorepo or tooling are models of success. It was an accidental, path-dependent kludge on top of Perforce, and there is really no reason to believe it’s a good idea, certainly not the mere fact that Google uses this approach.
- gefh 8y agoDo you wonder whether he is a reliable narrator?
- mlthoughts2018 8y agoI don’t, but it’s fair to ask. He was unequivocally the best senior manager I’ve worked with. Extremely technically smart but skilled at letting people under him work autonomously, good communicator, cared a lot about pushing best practices past bureaucratic barriers. His description of Google made it seem like it had the same dysfunction every place has. And the monorepo was a totally mundane, garden variety eyesore kind of in-house framework that you’ll find anywhere. I think he recognized the usefulness of just working with it and picking battles. He was just dumbfounded that any outsider would see the monorepo project and think it possibly had any relevance for anyone else. It was just a Google-history-specific frankenstein sort of thing that got wrangled with tooling later. The supposed benefits are all just retrofitted on.
- malkia 8y agoHere is the video (with Rachel Potvin), predating the article by some months: https://www.youtube.com/watch?v=W71BTkUbdqE https://www.youtube.com/watch?v=W71BTkUbdqE
- testcross 8y agoI don't understand why gitlab/github/bitbucket don't provide better tools for monorepo. This is a topic pretty trendy. But there is absolutely no tools helping with control access, good ci, ...
- malkia 8y agoWhat's missing in these is cross-reference, which is not possible without somewhat established BUILD system (caps "pun-intened") - e.g. like bazel/build, then a source code indexer, etc, etc. This becomes very critical for doing reviews, since it allows you to "trace" things without running them, apart from many other things. For example large scale refactorings looking for usages of functions, and other examples like it. Why githab/gitlab/etc. can't do it? Well because hardly there could be one encompassing BUILD system to generate correctly this index.
- testcross 8y agoThey can create a standard file format that has to be generated by build system. github is in a pretty powerful position. They can create even a shitty version of it and people will follow. I've been thinking about a tool like this for a long time. A way to attach to each commit not only the diff in the code, but also the list of places affected by the changes (usages of functions that are modified for example). Then during review we wouldn't have only a stupid diff. We would have a list of place to check to be sure that the changes make sense in the context of the project.
- malkia 8y agoEven if they can, it's one thing indexing your own source files every night, another indexing a much bigger amount + massive amounts of branches, clones, etc. (I'm talking about github) - e.g. not practical - as there is no no clear way to say which branch (from git) must be indexed (obviously not all) - e.g. there is no encompassing "standard" saying so. That by itself is another BIG PLUS for mono-repo (and "mono"-rules) - things are done one (opinionated) way, trunk based development - but thus giving you things that you won't be able to have normally. Now indexing source file is not an easy and cheap task - it's basically a huge MapReduce done over several hours (just guessing), so there must be a reason for this to be done.
- makecheck 8y agoThis is clearly detrimental to external projects such as Go packaging, since their own developers will never be looking at dependency problems in the same way as outside groups. Monorepo also bugs me because there will always be some external package you need, and invariably it’s almost impossible to integrate due to years of colleagues making internal-only things assume everything imaginable about the structure and behavior of the monorepo. There will be problems not handled, etc. and it leads to a lot of NIH development because it’s almost easier in the end. Also, it just feels risky from an engineering perspective: if your repository or tools have any upper limits, it seems like you will inevitably find them with a humongous repo. And that will be Break The Company Day because your entire process is essentially set up for monorepo and no one will have any idea how to work without it.
- topspin 8y ago> This is clearly detrimental to external projects such as Go packaging Indeed. Google's monorepo means the largest cohort of Go programmers in the world are mostly indifferent to composing packages in the usual (cpan/maven/composer/npm/nuget/cargo/swift/pip/rubygems/bower/etc) manner. Non-Google Go programmers have been left to schlep around with marginal solutions for years, although in the last few months we begin to see progress here[1]. This was the #1 discouragement I experienced when experimenting with Go. Google's monorepo may be wonderful from Google's perspective but I don't think it's been a win for Go. * yes I know some of these are also build systems and provide many other capabilities, some of which are arguably detrimental. Versioned, packaged, signed dependencies and thus repeatable build artifacts is the point. [1] https://github.com/golang/go/issues/24301 https://github.com/golang/go/issues/24301
- robaato 8y agoWhat about Android and 800-1,000 git repos?! Have seen the pain trying to manage that across larger teams (e.g. thousands of devs) - and no the "repo" tool is not sufficient.
- nwlieb 8y agoI'm very curious, what pain did you see with the repo.py tool?
- the_arun 8y agoSeems like Google uses its own custom Source Control & tools - https://www.quora.com/What-version-control-system-does-Google-use-and-why https://www.quora.com/What-version-control-system-does-Googl....
- tzhenghao 8y agoHaving worked at different companies adopting both monorepo and the multiple repos approach, I find monorepo a better normalizer at scale in consolidating all "software" that runs the company. Just like what many commenters here have mentioned, the monorepo approach is a forcing function on keeping compatibility issues at bay. What you don't want is to end up in a situation where teams reinvent their own wheels instead of building on top of existing code, and at scale, I think the multiple repo approach tends to breed such codebase smell. [1] I'm sure 8000 repos is living hell for most organizations. [1] - https://www.youtube.com/watch?v=kb-m2fasdDY https://www.youtube.com/watch?v=kb-m2fasdDY
- shiift 8y agoI really liked that talk! Lots of relevant information and I can definitely relate, working at a Amazon. Wouldn't say that we are hurt by all of the same problems (we have solutions that work very well for some of them), but we definitely are aware of them.
- carapace 8y agohttps://en.wikipedia.org/wiki/Conway%27s_law https://en.wikipedia.org/wiki/Conway%27s_law > "organizations which design systems ... are constrained to produce designs which are copies of the communication structures of these organizations." Interestingly, in light of the above adage, this massive repo is organized (if that's the word for it) like a bazaar or flea market. (Rather than like a phone book https://en.wikipedia.org/wiki/Yellow_pages https://en.wikipedia.org/wiki/Yellow_pages )
- fizixer 8y agoI don't care about that. For me this is incomprehensible: Why the eff does Google have billions of lines of code in their repo? I hope they are not counting revisions (e.g., if a single 1 million project has 100 revisions, that's 1 million, not 100 million). I have heard that they do count generated code (so it's not all handwritten code). In that case again, I have two things to say: - that's a bad metric. I could overnight generate a billion lines of code with each line a printf of number_to_word of numbers from 1 to a billion. They want to measure the size of the repo? They should tell us the gigabytes, terabytes etc. But when it's lines of code, it's cheezy and childish to blow up the measure by including lines of generated code. - But more importantly, I hope the generated code is 90% or more of that repository. Because any less than that would mean that Google engineers have handwritten 100 million or more lines of code through out the lifetime of the company, in which case I have to ask: what bloated mess do you have on your hands? I thought you guys were the top engineers of the world.
- wrayjustin 8y ago> includes approximately one billion files ... > including approximately two billion lines of code _also_ > in nine million unique source files I should insert a joke about how well the system would do if each source file contained more than two lines of code. But seriously, this summary could use some work.
- rpcastagna 8y agoBinary files (arbitrary example: images used for golden screenshots in tests) have no line counts and are likely skewing the numbers here -- in the way you're (logically) looking to interpret them at least. From a system design perspective, being able to handle a large number of files regardless of type is an interesting challenge, as is being able to handle a large number of highly indexed text files. All three of those statistics seem potentially interesting for different audiences that might read this paper.
- tflinton 8y agoA repo including configuration and data. How about we stop considering google an engineering leader and just a search leader?
- tflinton 8y agoA repo with configuration, secrets and data? Can we stop considering google an engineering leader and just a search algorithm leader?
- IloveHN84 8y agoThe giant monorepo works only if you're using SVN, with Git it would be tremendous
- therealmarv 8y agoUnless you change git like Microsoft did: https://blogs.msdn.microsoft.com/bharry/2017/05/24/the-largest-git-repo-on-the-planet/ https://blogs.msdn.microsoft.com/bharry/2017/05/24/the-large...
- HNNewer 8y agosure, but GitHub / GitLab / Bitbucket don't offer it
- axaxs 8y agoSorry, but as someone who has been in orgs that do both, mono repo is a mistake. Constant needs to pull unrelated changes before pushing, pipelines requiring to grab the whole repo for dependencies, etc. I understand the arguments for mono repo, but never think it's nothing that outweighs the cons.
- robaato 8y agoWell those are issues around having a git mono repo - where the repo is the unit of change - you get it or you don't. With mono repos such as SVN or Perforce you just work on whatever subset you want.
- tsycho 8y agoIt's not just devops that you need to pull off a large monorepo; the other big thing is a strong testing culture. You have to be able to rely on unit tests from across the code base being a sufficient indicator of whether your commit is good. AND a presubmit process that can compute which parts of the monorepo get affected by your diff, and run tests against them automatically before committing your diff. Google not only has the above but also has a strong pre-submission code review process which catches large classes of bugs in advance.
- timkrueger 8y agoWe work with an monorepo since Septemeber 2017. I wrote about the migration: https://timkrueger.me/a-maven-git-monorepo/ https://timkrueger.me/a-maven-git-monorepo/ Our developers like it, because they can use 'mkdir' to create a new component, search threw the complete codebase with 'grep' and navigate with 'cd'.
- jgibson 8y agoIs it just me, or are a lot of people here conflating source control management and dependency management? The two don't have to be combined. For example, if you have Python Project X that depends on Python Project Y, you can either have them A) in different scm repos, with a requirements.txt link to a server that hosts the wheel artifact, B) have them in the same repo and refer to each other from source, or C) have them in the same repository, but still have Project X list its dependency of project Y in a requirements.txt file at a particular version. With the last option, you get the benefit of mono-repo tooling (easier search, versioning, etc) but you can control your own dependencies if you want. edit: I do have one question though, does googles internal tool handle permissions on a granular basis?
- bananarepdev 8y agoMaybe this is a reflection of modern tools using the version control system to store built artifacts, like npm and "Go get" do. Anyway, depending on the programming language, you can have a monorepo and still bind your modules with artifact dependecy, not necessarily depending on the code itself.
- maccard 8y agoI can't comment specifically on Google's tool, but I know it's based on perforce. perforce does have granular permissions - https://www.perforce.com/perforce/r15.1/manuals/p4sag/chapter.protections.html https://www.perforce.com/perforce/r15.1/manuals/p4sag/chapte...
- justicezyx 8y agoSingle repo is one design that coherently addresses source control management and dependency management. The key is to let the repo be a single comprehensive source of data for building arbitrary artifacts.
- evfanknitram 8y agoI don't know what this means. How is "single repo" a "design" and how does this design dictate dependency management? Yes, if you have a single repo then that would be a single source of data for building your stuff. That seems redundant.
- stevesimmons 8y agoMy company has a 50m LOC Python codebase in a monorepo. It works really well, given the rate of change of thousands of developers globally. That is only possible because of the significant investment in devtools, testing and the deployment infrastructure. Here is "Python at Massive Scale", my talk about it at PyData London earlier this year: https://youtu.be/ZYD9yyMh9Hk https://youtu.be/ZYD9yyMh9Hk
- jamesmiller5 8y agoI wish more developers knew of the wonderful "repo" tool[0] developed by the Android devs which allows a monorepo _perspective_ of many git repositories. Breakdown of the repo tool and example manifest files http://blog.udinic.com/2014/05/24/aosp-part-1-get-the-code-using-the-manifest-and-repo/ http://blog.udinic.com/2014/05/24/aosp-part-1-get-the-code-u... [0] https://source.android.com/setup/develop/repo https://source.android.com/setup/develop/repo
- joe_fishfish 8y agoThis is probably a stupid question, but I couldn't find an answer. Does this mean Google keeps all of its different products in all their different languages and environments in one repo? So like, Android lives in the same repo as Gmail, which is the same repo as all the Waymo code and the Google search engine code as well? That seems insane to me.
- paulddraper 8y agoVersion controlled repositories are like business offices. You can have your entire company in one location, or the entire company in separate locations. The most important thing is the logical rather than physical organization: team structure, executive leadership, inter-org dependencies, etc. You can achieve autonomy and good structure with or without separate locations. A single location reduces barriers, but at some point multiple locations can solve physical and logistical challenges. General rule of thumb is to own and operate office space in a few locations as possible, but at some point you have to take drastic measures one way or another. (Notice that Google had to invent their own proprietary version control system just for their monorepo. And not even Google actually uses a single repo as the source of truth: e.g. Chromium and Android.)
- alexeiz 8y ago> Trunk-based development. ... is beneficial in part because it avoids the painful merges that often occur when it is time to reconcile long-lived branches. Development on branches is unusual and not well supported at Google, though branches are typically used for releases. This sounds like the SVN model to me where branches are cumbersome and therefore they are very rare. After getting used to the Git branching model where branches are free and merges are painless, it would be very hard to go back to the old development model without branches.