10 ms·
I feel terrible for anyone who sees this and thinks, “ah! I should move to a monorepo!” I’ve seen it several times, and the thing they all seem to overlook is t
by hobls 8y ago
I feel terrible for anyone who sees this and thinks, “ah! I should move to a monorepo!” I’ve seen it several times, and the thing they all seem to overlook is that Google has THOUSANDS of hours of effort put into the tooling for their monorepo. Slapping lots of projects into a single git repo without investing in tooling will not be a pleasant experience.
- nine_k 8y agoMost folks who consider a monorepo don't have billions of lines of code, and often not even millions. Linux kernel is a monorepo.
- threeseed 8y agoLinux kernel is one functional piece of work though. Imagine if we combined KDE, Gnome, Linux Kernel, ZFS etc all in the one monorepo.
- dmoy 8y agoAnd gnucash, libreoffice, a couple copies of android, three other things that forked the linux kernel, and then all of apache to boot.
- rhencke 8y agoAnd we'll call it something crazy, like a Linux distribution!
- nine_k 8y agoWhere there are obvious and pronounced functional boundaries, often backed by administrative boundaries, separate repos makes total sense. Otherwise, it's an optimization; see "premature optimization" for cautions.
- pcwalton 8y agoThere's also the fact that monorepos have issues when you don't have one organization responsible for all the code. The Linux kernel and NetHack don't live in the same repository for good reason.
- dekhn 8y agoI dunno, the BSD distribution included a wide gamut of games along with the kernel source in the same tree. In fact, NetHack is derived from Hack which itself is derived from Rogue, which was distributed within BSD. And BSD represented a cross-organization responsibility (see the history of AT&T and BSD).
- akvadrako 8y agoIt does seem like the Linux model scales better than the BSD model.
- trasz 8y agoWasn't it because it was the same group of people who worked on both? And when it ceased to make sense, the games were split off - the only remaining ones in FreeBSD are things like banner(6) or pom(6).
- pcwalton 8y agoFine, replace NetHack with Quake 3. :)
- rco8786 8y agoAre you suggesting that there’s a solution to managing large amounts of code that doesn’t involve large amounts of tooling?
- hobls 8y agoThere’s an important distinction between lots of code and lots of projects. I agree; if you have a ton of code you’d better invest in tooling. But if you just have several normally sized projects, a monorepo can make your life much more difficult than simply using several repos.
- rco8786 8y agoSure. You could also replace the term “monorepo” with “separate repos” and your statement would be just as valid. Either way you go has pros and cons.
- hobls 8y agoAgreed. I think the thing is that GitHub basically supports the separate repo approach fairly well out of the box. Using only a single repo requires more thought around your strategy, especially if you have multiple teams.
- acomjean 8y agoBut what are the alternatives to the monorepo in git? All the ways of splitting code up and deploying multiple git repos for one project seem terrible.
- forrestthewoods 8y ago> what are the alternatives to the monorepo in git? A monorepo in Perforce!
- threeseed 8y agoI don't understand why people are against this. You can have per repo branches/tags, the history is clean and relevant, it's easy to triage breakages, easy for different apps to have different versions of code etc. Plus for CI/CD it's trivial to just have one Jenkins jobs per repo as well and simple Git commit triggers. The entire programming world revolves around libraries and yet when it comes to our own code we are afraid of them ? Strange.
- reificator 8y ago> multiple git repos for one project If it's one project, it's not a monorepo. It's a repo.
- acomjean 8y agoWe wanted to have a "common" subsystem that was common across projects. Being able to add and work on the common area and new projects at the same time was important. Pushing the common area back and being able to deploy to the older projects and test was important. This seems difficult in git. There are "submodules" and "subtrees" but none seemed particularly great and as far as I could tell each came with a bunch of caveats. I'll admit my Git skills aren't great, but I've used a variety of source control and tried to suss out the best way to deal with a small team. We ended up using "git subrepo" which is an add on thing I don't love, but it works. part of the motivation is "common" and "project 2" are to be open sourced, but "project 1" which also uses "common" isn't.
- deleted 8y ago
- mr_tristan 8y agoI sense that Google invests much more in it's infrastructure then most companies make in revenue. I've worked with monorepos, and I'd be loathe to recommend it as well; the combination of culture shift and tooling it takes to keep a monorepo system running makes most CD processes you see today look like child's play. There is a lot of very good free software that supports most of the open source approach to CD these days; but very, very little freely available monorepo tooling. Just check out https://github.com/korfuri/awesome-monorepo https://github.com/korfuri/awesome-monorepo - it's a quick read. I haven't found many other notably superior compilations. Compared with available OSS workflows and tooling, it's rather sparse, filled with bespoke approaches everywhere.
- ryancox 8y agoAgreed about the lack of monorepo tooling. There's just not that much out there. A couple of other links I didn't see in the awesome-monorepo: - https://github.com/facebookexperimental/mononoke https://github.com/facebookexperimental/mononoke - I hear this is a real thing and not a science fair project - https://github.com/bors-ng/bors-ng https://github.com/bors-ng/bors-ng - Needed in a monorepo to handle high arrival rate of commits / merges
- golangnews 8y agoWhat problems specifically did you see? Was this because the repo was too large? I understand at google scale you'd need lots of tooling but why at a smaller scslr of merging a dozen small repos?
- mr_tristan 8y agoThe biggest problems are always cultural. Most monorepo workflows really reinforce constant integration, and once you have separate teams with separate managers, I've always witnessed constant conflict that ended up trying to establish spheres of control. It's bizarre - but it's something I've seen at pretty much every place I've worked at. With all that integration, your single CI toolchain is front and center since everyone's success or failure is tied to it. While projects like bazel exist, how many developers do you know work with bazel every day? I no nobody who does. And most want documented IDE support and ease of use, not some optimal CI workflow. I've found gradle to be OK, but even that kind of pushes everyone toward using Jetbrains tooling. In the end, almost real monorepos have significant custom CI tooling that wires together different toolchains, and, they may have to maintain custom tooling for use in developer machines. And that custom tooling can get expensive to maintain as the project scales up.
- hobs 8y agoheh, thousands, its probably at least an OOM greater, if not two.
- waterhouse 8y agoMy brain sees OOM and thinks "out of memory", which might be applicable too.
- hobls 8y agoHah, you know I started with millions and then did some fuzzy math and started debating team sizes (since I know google often doesn’t have giant teams), ended up somewhere in the hundreds of thousands, and rounded to thousands. But I made it all caps so you know it’s the SERIOUS kind of thousands. :P
- adamrt 8y agoWe moved to a monorepo about 2 years ago and it has been nothing but success for us. We have quite a few projects but only 4 major applications. Maybe it is that a few of our projects intertwine a bit so making spanning changes in separate repositories was a pain. Doing separate PRs, etc. Now changes are more atomic. Our entire infrastructure can be brought up in development with a single docker-compose file and all development apps are communicating with each other. I don't think we've had any issues that I can recall. We are a reasonably small team though, so maybe that is part of it.
- hobls 8y agoA single team is really helpful. Where I’ve seen it get particularly unhelpful is with multiple teams. I’m also not opposed to the concept, I just think it requires work to do correctly.
- bedros 8y agohow do you create branches in mono repo? for example I want to use branch rev5 from project A and rev3 from project B how I do that in a mono repo, I could not do it in HG, but sure about GIT
- Tloewald 8y agoCan’t you create a branch and merge the two branches you are interested in into that?
- bedros 8y agomy understanding, if you branch, you branch the entire repo, (not sure about some special case extensions ) if you have two projects stored in a single repo, you are forced to use whatever rev at for each project at a point of branch rev 5543 for example
- ScottBurson 8y agoIn Perforce, which is more or less what Google is using, you can branch any directory within the repo. (You would never branch the whole repo; that makes no sense.) So if you wanted to construct a directory with one version of one subdirectory, and a different version of another subdirectory, that's quite straightforward.
- mnm1 8y agoSlapping a whole bunch of projects into multiple repos with dependencies isn't a pleasant experience either. What is the solution then? I certainly don't want to host my own npm/composer/maven/clojars repos or even use those dependency managers to manage my own code which constantly changes and relies on multiple libraries both on the backend and frontend. I've tried this and, at least with a small team of two, it's not a pleasant experience at all. So how can I solve this problem? Cause the monorepo is very enticing after dealing with multiple repos and multiple dependencies pulled through dependency managers that clearly do not do well with dependencies that are constantly in flux.
- shakna 8y agoSubmodules? [0] Easy to use, cutting edge updates. [0] https://git-scm.com/book/en/v2/Git-Tools-Submodules https://git-scm.com/book/en/v2/Git-Tools-Submodules [1] https://www.mercurial-scm.org/wiki/Subrepository https://www.mercurial-scm.org/wiki/Subrepository
- glandium 8y agoSubmodules are cutting edge and have cutting edges. The user experience on some corner cases can be painful. Example: if you happen to have unrelated conflicts when you rebase some patch across a submodule update, you're most likely going to end up committing a reversal of the submodule update.
- flukus 8y ago> and the thing they all seem to overlook is that Google has THOUSANDS of hours of effort put into the tooling for their monorepo That's a huge understatement. They haven't just slapped a few scripts on top of git/svn, they've created their own proprietary scm to manage all of this. They've thrown more at this beast than most companies will throw at their actual product. I'm also not convinced they haven't reinvented individual repositories inside this monorepo, it sounds like you can create "branches" of just your code and share them with other people with committing to the trunk, this is essentially an individual repository that will be auto deployed when you merge to master.
- joshuamorton 8y agoYour last paragraph doesn't sound like anything at Google. Most engineers will never use branches at all, and even fewer will use branches that merge into trunk (instead of away from it).
- flukus 8y agoThere is a set of code changes locally and those changes are bundled off to the test server to run the full test suite? That's a branch. Now let's say I break a project sharing this code and because I'm not an expert in all 2 billion LoC and 3000 projects google is running I need to enlist some help in fixing what I broke. Presumably there is a way for the developers on that downstream project to pull in my change set? That's a shared branch. Now assuming I can get all of these planets aligned correctly I'm going to need to take this set of changes and put it into the master version aren't I? That's merging my branch into trunk.
- ebikelaw 8y agoYour mental model of how this works within Google is completely foreign to me. I think you've made an unfounded assumption somewhere.
- joshuamorton 8y agoYeah that second thing doesn't exist. That first thing doesn't really exist the way you conceptualize either, I don't think.
- maxpert 8y agoSpot on! I've seen org wide mono repos at Microsoft and they had their custom tooling and build systems built on top of SourceDepot.
- robaato 8y agoWhich is just rebadged Perforce :)
- ma2rten 8y agoGoogle used Perforce for a very long time before they built their own version control system.
- shados 8y agoBoth monorepo or "micro repo" end up falling apart at scale without some devops work involved. Either will work if you only have a few dozen projects. Neither will work once you hit 10s of millions of lines of code. But people seem to forget that it wasn't that long ago that git didn't exist, making multiple repos was a pain in the butt. Managing multiple repos locally was hell. Monorepos were the norm. Then as the state of version control ramped up, and making repos became easy, and having so much code in one repo had performance issues (overnight CVS/SourceSafe/SVN pull on your first day at work anyone? Branches that take hours to create?), people started making repos per project. The micro-service fad made that a no-brainer. Now, for companies like Facebook and Google, or really any company that wrote code before the modern days and has a non-trivial amount of it, switching was not exactly a simple matter. So they just poured their energy into making the monorepo work. They're not the only ones to do it either (though not everyone has to do it at Google, Facebook or Microsoft scale, obviously, so its a bit easier for most). And so it works. And then people forget how to make distributed repos work and claim things like "omg I have to make 1 PR per repo when making breaking changes!", as if it was a big deal or it wasn't a solved problem.
- hobls 8y agoAbsolutely! At some point you must invest in your tools. (Early, in my opinion.) I think the clarification I’d offer is that in the age of GitHub the “standard” model is multiple repos, so you’re actually giving up some tooling if you just shove everything in a single repo. (I’m also not sure I’d generally categorize tools work as as “dev ops,” though I can certainly see how they end up intertwined.)
- lclarkmichalek 8y agoI've seen hardly any tools to manage dependencies across multiple repos. Modifying multiple repos at the same time isn't an issue I see many resources devoted to, and managing those cross repo versions is almost never done well. In comparison, both buck and bazel offer pretty mature monorepo management tooling. On the VCS front, you can take native git/HG a long way.
- fcarraldo 8y ago
- georgewfraser 8y agoI kinda disagree, we’re a dev team of 30, 3.5 years in, 150k lines of code and we’ve always had a monorepo. We had to switch from maven to bazel after about 2 years because test times got out of control; bazel has been about 50% more annoying than maven but the incremental builds work perfectly.
- timkrueger 8y agoInteresting. Do you have wrote something about that migration?
- georgewfraser 8y agoNo, though that would be a good blog post. We tried to make multi-module maven work for a while, eventually gave up and wrote some scripts that would convert maven to bazel, using many assumptions that applied only to our particular case. We did the cutover in one day but kept maven around for a couple weeks in case we decided to bail on bazel. It worked out; we even found CircleCI works great. I would say the weak link in the bazel ecosystem is the IntelliJ plugin, which is very functional but also very slow.
- wirrbel 8y agoSame line of thinking, just different conclusions. I feel terrible for anyone trying to run a company with open-source style independent repos. On a popular github project, you have MANY potential contributors that will tell you if a PR, or a release candidate break API compatibility, etc. There are thousands of hours in open source dedicated to fixing integration issues due to the (unavoidable) poly-repo situation. Monorepos in companies are relatively simple. You need to dedicate some effort in your CI and CD infrastructure, but you'll win magnitudes by avoiding integration issues. Enough tooling is out there already to make it easy on you. Monorepos' biggest problem in an org is the funding, as integration topics are often deprioritized by management, and "we spend 10k per year on monorepo engineering" for some reason is a tough sell for orgs, who seem to prefer to "spend 5k for each of the 5 teams so that they maintain their own CD ways and struggle integrating which incurrs another 20k that just is not explicitly labeled as such". Developer team dynamics also play a role. I have observed the pattern now multiple times (N=3): * Developers have a monolithic repo, that has accumulated a few odd corners over time. * The feeling builds up that this monolithic repo needs to be modularized. * It is split up into libraries (or microservices), this is kind of painful, but feels liberating at first (now finally John does not break my builds anymore) * Folks realize: John doesn't break my builds anymore, but now I need to wait for integration on the test system to learn if he broke my code, and sometimes I only learn it in production. * people start posting blog posts on monorepos That pattern takes 2-3 years to play out, but I have seen it on every job I worked.
- marmaduke 8y agoworking as dev with academic teams, I usually use many repos for "damage control" as git-ignorant scientists will dump irrelevant files into a repo. with that in mind, is monorepo is a universally good approach or is more dependent on good behavior of team members than polyrepo?
- jakoblorz 8y agoI don't think that there is the one size fits all solution especially if you can't expect basic knowledge about git
- 8y ago
- baybal2 8y agoDumping all code in a single repo, even for a 30 man development shop was really tough. Doing so for a company of few thousands must be truly crazy. I advice Google to replace the person in their internal IT who came up with that idea.
- foota 8y agoReally probably closer to millions of hours.
- 013a 8y agoI call this "Google Imposter Syndrome". Because Google (insert Facebook, Apple, Amazon, etc) has success with Monorepos (insert gRPC, Go, Kubernetes, React/Native, etc), it must be a great idea, we should do it. You see this everywhere. Also known as an Appeal to Authority. My personal opinion: very few companies will hit a point where sheer volume of code or code changes makes a monorepo unwieldy. Code volume is a Google-problem. But every company will have problems with Github/Gitlab/whatever tooling with multiple repos; coordinating merges/deploys across multiple projects, managing issues, context switches between them, etc. And every company will also have problems with CI/CD in a monorepo. Point being... there are problems with both, and there are benefits to both. I don't think one is right or wrong. I personally feel that solving the problems inherent to monorepos, at average scale, is easier than solving the problems inherent to distributed repos. The monorepo problems are generally internal technical, whereas the distributed repo problems are generally people-related and tooling outside of your control.
- rubenbe 8y agoI've seen multiple companies struggling with maintaining interdependencies between multiple repos. It often results in an expensive custom solution. As a general guideline I'd say "when in doubt, put the code in a single repo"
- kqr 8y agoSomeone at some point said "Google may not be successful for the interview practises they use; they're big enough that they could very well be successful despite the interview practises they use." It stuck with me, and is applicable to so many things. Including, maybe, this?
- kungtotte 8y agoAnother question is just the sheer scale of the FAANG companies, making things work at that scale is likely to be counterintuitive sometimes. I just looked it up, Facebook has 2.2 billion users monthly. That's almost a third of the entire planet. Shit that makes sense for them won't make sense for 99% of everyone else.
- 8y ago
- w_t_payne 8y agoTooling is required for coordinating configuration management on multiple repositories too. Also, why isn't such tooling available as open source? I'm trying to do my bit, but we could do with more effort being put into this, somehow.
- pavbelshippable 8y agoMaybe relatable, Even we believe mono repos are the right choice for teams that want to ship code faster. There are concerns that this doesn't scale well, but these are largely unfounded. Companies like Twitter, Google, Facebook run massive monolithic repos with 1000s of developers. With mono repos you will have, > Better developer testing: Developers can easily run the entire platform on their machine and this helps them understand all services and how they work together. This has led our developers to find more bugs locally before even sending a pull request. > Reduced code complexity: Senior engineers can easily enforce standardization across all services since it is easy to keep track of pull requests and changes happening across the repository. > Effective code reviews: Most developers now understand the end to end platform leading to more bugs being identified and fixed at the code review stage. > Sharing of common components: Developers have a view of what is happening across all services and can effectively carve out common components. Over a few weeks, we actually found that the code for each microservice became smaller, as a lot of common functionality was identified and shared across services. > Easy refactoring: Any time we want to rename something, refactoring is as simple as running a grep command. Restructuring is also easier as everything is neatly in one place and easier to understand. The results? Our productivity has increased at least 5x. The overall experience we have written over here http://blog.shippable.com/our-journey-to-microservices-and-a-mono-repository http://blog.shippable.com/our-journey-to-microservices-and-a...
- EnderMB 8y agoI've seen this a few times in the .NET world, mainly as a carry-over from Subversion when we had moved to Mercurial and git. Some mad genius in a company will write a fuck-ton of helper classes and utilities that take the heavy lifting out of everything remotely hard, to the point where you almost never need to touch a third-party API for a CMS, email send service, or cloud-hosting provider. Instead of supplying these as private NuGet packages to be installed into an application, they sit in solutions in their entirety, in case they are needed. That application then goes to a new developer team, and they have zero idea why there are millions of lines of code and dozens of projects for a basic website that doesn't really seem to do anything. It's a nice idea, but it has resulted in some very tightly coupled applications. I remember one time where a new developer changed some code in one of the utilities that handled multi-language support, and for some reason our logs reported that the emails were broke.
- keerthiko 8y agoWhat I can advise against is repo partitioning prematurely. I have been on multiple teams that have thought "Oh this will be a common library for all our projects" or "this is a sample project" or "this is the android version and this is the iOS version" and split projects up into different repos, only to wind up with crazy dependencies between repos which have fallen out of sync or require another repo to be on a specific branch/hash to work correctly, causing all kinds of chaos. Split your repos by dependencies, and once your system architecture is kind of fleshed out. Just use branches on the same repo until then.