16 ms·
Monorepo – Our Experience
- deleted 2y ago[deleted]
- siva7 2y agoOk, but the more interesting part - how did you solve the CI/CD part and how does it compare to a multirepo?
- CharlieDigital 2y agoMost CI/CD platforms will allow specification of targeted triggers. For example, in GitHub[0]: name: ".NET - PR Unit Test" on: ## Only execute these unit tests when a file in this directory changes. pull_request: branches: [main] paths: [src/services/publishing/**.cs, src/tests/unit/**.cs] So we set up different workflows that kick off based on the sets of files that change. [0] https://docs.github.com/en/actions/writing-workflows/workflow-syntax-for-github-actions#onpushpull_requestpull_request_targetpathspaths-ignore https://docs.github.com/en/actions/writing-workflows/workflo...
- victorNicollet 2y agoI'm not familiar with GitHub Actions, but we reverted our migration to Bitbucket Pipelines because of a nasty side-effect of conditional execution: if a commit triggers test suite T1 but not T2, and T1 is successful, Bitbucket displays that commit with a green "everything is fine" check mark, regardless of the status of T2 on any ancestors of that commit. That is, the green check mark means "the changes in this commit did not break anything that was not already broken", as opposed to the more useful "the repository, as of this commit, passes all tests".
- ants_everywhere 2y agoisn't that generally what you want? the check mark tells you the commit didn't break anything. if something was already broken it should have either blocked the commit that broke it or there's a flake somewhere that you can only locate by periodically running tests independent of any PR activity.
- daelon 2y agoIs it a side effect if it's also the primary effect?
- plorkyeran 2y agoI would find it extremely confusing and unhelpful if tests in the parent commit which weren't rerun for a PR because nothing relevant was touched marked the PR as red. Why would you even want that? That's not something which is relevant to evaluating the PR and would make you get in the habit of ignoring failures. If you split something into multiple repositories then surely you wouldn't mark PRs on one of them as red just because tests are failing in a different one?
- victorNicollet 2y agoI suppose our development process is a bit unusual. The meaning we give to "the commit is green" is not "this PR can be merged" but "this can be deployed to production", and it is used for the purpose of selecting a release candidate several times a week. It is a statement about the entire state of the project as of that commit, rather than just the changes introduced in that commit. I can understand the frustration of creating a PR from a red commit on the main branch, and having that PR be red as well as a result. I can't say this has happened very often, though: red commits on the main branch are very rare, and new branches tend to be started right after a deployment, so it's overwhelmingly likely that the PR will be rooted at a green commit. When it does happen, the time it takes to push a fix (or a revert) to the main branch is usually much shorter than the time for a review of the PR, which means it is possible to rebase the PR on top of a green commit as part of the normal PR acceptance timeline.
- plorkyeran 2y agoGoing off the PR status to determine if the end result is deployable is not reliable. A non-FF merge can have both the base commit and the PR be green but the merged result fail. You need to run your full test suite on the merged result at some point before deployment; either via a commit queue or post-merge testing.
- victorNicollet 2y agoI agree ! We use the commit status instead of the PR status. A non-FF merge commit, being a commit, would have its own status separate from the status of its parents.
- hk1337 2y agoEven AWS CodeBuild (or CodePipeline) allows you to do this now. It didn't before but it's a fairly recent update.
- CharlieDigital 2y agoAs a prior user of AWS Code*, I can appreciate that you qualified that with "Even" LMAO
- devjab 2y agoI don’t think CI/CD should really be a big worry as far as mono-repositories go as you can setup different pipelines and different flows with different configurations. Something you’re probably already doing if you have multiple repos. In my experience the article is right when it tells you there isn’t that big of a difference. We have all sorts of repositories, some of which are basically mono-repositories for their business domain. We tend to separate where it “makes sense” which for us means that it’s when what we put into repositories is completely separate from everything else. We used to have a lot of micro-repositories and it wasn’t that different to be honest. We grouped more of them together to make it easier for us to be DORA compliant in terms of the bureaucracy it adds to your documentation burden. Technically I hardly notice.
- JamesSwift 2y agoIn my limited-but-not-nothing experience working with mono vs multi repo of the same projects, CI/CD definitely was one of the harder pieces to solve. Its highly dependent on your frameworks and CI provider on just how straightforward it is going to be, and most of them are "not very straightforward". The basic way most work is to run full CI on every change. This quickly becomes a huge speedbump to deployment velocity until a solution for "only run what is affected" is found.
- bluGill 2y agoThe problem with "only run what is affected" is it is really easy to have something that is affected but doesn't seem like it should be (that is whatever tools you have to detect is it affected say it isn't). So if you have such a system you must have regular rebuild everything jobs as well to verify you didn't break something unexpected. I'm not against only run what is affected, it is a good answer. It just has failings that you need to be aware of.
- JamesSwift 2y agoYeah thats a good point. Especially for an overly-dynamic runtime like ruby/rails, theres just not usually a clean way to cordon off sections of code. On the other hand, using nx in an angular project was pretty amazing.
- victorNicollet 2y agoWouldn't CI be easier with a monorepo ? Testing integration across multiple repositories (triggered by changes in any of them) seems more complex than just adding another test suite to a single repo.
- bluGill 2y agoPros and cons. Both can be used successfully, but there are different problems to each. If you have a large project you will have a tool teams to deal with the problems of your solution.
- habosa 2y agoI think this is one of the big reasons so many companies end up using Bazel as part of scaling their monorepos. It has many faults but one thing it does perfectly is “build and test everything affected by this commit”
- IshKebab 2y agoYou use a build system that sandboxes dependencies (Bazel, Buck2, Please, etc.) so that you know there are no undeclared dependencies. Then you can simply query the build system "what could have been affected by this change" and only build/test those things.
- xyzzy_plugh 2y agoWithout indicating my personal feelings on monorepo vs polyrepo, or expressing any thoughts about the experience shared here, I would like to point out that open-source projects have different and sometimes conflicting needs compared to proprietary closed-source projects. The best solution for one is sometimes the extreme opposite for the other. In particular many build pipelines involving private sources or artifacts become drastically more complicated than their those of publicly available counterparts.
- bunderbunder 2y agoI've also seen this with branching strategies. IMO the best branching strategy for open source projects is generally the worst one for commercial projects, and vice versa.
- adastra22 2y agoWhich strategies are you assigning to each?
- b5hi 2y agothis should be the top comment
- magicalhippo 2y agoWe're transitioning from a SVN monorepo to Git. We've considered doing a kind of best-of-both-worlds approach. Some core stuff into separate libraries, consumed as nuget packages by other projects. Those libraries and other standalone projects in separate repos. Then a "monorepo" for our main product, where individual projects for integrations etc will reference non-nuget libraries directly. That is, tightly coupled code goes into the monorepo, the rest in separate repos. Haven't taken the plunge just yet tho, so not sure how well it'll actually work out.
- dezgeg 2y agoIn my experience this turns to nightmare when (not if, when) there is need to make changes to the libraries and app at the same time. Especially with libraries it's often necessary to create a client for an API at the same time to really know that the interface is any good.
- magicalhippo 2y agoThe idea is that the libraries we put in nuget are really non-project-specific. We'll use nuget to manage library versions rather than git submodules, so hopefully they can live fine in a separate repo. So updating them at the same time shouldn't be a huge deal, we just make the change in the library, publish the nuget package, and then bump the version number in the downstream projects that need the change. Ideally changes to these libraries should be relatively limited. For things that are intertwined, like an API client alongside the API provider and more project-specific libraries, we'll keep those together in the same repo. If this is what you're thinking of, I'd be interested in hearing more about your negative experiences with such a setup.
- adastra22 2y ago...just curious, what's your line of work where you're still using SVN?
- magicalhippo 2y agoCompany makes B2B software. Our main product is a traditional Win32 application with over 25 year old code in production, though most of it is much more recent. SVN worked well for us, being a relatively small team until recently, so no real need to change until now. The primary driver for the move to Git has been compliance. We just can't have on-prem servers with critical things like code anymore, and there's effectively just one cloud-based SVN offering that had ISO27001 etc. So incentive to move to Git just got a lot stronger.
- CharlieDigital 2y ago> Moving to a monorepo didn't change much, and what minor changes it made have been positive. I'm not sure that this statement in the summary jives with this statement from the next section: > In the previous, separate repository world, this would've been four separate pull requests in four separate repositories, and with comments linking them together for posterity. > > Now, it is a single one. Easy to review, easy to merge, easy to revert. IMO, this is a huge quality of life improvement and prevents a lot of mistakes from not having the right revision synced down across different repos. This alone is a HUGE improvement where a dev doesn't accidentally end up with one repo in this branch and forgot to pull this other repo at the same branch and get weird issues due to this basic hassle. When I've encountered this, we've had to use another repo to keep scripts that managed this. But this was also sometimes problematic because each developer's setup had to be identical on their local file system (for the script to work) or we had to each create a config file pointing to where each repo lived. This also impacts tracking down bugs and regression analysis; this is much easier to manage in a mono-repo setup because you can get everything at the same revision instead of managing synchronization of multiple repos to figure out where something broke.
- notwhereyouare 2y agoironically was gonna come and comment on that same second block of text. We went from monorepo to multi-repo at work and it's been a huge set back and disappointment with the devs because it's what our contractors recommended. I've asked for a code deploy and everything and it's failed in prod due to a missing check in
- CharlieDigital 2y ago> ...because it's what our contractors recommended It's sad when this happens instead of taking input from the team on how to actually improve productivity/quality. A startup I joined started with a multi-repo because the senior team came from a FAANG where this was common practice to have multiple services and a repo for each service. Problem was that it was a startup with one team of 6 devs and each of the pieces was connected by REST APIs. So now any change to one service required deploying that service and pulling down the OpenAPI spec to regenerate client bindings. It was so clumsy and easy to make simple mistakes. I refactored the whole thing in one weekend into a monorepo , collapsed the handful of services into one service, and we never looked back. That refactoring and a later paper out of Google actually inspired me to write this article as a practical guide to building a "modular monolith": https://chrlschn.dev/blog/2024/01/a-practical-guide-to-modular-monoliths/ https://chrlschn.dev/blog/2024/01/a-practical-guide-to-modul...
- memsom 2y agomonorepos are appropriate for a single project with many sub parts but one or two artifacts on any given release build. But they fall apart when you have multiple products in the monorepo, each with different release schedules. As soon as you add a second separate product that uses a different subset of any code in the repo, you should consider breaking up the monorepo. If the code is "a bunch of libraries" and "one or more end user products" it becomes even more imperative to consider breaking down stuff.. Having worked on monorepos where there are 30+ artifacts, multiple ongoing projects that each pull the monorepo in to different incompatible versions, and all of which have their own lifetime and their own release cycle - monorepo is the antithesis of a good idea.
- munksbeer 2y agoNo offense but I think you're doing monorepos wrong. We have more than 100 applications living in our monorepo. They share common core code, some common signals, common utility libs, and all of them share the same build. We release everything weekly, and some things much more frequently. If your testing is good enough, I don't see what the issue is?
- bluGill 2y ago> If your testing is good enough, I don't see what the issue is? Your testing isn't good enough. I don't know who you are, what you are working on, or how much testing you do, but I will state with confidence it isn't good enough. It might be acceptable for your current needs, but you will have bugs that escape testing - often intentional as you can't stop forever to fix all known bugs. In turn that means if anything changes in your current needs you will run into issues. > We release everything weekly, and some things much more frequently. This is a negative to users. When you think you will release again next so who cares about bugs it means your users see more bugs. Sure it is nice that you don't have to break open years old code anymore, but if the new stuff doesn't have anything the user wants is this really a good thing?
- munksbeer 2y ago> Your testing isn't good enough. I don't know who you are, what you are working on, or how much testing you do, but I will state with confidence it isn't good enough. Yes, that is true. No amount of testing can prevent bugs in a complex enough project. But this is no different in a monorepo or multirepo. I apologise but I don't think I've understood your point.
- h1fra 2y agoI think the big issue around monorepo is when a company puts completely different projects together inside a single repo. In this article almost everything makes sense to me (because that's what I have been doing most of my career) but they put their OTP app inside which suddenly makes no sense. And you can see the problem in the CI they have dedicated files just for this App and probably very few common code with the rest. IMO you should have one monorepo per project (api, frontend, backend, mobile, etc. as long as it's the same project) and if needed a dedicated repo for a shared library.
- fragmede 2y ago> you should have one monorepo per project (api, frontend, backend, mobile, etc. as long as it's the same project) that's not a monorepo! Unless the singular "project" is stuff our company ships, the problem you have is of impedance mismatch between the projects, which is the problem that an actual monorepo solves. for swe's on individual projects who will never have the problem of having to ship a commit on all the repos at the "same" time, yeah that seems fine, and for them it is. the problem comes as a distributed systems engineer where, for whatever reason, many or all the repos need to be shipped at the ~same time. or worse - A needs to ship before B which needs ship before C but that needs to ship before A, and you have to unwind that before actually being able to ship the change.
- hk1337 2y ago> that's not a monorepo! Sure it is! It's just not the ideal use case for a monorepo which is why people say they don't like monorepos.
- vander_elst 2y ago"one monorepo per project (api, frontend, backend, mobile, etc. as long as it's the same project) and if needed a dedicated repo for a shared library." They are literally saying that multiple repos should be used, also for sharing the code, this is not monorepo, these are different repos.
- 2y ago
- gregmac 2y agoTo me, monorepo vs multi-repo is not about the code organization, but about the deployment strategy. My rule is that there should be a 1:1 relation between a repository and a release/deployment. If you do one big monolithic deploy, one big monorepo is ideal. (Also, to be clear, this is separate from microservice vs monolithic app: your monolithic deploy can be made up of as many different applications/services/lambdas/databases as makes sense). You don't have to worry about cross-compatibility between parts of your code, because there's never a state where you can deploy something incompatible, because it all deploys at once. A single PR makes all the changes in one shot. The other rule I have is that if you want to have individual repos with individual deployments, they must be both forward- and backwards-compatible for long enough that you never need to do a coordinated deploy (deploying two at once, where everything is broken in between). If you have to do coordinated deploys, you really have a monolith that's just masquerading as something more sophisticated, and you've given up the biggest benefits of both models (simplicity of mono, independence of multi). Consider what happens with a monorepo with parts of it being deployed individually. You can't checkout any specific commit and mirror what's in production. You could make multiple copies of the repo, checkout a different commit on each one, then try to keep in mind which part of which commit is where -- but this is utterly confusing. If you have 5 deployments, you now have 4 copies of any given line of code on your system that are potentially wrong. It becomes very hard to not accidentally break compatibility. TL;DR: Figure out your deployment strategy, then make your repository structure mirror that.
- aswerty 2y agoThis mirrors my own experience in the SaaS world. Anytime things move towards multiple artifacts/pipelines in one repo; trying to understand what change existed where and when seems to always become very difficult. Of course the multirepo approach means you do this dance a lot more: - Create a change with backwards compatibility and tombstones (e.g. logs for when backward compatibility is used) - Update upstream systems to the new change - Remove backwards compatibility and pray you don't have a low frequency upstream service interaction you didn't know about While the dance can be a pain - it does follow a more iterative approach with reduced blast radiuses (albeit many more of them). But, all in all, an acceptable tradeoff. Maybe if I had more familiarity in mature tooling around monorepos I might be more interested in them. But alas not a bridge I have crossed, or am pushed to do so just at the moment.
- syndicatedjelly 2y agoSome thoughts: 1) Comparing a photo storage app to the Linux kernel doesn't make much sense. Just because a much bigger project in an entirely different (and more complex) domain uses monorepos, doesn't mean you should too. 2) What the hell is a monorepo? I feel dumb for asking the question, and I feel like I missed the boat on understanding it, because no one defines it anymore. Yet I feel like every mention of monorepo is highly dependent on the context the word is used in. Does it just mean a single version-controlled repository of code? 3) Can these issues with sync'ing repos be solved with better use of `git submodule`? It seems to be designed exactly for this purpose. The author says "submodules are irritating" a couple times, but doesn't explain what exactly is wrong with them. They seem like a great solution to me, but I also only recently started using them in a side project
- datadrivenangel 2y agoMonorepo is just a single repo. Yup. Git submodules have some places where you can surprisingly lose branches/stashed changes.
- syndicatedjelly 2y agoOne of my repos has a dependency on another repo (that I also own). I initialized it as a git submodule (e.g. my_org/repo1 has a submodule of my_org/repo2). Git submodules have some places where you can surprisingly lose branches/stashed changes. This concerns me, as git generally behaves as a leak-proof abstraction in my experience. Can you elaborate or share where I can learn more about this issue?
- datadrivenangel 2y agoFrom the git-scm book: "The other main caveat that many people run into involves switching from subdirectories to submodules. If you’ve been tracking files in your project and you want to move them out into a submodule, you must be careful or Git will get angry at you. " Though apparently newer versions of git are better about not losing submodule branches, so my concerns were outdated.
- mgaunard 2y agoDoing modular right is harder than doing monolithic right. But if you do it right, the advantage you get is that you get to pick which versions of your dependencies you use; while quite often you just want to use the latest, being able to pin is also very useful.
- lukewink 2y agoYou can still publish packages and pull them down as (pinned) dependencies all within a monorepo.
- mgaunard 2y agothat's a terrible and arguably broken-by-design workflow which entirely defeats the point of the monorepo, which is to have a unified build of everything together, rather than building things piecemeal in ways that could be incompatible. For C++ in particular, you need to express your dependencies in terms of source versions, and ensure all of the build artifacts you link together were built against the same source version of every transitive dependency and with the same flags. Failure to do that results in undefined behaviour, and indeed I have seen large organizations with unreliable builds as a manner of routine because of that. The best way to achieve that is to just build the whole thing from source, with a content-addressable-store shared with the whole organization to transparently avoid building redundant things. Whether your source is in a single repo or spread over several doesn't matter so long as your tooling manages that for you and knows where to get things, but ultimately the right way to do modular is simply to synthesize the equivalent monorepo and build that. Sometimes there is the requirement that specific sources should have restricted access, which is often a reason why people avoid building from source, but that's easy to work around by building on remote agents. Now for some reason there is no good open-source build system for C++, while Rust mostly got it right on the first try. Maybe it's because there are some C++ users still attached to the notion of manually managing ABI.
- move-on-by 2y agoThis is something I’ve been preaching at work and it just falls on deaf ears. We aren’t even doing a great job with monolithic, why people think modular is going to _improve_ things is beyond my comprehension. I’ve pretty much decided no one actually cares and they just see it as an opportunity for résumé building. Someone want to enlighten me?
- stackskipton 2y agoAs DevOps/SRE type person that occasionally gets stuck with builds, Monorepos world well if company will invest in the build process. However, many companies don't do well in this area and Monorepo blast radius becomes much bigger so individual repos it is. Also, depending on the language, building private repo is easy enough to keep all common libraries in.
- stillbourne 2y agoI like to use the monorepo tools without the monorepo repo. If that makes any god damn sense. I use NX at my job and the monorepo was getting out of hand, 6 hour pipeline builds, 2 hours testing, etc. So I broke the repo into smaller pieces. This wouldn't have been possible if I wasn't already using the monorepo tools universally through the project but it ended up working well.
- KaiserPro 2y agoMonorepos have their advantages, as pointed out, one place to review, one place to merge. But it can also breed instability, as you can upgrade other people's stuff without them being aware. There are ways around this, which involve having a local module store, and building with named versions. Very similar to a bunch of disparate repos, but without getting lost in github (github's discoverability was always far inferior to gitlab) However it has its draw backs namely that people can hold out on older versions than you want to support.
- dkarl 2y ago> But it can also breed instability, as you can upgrade other people's stuff without them being aware This is why Google embraced the principle that if somebody breaks your code without breaking your tests, it's your fault for not writing better tests. (This is sometimes known as the Beyonce rule: if you liked it, you should have put a test on it.) You need the ability to upgrade dependencies in a hands-off way even if you don't have a monorepo, though, because you need to be able to apply security updates without scheduling dev work every time. You shouldn't need a careful informed eye to tell if upgrades broke your code. You should be able to trust your tests.
- KaiserPro 2y agoI mean yes, that is an approach. However that only really works if people are willing to put effort into tests. It also surmises that people are able to accuratly and simply navigate dependencies. The issue is that monorepos make it trivial to add dependencies, to the point where if I use a library to get access to our S3-like object storage system, it ends up pulling a massive chain of deps culminating in building caffe binaries (yes, as in the ML framework.) I cannot possibly verify that massive dependency chain, so putting a test on some part which fails is an exercise in madness. It requires a culture of engineering discipline that I have yet to see at scale.
- msoad 2y agoI love monorepos but I'm not sure if Git is the right tool beyond certain scale. Where I work doing a simple `git status` takes seconds due to the size of the repo. There has been various attempts to solve Git performance but so far this is nothing close to what I experienced at Google. The Git team should really invest in tooling for very large repos. Our repo is around 10M files and 100M lines of code and no amount of hacks on top of Git (cache, sparse checkout etc etc) is not really solving the core problem. Meta and Google have really solved this problem internally but there is no real open source solution that works for everyone out there.
- dijit 2y agoI’m secretly hoping that google releases piper (and Mondrian); the gaming industry would go wild. Perforce is pretty brutal, and the code review tools are awful - but its still the undisputed king of mixed text and binary assets in a huge monorepo.
- habosa 2y agoMondrian! That’s a name I haven’t heard in a while. Google uses Critique for code review these days. I tried to bring the best of Critique to GitHub with https://codeapprove.com https://codeapprove.com but you’re right there’s a lot that just doesn’t work on top of git.
- ralph84 2y agoThere were rumors that at one point piper included source code licensed from Perforce. That could make open sourcing it more difficult if any of that code is still hanging around.
- vlovich123 2y agoMeta open-sourced their complete stack: https://github.com/facebook/sapling https://github.com/facebook/sapling Microsoft released Scalar (https://github.com/microsoft/scalar https://github.com/microsoft/scalar) although it's not a complete stack yet but it is already planning on releasing the backend components eventually. Have you tried Sapling? It has EdenFS baked in so it'll only materialize the files you touch and operations are fast because it has a filesystem watcher for activity so it doesn't need to do a lot of work to maintain a view of what has been invalidated.
- paxys 2y agoAll the pitfalls of a monorepo can disappear with some good tooling and regular maintenance, so much so that devs may not even realize that they are using one. The actual meat of the discussion is – should you deploy the entire monorepo as one unit or as multiple (micro)services?
- marcosdumay 2y agoThat's the thing. All the pitfalls of multi-repos also disappear with good tooling and regular maintenance. Neither one has an actual edge. Yet you can find countless articles from people talking about their experience. Take those as a hint about what kind of tooling you need, not about their comparative qualities.
- bobim 2y agoStarted to use a monorepo + worktrees to keep related but separated developments all together with different checkouts. Anybody else on the same path?
- __MatrixMan__ 2y agoEvery monorepo I've ever met (n=3) has some kind of radioactive DMZ that everybody is afraid to touch because it's not clear who owns it but it is clear from its quality that you don't want to be the last person who touched it because then maybe somebody will think that you own it. It's usually called "core" or somesuch. Separate repos for each team means that when two teams own components that need to interact, they have to expose a "public" interface to the other team--which is the kind of disciplined engineering work that we should be striving for. The monorepo-alternative is that you solve it in the DMZ where it feels less like engineering and more like some kind of multiparty political endeavor where PR reviewers of dubious stakeholder status are using the exercise to further agendas which are unrelated to the feature except that it somehow proves them right about whatever architectural point is recently contentious. Plus, it's always harder to remove something from the DMZ than to add it, so it's always growing and there's this sort of gravitational attractor which, eventually starts warping time such that PR's take longer to merge the closer they are to it. Better to just do the "hard" work of maintaining versioned interfaces with documented compatibility (backed by tests). You can always decide to collapse your codebase into a black hole later--but once you start on that path you may never escape.
- zaphar 2y agoSince we are indulging in generalizations from our past. With separate repos you end up with 10 "cores" that are radioctive DMZ's everybody is afraid to touch. And those "disciplined" public API's will be universally hated by everyone who consumes them. Neither a monorepo nor separate repos will result in people being disciplined. If you already have the discipline to do separate repositories correctly then you'll be fine with a monorepo. So I guess it's six on one hand, half dozen in the other.
- __MatrixMan__ 2y agoNo I think there's a clear difference. I've seen this several times: Somebody changes teams and now they're no longer responsible for a bit of code, but then they learn that it is broken in some way, and now they're sneaking in commits that--on paper--should now be handled by somebody else. Dev's *like* to feel ownership of reasonably sized chunks of code. We like to arrange it in ways that is pleasing for us to work on later down the road. And once we've made those investments, we like to see them pay off by making quick easy changes that make users happy. Sharing a small codebase with three or four other people and finding ways to make each other's lives easier while supporting it is *fun* and it makes for better code too. But it only stays fun if you have enough autonomy that you can really own it--you and your small team. Footguns introduced need to be pointed at your feet. Automation introduced needs to save you time. If you've got the preferences of 50 other people to consider, and you know that whatever you do you're going to piss off some 10 of them or another... the fun goes away. This is simple: > we own this whole repo and only this 10% of it (the public interface) needs to make external stakeholders happy, otherwise we just care about making each other happy. ...and it has no space in it for there to be any code which is not clearly owned by somebody. In a monorepo, there are plenty of places for that.
- alphazard 2y agoThe classic micro/multi repo mistake is reaching for more repos when you really need better tooling and permissions on single repo. People have probably wasted millions of engineer-hours across the industry with multiple repos, all because GitHub doesn't have expressive path-level permissions.
- gorgoiler 2y agoRepository boundaries are affected far more by the social structure of your organisation than anything technical. Do you want hard boundaries between teams — clear responsibilities with formal ceremony across boundaries, but at the expense of living with inflexibility? Do you want fluidity in engineering, without fixed silos and a flat org structure that encourages anyone to take on anything that’s important to the business right now, but with the much bigger overhead of needing strong people leaders capable of herding the chaos? I’m sure there are dozens of other examples of org structures and how they are reflected in code layout, repo layout, shared directories, dropboxes, chat channels, and email groups etc.
- drbojingle 2y agoa lot of comments here seem to think that mono-repo has to mean something about deployment. I just don't want to have to run git fetch and 5 different repos to get everything I need and that's good enough reason for me to use one.
- klabb3 2y agoNot to mention when you need to make cross repo changes in development, and you have to set up a whole web of local repointing in package manifests. Repo is a hard boundary. Sometimes you need one. But to create a boundary when you don’t have to, on things that are deeply interconnected and have nowhere near a stable api? Utter madness imo.
- __MatrixMan__ 2y agoPresumably your language has a package manager which can do that for you? (or you could distribute it as an image). I guess you do have to decide which version you should depend on, but putting that power in your hands is sort of the point.
- bob1029 2y agoThe #1 benefit for me regarding the monorepo strategy is that when someone on the team refers to a commit hash, there is exactly one place to go and it provides a consistent point-in-time snapshot of everything. Ideally, all of the commits on master are ~good, so you have approximately a perfect time machine to work with. I have solved more bugs looking at diffs in GitHub than I have in my debugger simply by having everything in one happy scrolly view. Being able to flick my mouse wheel a few clicks and confirm that the schema does indeed align with the new DTO model props has saved me countless hours. Confirming stuff like this across multiple repos & commits can encourage a more lackadaisical approach. This also dramatically simplifies things like ORM migrations, especially if you require that all branches rebase & pass tests before merging. I agree with most of the hypothetical caveats, but if you can overcome them even with some mild degree of suffering, I don't see why you wouldn't fight for it.
- hu3 2y agoI've seen teams take this even further and vendor all dependencies. This way a commit hash contains even the exact third party code involved.
- yurishimo 2y agoThis is perfectly fine if your language of choice doesn’t have a robust package manager that supports version pinning. But then you need to enforce pinning across your org which could prove its own challenge.
- hu3 2y agoWhy only restrict to these cases? One of the teams vendored npm and go packages which are robust. Code always ran inside Docker. Being able to just clone and run simplified their flow.
- IshKebab 2y agoI don't see how that relates to monorepos. Even using submodules a single commit hash will specify the commit hashes of every submodule (usually; you can actually set up submodules to point to a branch instead of a commit but I've never seen anyone do this). There are definitely huge benefits to monorepos but I don't see how this is one.
- akoboldfrying 2y agoThis prompted a shower thought: Isn't N separate repos actually strictly worse than a monorepo with N completely independent long-lived branches, where each person checks out all the ones they need to work on under separate folders with `git worktree add`? I can think of only 2 ways that the multiple-branch monorepo is worse: 1. If the monorepo is large, everyone has to deal with a fat .git folder even if they have only checked out a branch with a few files. 2. Today, everyone expects different branches in a repo to contain "different versions of the same thing", not "a bunch of different things". But this is purely convention. The only real benefit that I can see of making a separate repo (over adding a new project directory to a "classic" monorepo) is the lower barrier to getting underway -- you can just immediately start doing whatever you want; the pain of syncing repos comes later. But this is also true when starting work under a new branch in the branch-per-project style monorepo: you can just create a branch from the initial commit, and away you go -- and if you need to atomically make changes across projects, just merge their branches first! What are the downsides I'm not seeing?
- adastra22 2y ago> a monorepo with N completely independent long-lived branches That's not a monorepo.
- akoboldfrying 2y agoIt's a single git repo. What would you call it?
- sunshowers 2y agoI'd consider continuous integration (people continuously merge their changes into main) to be one of the defining characteristics of a monorepo.
- adastra22 2y ago“Repo” here really means branch. A monorepo is where you put everything into a single branch.
- oslem 2y agoAlright everyone, we’ve trained for this. Grab your popcorn and get a good seat. It’s the return of the of great mono/poly repo debate.
- yen223 2y ago> Moving to a monorepo didn't change much, and what minor changes it made have been positive. It's pretty refreshing to see an experience report whose conclusion is "not much has changed", even though in practice that's the most common result for any kind of process change.
- dboreham 2y agoThe meta-syndrome here is: one person fixing a problem they have, and thereby making problems other people have worse. Often first person doesn't have a good awareness of the full melange of problems and participants.
- LunicLynx 2y agoThere is only one concept of a monorepo. And that is the google approach. This is a project repo and in a project repo things should stay together. Your tooling must be different for it to work. So using git for it will not have a positive result.
- ofrzeta 2y agoI am not convinced Git submodules are so bad. Obviously it's a bit more work than a monorepo but it actually works quite nice to have the parent repo pin the commits of the submodules. So you can just update, say, the "frontend" when fixing a bug, update the submodule and commit the hash to the parent. lgtm.
- Abismith 2y ago[dead]
- vekker 2y agoI like monorepos as a developer, but as a founder, monorepos have one massive downside: if you want to hire outside help, you have to share everything. While in some cases, the complete context is helpful for the job, in other cases, and I realize this may be pure paranoia but, you may not want to share the complete picture.
- IshKebab 2y agoIs that really an issue? Huge companies like Microsoft and Google use monorepos, and they hire tens of thousands of people and contractors all with access to the code. I think it's a natural fear but the reality is that a) most people don't leak source code, and b) access to source code isn't really that valuable. Most source code is too custom to be useful to most other people, and most competitors (outside China at least) wouldn't want to steal code anyway. Actually I did find this answer on how Google does it and apparently they do support some ACLs for directories in their monorepo. Microsoft uses Git though so I'm not sure what they do. https://www.quora.com/If-Google-has-1-big-monorepo-how-do-they-keep-certain-code-private-if-they-work-with-outside-contractors-etc https://www.quora.com/If-Google-has-1-big-monorepo-how-do-th...
- bob1029 2y ago> and b) access to source code isn't really that valuable This is a very important lesson. Once you learn that The Moat is more about the customers & trust, you stop worrying so much about every last possible security vector into your text files. Treating a repository like a SCIF will put a lot of friction on getting things done. If you simply refrain from placing production keys/certs/secrets in your source code, nothing bad will likely occur with a broad access policy. The chances that your business has source code with any intrinsic market value is close to zero. That is how much money you should spend on defending it.
- photonthug 2y agoMonorepos: if you don’t have it, everyone wants it, and if you do have it, no one likes it. There are solid, good faith arguments for each way. There are very real benefits to switching in some circumstances, but those reasons are themselves fragile and subject to frequent change based on cultural or operational shake ups. So in the end this is a yaml vs json type of argument mostly, and if you’re thinking about rioting over this there is a very good chance you could find a better hill to die on.
- dan-robertson 2y agoI have a monorepo and like it. I think at some size you get enough people maintaining the repo that it becomes good? I also never want yaml, fwiw.
- photonthug 2y ago> I also never want yaml, fwiw. Fair enough, it’s just that if you’re my coworker or employee I’m wondering if you don’t have something more important to worry about =D
- DrBazza 2y agoIt's kind of funny that the wisdom in software development is to program against interfaces and not implementations. And yet here we are with monorepos, doing the big-ball of mud approach. I've worked on several multi-repo systems and several monorepos. I have a weak preference for monorepos for some of the reasons given, especially the spread of pull requests, but that's almost a 'code smell' in some respects. Monorepos that I've contributed to that have worked well: mostly one language (but not always), a single top-most build command that builds everything, complete coverage by tests in all dimensions, and the repo has tooling around it (ensure code coverage on check in, and so on). Monorepos that I've contributed to that haven't: opposites of the previous points. Multi-repos that have worked well: well abstracted and isolated, some sort of artefact repository (nexus, jfrog, whatever) as the layer between repos, clear separation of concerns. Multi-repos that have not worked well: again, opposites of the previous, including git submodules (please, just don't), code duplication, fragile dependencies where changing any repo meant all had to change.
- deleted 2y ago[deleted]
- indulona 2y agomonorepo is the way to go if the code portrays to the entire application as a whole. otherwise, if there are applications that are not connected in any way, it makes absolutely no sense to pull them together. it's really not a rocket science. some people just prefer to ice skate uphill, i guess, and have to make simple things complicated.
- default-kramer 2y ago> Refactoring across repository boundaries requires much more activation energy as compared to spotting and performing gradual refactorings across folder boundaries. Technically it is the same, but the psychological barriers are different. I love the "activation energy" metaphor here. But I don't agree that "technically it is the same." At my current job, we have more than 100 minirepos and I am unable to confidently refactor the system like I normally would in a monorepo. It's not merely the psychological barrier. It's that I am unable to find all the call sites of any "published" function. Minirepos create too many "published" functions in the form of Nuget packages. Microservices create too many "published" functions in the form of API endpoints. In either case, "Find All References" no longer works; you have to grep, but most names are not unique enough. For this reason, the kind of refactoring that keeps a codebase healthy happens at a much lower rate than all the other projects I've worked on.
- mcnichol 2y agoMonolith vs Microservice argument all over again. Tradeoffs for mono are drivers of micro and vice versa. Looking at the GitHub insights it becomes pretty clear there are about two key devs that commit or merge in PRs to main. I'm guessing this is also whom the code reviews happen etc. Comparing itself to Linux where the number of recurring contributors are more by orders of magnitude just reeks of inexperience. I'm being tough with my words because at face value, the monorepo argument works but it ends in code-spaghetti and heartache when things like developer succession, corporate strategy, market conditions throw wrenches in the gears. Not for nothing I think a monorepo is perfectly fine when you can hold the dependency graph (that you have influence over) in your head. Maybe there's a bit of /rant in this because I'm tired of hearing the same problem with solutions that are spun as novel ideas when it's really just: "Pre-optimization is the root of all evil." You don't need to justify using a monorepo if you are small or close to single threaded in sending stuff into main. It's like a dev telling me: "I didn't add any tests to this and let me explain why..." The explanation is the admission in my mind but maybe I'm reading into it too much. Article is nicely written and an enjoyable read but the arguments don't have enough strength to justify. You are using a monorepo, that's okay. Until it's not, that's okay too.