14 ms·
The Ingredients of a Productive Monorepo
- jph 1y agoGood practical article, thank you. I've added the link to my monorepo-vs-polyrepo guide here: https://github.com/joelparkerhenderson/monorepo-vs-polyrepo/ https://github.com/joelparkerhenderson/monorepo-vs-polyrepo/
- swgillespie 1y agothanks!
- Flux159 1y agoThis definitely tracks with my experience in big tech - managing large scale build systems ends up taking a team that works on the build system itself. The underlying repo technology itself needs to work at scale & that was with a virtual file system that downloaded source files on demand when you needed to access them. One thing that this article didn't mention is that most development was done either on your development server running a datacenter (think ~50-100 cores) - or on an "on demand" machine that was like a short lived container that generally stayed up to date with known good commits every few hours. IDE was integrated with devservers / machines & generally language servers, other services were prewarmed or automatically setup via chef/ansible, etc. Rarely would you want to run the larger monorepos on your laptop client (exception would generally be mobile apps, Mac OS apps, etc.).
- swgillespie 1y agoYeah - I worked on that build team probably at the same place you did! I think for a lot of users it's more important that the monorepo devenv be reproducible than be specifically local or specifically remote. It's certainly easier to pull this off when it's a remote devserver that gets regularly imaged.
- codethief 1y ago> Yeah - I worked on that build team probably at the same place you did! I did not work at that place but the story sounds very familiar – I believe there might have been a blog post about that remote development environment here on HN some time ago?
- zer00eyz 1y ago> One thing that this article didn't mention is that most development was done either on your development server running a datacenter (think ~50-100 cores) I have done this for many small teams as well. It remains pretty hard to get engineers to stop "thinking localy" when doing development. And with what modern hardware looks like (in terms of cost and density) it makes a lot of sense to find a rack some where for your dev team... It's easy enough to build a few boxes that can run dev, staging, test and what ever other on demand tooling you need with room to grow. When you're close to your infrastructure and it looks that much like production, when you have to share the same playground the code inside a monorepo starts to look very different. > managing large scale build systems ends up taking a team that works on the build system itself This is what stops a lot of small teams from moving to monorepo. The thing is, your 10-20 person shop is never going to be google or fb or ms. They will never have large build system problems. Maintaining all of it MIGHT be someone's part time job IF you have a 20 person team and a very complex product. Even that would be pushing it.
- tayo42 1y agoWorking with a well maintained mono repo is so nice, any other workflow just sucks to go back to. Working with a "lets do a monorepo" monorepo, where who ever set it up didn't understand the points in this article and more is a nightmare. I think this is a business opportunity, if someone could sell the polished monorepo experience and tools to companies with engineering organizations but can't pull off a successful "we need to fork git" project to support their developers.
- halflife 1y agoIt is a business opportunity, NX is offering it. In my previous startup, I started developing from the get go with NX, it became a huge velocity boost to our team. With 15 person RND we had standards that a 100 person RND didn’t accomplish. In my new company (which has bought the startup), they tried the “let’s do a monorepo” approach. It is a catastrophe. I am now in the process of migrating them to NX with great results.
- mierz00 1y agoLikewise, we’re using NX at my company and it has been a great experience. Previous mono repo experiences were nothing short of a nightmare so it’s refreshing to see tooling come so far.
- baq 1y agoThey’re on the right path but still have a lot to learn in the testing department. Source: my job’s monorepo is running nx, but I’m not in developer productivity; used to work with a large codebase with an accompanying test suite of thousands of hours of a single box and it’s kinda like watching people rediscovering the roundness required to make the wheel.
- lxe 1y agoSo there are 2 kinds of big tech monorepos. One is the kind described in the article here: "THE" monorepo of the (mostly) entire codebase, requiring custom VCS, custom CI, and a team of 200 engineering supporting this whole thing. Uber and Meta and I guess Google do it this way now. It takes years of pain to reach to this point. It usually starts with the other kind of "monorepo": The other kind is the "multirepo monorepo" where individual teams decide to start clustering their projects in monorepos loosely organized around orgs. The frontend folks want to use Turborepo and they hate Bazel. The Java people want to use Bazel and don't know that anything else really exists. The Python people do whatever the python people do these days after giving up on Poetry, etc... Eventually these might coalesce into larger monorepos. Either approach costs millions of dollars and millions of hours of developers' time and effort. The effort is largely defensible to the business leaders by skillful technology VPs, and the resulting state is mostly supported by the developers who chose to forget the horror that they had to endure to actually reach it.
- Pawka 1y agoIt's worth noting that most monorepos won't reach the same size as repositories from Google, Uber, or other tech giants. Some companies introduce new services every day, but for some, the number of services remains steady. If a company has up to 100 services, there won't be VCS scale problems, LSP will be able to fit the tags of the entire codebase in a laptop's memory, and it is probably _almost_ fine to run all tests on CI. TL;DR not every company will/should/plan to be the size of Google.
- CamouflagedKiwi 1y agoI do think the 'run all tests on CI' part is not that fine, it bites a lot earlier than the others do. Git is totally fine for a few hundred engineers and 100ish services (assuming nobody does anything really bad to it, but then it fails for 10 engineers anyway), but running all tests rapidly becomes an issue even with tens of engineers. That is mitigated a lot by a really good caching system (and even more by full remote build execution) but most times you basically end up needing a 'big iron' build system to get that, at which point it should be able to run the changed subset of tests accurately for you anyway.
- baq 1y agoPerfect write up. Rarely do I nod and murmur ’yes’ and ‘finally someone has written about it’ alternatively on each paragraph.
- AlotOfReading 1y agoOne thing I don't usually see discussed in monorepo vs multi repo discussions is there's an inverse Conway's law that happens: choosing one or the other will affect the structure of your organization and the way it solves problems. Monorepos tend to invite individual heroics among common infrastructure teams, for example. Because there are so many changes going in at once, anything touching a common area has a huge number of potential breakages, so effort to deliver even a single "feature" skyrockets. Doing the same thing in a multi-repo may require coordinating several PRs over a couple of weeks and some internal politics, but that might also be split among different developers who aren't even on a dedicated build team.
- makeitdouble 1y agoIs your underlying assumption that the organization doesn't want to go one way or the other in the first place and is nudged by the technical choice afterwards ? I think most of the time the philosophical decision (more shared parts or better separation) is made before deciding how you'll deal with the repos. Now, if an org changes direction mid-way, the handling of the code can still be adapted without fundamentally switching the repo structure. Many orgs are multi-repo but their engineers have access to almost all of the code, and monorepo teams can still have strong isolation of what they're doing, up to having different CI rules and deployment management.
- TeMPOraL 1y agoI think GP's claiming it's a feedback loop, not one-directional relationship. Communication structure of an organization ends up reflected in the structure of systems it designs, and at the same time, the structure of a system influences the communication structure of the organization building it. This makes sense if you consider that: 1) Changes to system structure, especially changes to fundamentals when the system is already being built, are difficult, expensive and time consuming. This gives system designs inertia that grows over time. 2) Growing the teams working on a system means creating new organizational units; the more inertia system has, the more sense it makes for growth to happen along the lines suggested by system architecture, rather than forcing the system to change to accommodate some team organization ideals. Monorepo/multirepo is a choice that's very difficult to change once work on building the system starts, and it's a choice you commit at the very beginning (and way before the choice starts to matter) - a perfect recipe for not a mere nudge, but a scaffolding the organization itself will grow around, without even realizing it.
- deleted 1y ago[deleted]
- ianpurton 1y agoI've never worked on a mono repo that has the whole organizations code in it. What are the advantages vs having a mono repo per team?
- AlotOfReading 1y agoOne of the big advantages is visibility. You can be aware of what other people are doing because you can see it. They'll naturally come talk to you (or vice versa) if they discover issues or want to use it. It also makes it much easier to detect breakages/incompatibilities between changes, since the state of the "code universe" is effectively atomic.
- lenkite 1y agoNot sure if I get it. If you are using a product like Github Enterprise, you are already quite aware of what other people are doing. You have a lot of visibility, source-code search, etc. If you have a CICD that auto-creates issues you already can detect breakages, incompatibilities, etc. State of the "code universe" being atomic seems like a single point of failure.
- jeffbee 1y agoGitHub search is insanely bad and it cannot do things like navigating to definitions between repos in an org.
- lenkite 1y agoIf you want code search and navigation over a closed subgraph of projects that build into an artifact - opengrok does the job reasonably well.
- eddd-ddde 1y agoImagine team A vendors into their repo team B's code and starts adding their own little patches. Team B has no idea this is happening, as they only review code in repo B. Soon enough team A stops updating their dependency, and now you have two completely different libraries doing the "same" thing. Alternatively, team A simple pins their dependency to team B's repo at hash 12345, then just, never updates... How is team B going to catch bugs that their HEAD introduces on team A's repo?
- pawanjswal 1y agoI felt like a pep talk and reality check rolled into one.
- vinnymac 1y agoI established monorepos for the last two large projects I operated. I’ve never heard such nice compliments from contributors in my whole career. It seems not only can it be a productivity booster but people genuinely love when things are easy to grok and painless. Multiple large monorepos in an organization are highly valuable imo, and should become more of a thing over time.
- wocram 1y agoWhy multiple manyrepos over a single monorepo?
- vinnymac 1y agoIt would require a blog post for me to answer this in detail, but overall it’s due to the fact that monorepos trend toward specific stacks and runtimes. They aren’t often as flexible as they’d have you believe, so you might find a low level programmer isn’t able to operate as efficiently when operating in a primarily NodeJS-based Monorepo. Plus it makes it easier to separate parts of your business that are distinct, simplifying the complexity of security, maintenance, and compliance. I don’t think one should have 100s of them, practically speaking less than 5 should be enough to model any business today. It can be freeing though, and allows developers to take advantage of the tools that best fit there needs so they can go back to getting shit done.
- jbverschoor 1y agoIs there a way to set permissions on certain directories / force partial clones. Not just a sparse clone.
- echelon 1y agoYou can set permissions on writes. Optional, per-directory OWNERS files are common, and most VCS frontends (Github, Bitbucket, etc.) can be configured to prevent merges without approval from the owning team(s) or DRI(s). PRs that intersect multiple teams' ownership would require handoff of everyone impacted. So a team updating the company-wide "requests library" (or an equivalent change), with a wide blast radius, would be notifying everyone impacted and getting their buy-in.
- Pawka 1y agoIt depends on the VCS you use. I don't know any ways to manage read permissions, such as allowing a person to checkout one directory but not another, though you can do that per branch on git. But there are many ways to manage write permissions - limit the directories to which engineers are allowed to push code. E.g. if you use Git, this can be done with Gitolite, which is a popular hosting server. Gitolite has very flexible hooks support, especially with so-called "Virtual Refs" (or VREFs)[1]. It is out of the box and has support to manage write permissions per write path [2]. You can go even further and use your own custom binary for VREF to "decide" if a user is allowed to push certain changes. One possible option - read incoming changed files, read metainformation from the repository itself (e.g., CODEOWNERS file at the root of the repo), and decide if push should be accepted. GitHub has CODEOWNERS [3], which behaves similarly. [1]: https://gitolite.com/gitolite/cookbook.html#vrefs https://gitolite.com/gitolite/cookbook.html#vrefs [2]: https://gitolite.com/gitolite/vref.html#quick-introexample https://gitolite.com/gitolite/vref.html#quick-introexample [3]: https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners https://docs.github.com/en/repositories/managing-your-reposi...
- jbverschoor 1y agoIt's mostly about read/access permissions. I'd like to stay away from any type of git hook tbh
- lihaoyi 1y agoI wrote a bit about monorepo tooling in this blog post. It covers many of the same points in the OP, but in a lot more detail. - https://mill-build.org/blog/2-monorepo-build-tool.html https://mill-build.org/blog/2-monorepo-build-tool.html People like to rave about Monorepos, and they are great if set up correctly, but there's a lot of intricacies that often goes on behind the scenes to make a Monorepo successful that it's easy to overlook since usually some "other" team (devops teams, devtools team, etc.) is shouldering all that burden. Still worth it, but most be approached with caution
- kfkdjajgjic 1y agoThe artikel doesn’t bringa it up, but I’ve seen several places where repos has been cut according to company silos, where applikation code was in a monorepo for all teams, IaC was in one monorepo for all teams, and ops was in one monorepo for all teams. It was not good at all.
- bluGill 1y agothat isn't a mono repo, it is a polyrepo setup with the wrost features of a polyrepo. I use a polyrepo setup and it works well, but we need careful attention to the repo split - and repo where too many different teams work together gets the worst features of a monorepo combined with the worst features of a poly-repo. We have a lot of tooling around making the different repos stay in sync.
- woile 1y agoI've been very happy with nix. I've been using nix in the reciperium.com monorepo, granted, it's only me, but I'm quite happy with having everything there. From docs, to the infra with terraform, to frontend and backend. The procedure for the CI is quite straightforward (nix build .#project), and caching the dependencies in the CI works quite okay. Even the secrets are there, encrypted using age (might not be the best, but good enough).
- wocram 1y agoNix is great until you're rebuilding multiple packages that are slow to build without incrementality (eg. C++ or Rust).
- cloogshicer 1y agoHere's what I never got about monorepos: Imagine you have an internal library and also two consumers of that library in the repo. But then you make breaking changes to the library but you only have time to update one of the consumers. Now how can the other consumer still use the old version of that library?
- code_biologist 1y agoThat's the neat part. They don't. Either the broken consumer updates their use, you update it for them to get your change shipped, or you add some backwards compatibility approach so your breaking changes aren't breaking.
- cloogshicer 1y agoThanks for the info! Seems like a big restriction to me.
- bluGill 1y agoit is the only sane thing to do. Allowing everyone to use their own fork means when a major bug is found you have to fix thousands of forks. If the bug is a security zero day you don't have time.
- cloogshicer 1y agoCouldn't you just leave the other consumer at the old release (presumably well tested, stable)? I don't see how being forced to upgrade all consumers is a good thing.
- bluGill 1y agoMapbe but each one now is another thing that you need to fix if a major issue is found. If it is only a few releases not a problem but it can get to hundreds and that becomes hard. Particularly if the fix can be cherry-picked cleanly to other branches.
- slippy 1y agoIt's also worth noting that in systems that get as large as Google's that you end up with commits landing around the clock. It gets so that it's impossible to test everything for an individual commit, so you have a 2nd kind of test that launches all tests for all branches and monitors their status. At Google, we called this the Test Automation Platform (TAP). One cool thing was that it continuously started a new testing run of all testable builds every so often -- say, 15 minutes, and then your team had a status based on the flaky test failures vs solid test failures of if anyone in any dependency broke your code. So if your code is testing fine, and someone makes a major refactor across the main codebase, and then your code fails, you have narrowed the commit window to only 15 minutes of changes to sort through. As a result, people who commit changes that break a lot of things that their pre-commit testing would be too large to determine can validate their commits after the fact. There's always some amount of uncertainty with any change, but the test it all methodology helps raise confidence in a timely fashion. Also decent coding practices include: Don't submit your code at the end of the day right before becoming unavailable for your commute...
- yc-kraln 1y agoThe answer, of course, is "it depends". We have something like ~40 repos in our private gitlab repo, and each one has its own CI system, which compiles, runs tests, builds packages for distribution, etc. Then there's a CI task which integrates a file system image from those ~40 repo's packages, runs integration tasks, etc. Many of those components communicate with each other with a flatbuffers-defined message, which of course itself is a submodule. Luckily, flatbuffers allows for progressive enhancement, but I digress--essentially, these components have some sort of inter-dependency on them which at the absolute latest surfaces at the integration phase. Is this actually a multi-repo, or is it just a mono-repo with lots of sub-modules? Would we have benefits if we moved to a mono-repo (the current round-trip CI time for full integration is ~35 minutes, many of the components compile and test in under 10s)? Maybe. Everything is a tradeoff. Anything can work, it's about what kinds of frustrations you're willing to put up with.
- nssnsjsjsjs 1y ago> Any operation over your repository that needs to be fast must be O(change) and not O(repo). This is a good thought! It actually needs to be O(1/commit rate) though, so that having the monorepo doesn't create long queues of commits. Or have some process batch passing ready to merge PRs into a combined PR and try to merge that. And best guess on the failing PR if it fails.
- atq2119 1y agoIf you go the batching route, bisection on failure makes it more like O(log(1/commit rate)).
- eddd-ddde 1y agoJust make presubmit a fraction of your postsubmit. Each change has fast operations while still having global testing. Then if postsubmit fails you just have to rerun the intersection of failing tests and affected tests on each change since the last green commit.
- spankalee 1y agoFor those of you working in Node and npm, npm has pretty good built-in support for monorepos now with the workspaces feature. The big missing thing is incremental builds, which I highly recommend looking at Google's Wireit project for: https://github.com/google/wireit/ https://github.com/google/wireit/ Wireit is the smallest change from plain npm that gets you a real dependency graph of scripts, caching (with GitHub Actions support), incremental script running, and services.
- spankalee 1y agoI love monorepos, but in large organizations they have a counter-intuitive incentive for teams to _not_ allow other teams to depend on them, which can _reduce_ code reuse - the opposite of what some adopters want. This issue is that users of a library can put almost infinite friction on the library. If the library team wants to make a change, they have to update all the use sites, but Hyrum's Law will get you because users will do the damndest things. So for the top organization, it's good if many other teams can utilize a great team's battle-tested library, but for the library team it's just liability (unless making common code is their job). In a place like Google you either end up with internal copies and forks, strict access control lists, or libraries that are slow as molasses to change.
- eddd-ddde 1y agoWell when making a library that's intended to be shared, you REALLY need to stop for a second and think about the API. Ideally APIs don't change, and when they do, you better have planned for large scale changes, or just use a new function and mark the old one deprecated. I don't think there's anything wrong with copy pasting some useful piece of code too, not everything has to be a library you depend on, for small enough things.
- kccqzy 1y agoAt least the benefit of a monorepo is that you can find all the use sites in the first place and correct these wrong uses. You can even correct them atomically if you so wish.
- ec109685 1y agoI would still say code is more likely to be reused in the monorepo versus trying to take an external dependency in the poly repo case. Just the ease of making a change to target your case is so much higher.
- wocram 1y agoAll software with dependencies needs to respect it's dependents. A monorepo doesn't really change anything about the relationship between a library and it's users, except that the library or the users are somewhat more empowered to change each other.
- zvr 1y agoGenuine question, because I've never worked somewhere with a monorepo infrastructure: is it really "one repo for all code in the organization" or "one repo for everything related"? In my organization we have around 70k internal git repos (and an order of magnitude fewer public ones), but of course not everything is related to everything else; we produce many distinct software products. I can understand "collect everything of a product to a single repo"; I can even understand going to "if there is a function call, that code has to be in the same repo". But putting everything into a single place... What are the benefits?
- scott01 1y agoIn game dev monorepo per product is often used, which includes game code, art assets, build system and tooling, as well as engine code that can receive project-specific patches. In Perforce, it's organised into streams, where development streams are regularly promoted to staging, then to release, etc. The benefit is the tooling, as the article mentioned. Everything in the repo is organised consistently, so I can make ad-hoc Python tools relying on relative paths knowing that my teammates have identical folder structure.
- anon7000 1y agoWhen you have N repos, you also have N ways of managing dependencies, N ways of doing local bin scripts and dev environment setups, N projects with various out of date & deprecated setups, N places to look when you need to upgrade a vulnerable dependency, N services which may or may not configure telemetry in a consistent way, N different CI & deployment workflows… It just gets very difficult to manage, especially if people frequently need to work across many repos. Plus, onboarding is a pain in the ass. Monorepo example: if I want to add a new Typescript package/library for internal NodeJS use, we have a boot strapping script that sets it up. And it basically: 1. Inherits a tsconfig that just works in the context of the repo 2. Jest is configured with our default config for node projects and works with TS out of the box. 3. Listing / formatting etc are all working out of the box. 4. Can essentially use existing dependencies the monorepo uses 5. Imports in existing code work immediately since it’s not an external dependency 6. CI picks up on the new typescript & jest configs and adds jobs for them automatically 7. Code review & collaboration is happening in the same spot 8. This also makes it easier to have devs managing the repo — for example, routine work like updating NodeJS is a lot easier when you know everything is using a nearly identical setup & is automatically verified in CI. One challenge I had to help solve in a previous job was that onboarding was difficult because we had a small number of large repos everyone worked in. The standards were slightly different across them. Npm, pnpm, and yarn were all in use. Deployment worked pretty differently among them. CI setups were unique, and each of the large projects had, if not a team, some number of people spending a lot of time just managing the project’s workflows. So many coordination things just get easier when there isn’t an opportunity to get out of sync. If you do separate repos, you can totally share config… but now it costs a dependency update PR to pull in that tiny update to the shared unit test config and now everything. It’s just guaranteed to get out of sync, and it’s hard to catch issues when you can’t validate a config change with all projects using it at the same time. So because it becomes trickier (and takes work) just to do the action of syncing multiple repo’s setups… inevitably, you end up with some “standards” that are loosely followed and a lot of slightly different setups that get hard to untangle the longer they grow. If you can accept the cost of context switching between repos, or if people don’t need to switch, maybe it’s ok… until something like a foundational dependency update (NodeJS, Typescript, React, something like that) needed for security becomes extremely difficult because you have a million different ways of configuring things and the JS ecosystem sucks
- deleted 1y ago[deleted]
- bob1029 1y agoThis thread is reminding me of a prior one about complexity merchants. I am seeing a lot of sentiment that there is somehow a technical sacrifice by moving to a monorepo. This is absolutely ludicrous unless you fail to grasp the power of a hierarchical file system. I don't see how a big mess like CI/CD is made easier by spreading it out to more points of configuration. To me the whole point of a monorepo is atomic commits for the whole org. The power of this is really hard to overstate when you are trying to orchestrate the efforts of lots of developers - contrary to many claims. Rebasing in one repo and having one big meeting is a hell of a lot easier than doing it N times. Even if the people on the team hate each other and refuse to directly collaborate. I still don't see the reason to not monorepo. In this scenario, the monorepo becomes a useful management and HR tool.
- cmrdporcupine 1y agoThe push to fragmentation and atomism is so strong with this generation of devs. The obsession with microservices, dozens of small repositories, splitting everything up from fear of "monoliths." What they're doing is creating a mass of complexity that is turning org-chart problems into future technical ones and at the same time not recognizing the intrinsic internal dependencies of the software systems they're building. Luckily my current job is not like this, but the last one was, and I couldn't believe the wasted hours spent doing things as simple as updating the fields in a protobuf schema file.
- bluGill 1y agoThat push to fragmentation is in large part because of hard lessons learned from the problems of a monolith. The answer is IMO somewhere in between. Microservices can get too tiny and thus the system becomes impossible to understand. However a monolith is impossible to understand as well. The real problem is you need good upfront architecture to figure out how the whole system fits together. However that is really hard to get right (and Agile discourages it - which is right for small projects where those architects add complex things to mitigate problems you will never have)
- 1y ago
- gorgoiler 1y agoAn unspoken truth of a monorepo is that everyone is committed to developing on trunk, and trunk is never allowed to be broken. The consequence of this is that execution must be configurable at runtime: feature flags and configuration options with old and new code alongside each other. You can have a monorepo and still fail if every team works on their own branch and then attempts to integrate into trunk the week before your quarterly release process begins. You can fail if a core team builds a brand new version of the product on master with all new tests such that everything is green on every commit but your code is unreleasable because customers aren’t ready for v2 and you need to keep that v1 compatability around.
- 946789987649 1y agoI didn't know places still had quarterly releases. That seems to like the one to resolve rather than a mono repo.
- bluGill 1y agonot all the world is a web site or even internet connetted. not all the world has no safety concerns. if you work in medical or aviation areas every release legally needs extensive - months - testing before you can release. If there are issuse found in that testing you start over. Not all tests can be automated. i work in agraculture. the entire month of July there will be nobody in the world using a planter or any of the software on it. there is no point in a release then. the lack of users means we cannot use automated rollback if the change somehow fails for customers - we could but it would be months of changes rolled back whe Brasil starts planting season.
- vegetablepotpie 1y agoEvery company that uses SAFe agile has quarterly, or bi-quarterly, releases [1]. [1] https://www.servicenow.com/docs/bundle/yokohama-it-business-management/page/product/agile-SAFe/task/create-SAFeprogramincrement.html https://www.servicenow.com/docs/bundle/yokohama-it-business-...
- gorgoiler 1y ago
- cormacrelf 1y ago> Meta has a sophisticated implementation of a target determinator on top of buck2, but I don’t believe it is open-source. It is: https://github.com/facebookincubator/buck2-change-detector https://github.com/facebookincubator/buck2-change-detector > Some tools such as bazel and buck2 discourage you from checking in generated code and instead run the code generator as part of the build. A downside of this approach is that IDE tools will be unable to resolve any code references to these generated files, since you have to perform a build for them to be generated at all in the first place Not an issue I have experienced. It's pretty difficult to get into a situation where your IDE is looking in buck-out/v2/gen/781c3091ee3/... for something but not finding it, because the only way it knows about those paths is by the build system building them. Seeing this issue would have to involve stale caches in still-running IDE after cleaning the output folder, which is a problem any size repo can have. In general, if an IDE can index generated code with the language's own build system, then it's not a stretch to have it index generated code from another one. The problem is more hooking up IDEs to use your build system in the first place. It's a real slog to support many IDEs. Buck recently introduced an MSBuild project generator where all build commands shell out to buck2. I have seen references to an Xcode one as well, I think there's something there for Android as well. The rust-analyzer support works pretty well but I do run a fork of it. This is just a few. There is a need (somewhat like LSP, but not quite) for a degree of standardization. There is a cambrian explosion of different build systems and each company that maintains one of them only uses one or two IDEs and integrates with those. If you want to use a build system for an IDE they don't support, you are going to have a tough time. Last I checked the best effort by a Language Server implementation at being build-system agnostic is gopls with its "gopackagesdriver" protocol, but even then I don't think anyone but Bazel has integrated with it: https://github.com/bazel-contrib/rules_go/wiki/Editor-and-tool-integration https://github.com/bazel-contrib/rules_go/wiki/Editor-and-to...
- rwieruch 1y agoOver the past four years, I’ve set up three monorepos for different companies as contract work. The experience was positive, but it’s essential to know your tools. Since our monorepos were used exclusively for frontend applications, we could rely entirely on the JavaScript/TypeScript ecosystem, which kept things manageable. What I learned is that a good monorepo often behaves like a “polyrepo in disguise.” Each project within it can be developed, hosted, and even deployed independently, yet they all coexist in the same codebase. The key benefit: all projects can share code (like UI components) to ensure a consistent look and feel across the entire product suite. If you're looking for a more practical guide, check out [0]. [0] https://www.robinwieruch.de/javascript-monorepos/ https://www.robinwieruch.de/javascript-monorepos/
- wocram 1y agoThis isn't a polyrepo in disguise. This is a monorepo done correctly.
- bittermandel 1y agoI firmly believe that us at Molnett(serverless cloud) going for a strict monorepo built with Bazel has been paramount to us being able to make the platform with a small team of ~1.5 full-time engineers. We can start the entire platform, Kubernetes operators and all, locally on our laptops using Tilt + Bazel + Kind. This works on both Mac and Linux. This means we can validate essentially all functionality, even our Bottlerocket-based OS with Firecracker, locally without requiring a personal development cluster or such. We have made this tool layer which means if I run `go` or `kubectl` while in our repo, it's built and provided by Bazel itself. This means that all of us are always on the same version of tools, and we never have to maintain local installations. It's been a HUGE blessing. It has taken some effort, will take continuous effort and to be fair it has been crucial to have an ex Google SRE on the team. I would never want to work in another way in the future. EDIT: To clarify, our repo is essentially only Golang, Bash and Rust.
- mattmanser 1y agoThe question here is why are you using micro service pattern and k8s with 2 Devs. That pattern is not designed for that small scale operation and adds tons of completely unnecessary complexity. And does it really matter what you go with when you've got 1.5 engineers? It's a non-problem at that scale as both engineers are intimately aware of how the entire build process works and can keep it in their head. At that scale I've done no repo at all, repo stored on Dropbox, repo in VCS, SVN, whatever, and it all still worked fine. It really hasn't added anything at all to your success. BTW, it's still common for developers to start entire repos on their own laptops with zero hassles in tons of dev shops that haven't been silly and used k8s with 2 developers. In fact at the start of my career I worked with 10 or so developers the shitty old MS one where you had to lock files so no-one else can use them. You'd checkout files to allow you to change them (very different to git checkout), otherwise they'd be ready only on your drive. And the build was a massive VB script we had to run manually with params. And it still worked. We got some moaning when we moved to SVN too at how much better the old system was. Which was ridiculous as you used to have to run around and ask people to unlock key files to finish a ticket, which was made worse as we had developer consultants who'd be out of office for days on end. So then you'd have to go hassle the greybeard who had admin rights to unlock the file for you (although he wasn't actually that old and didn't have a beard).
- chrismatic 1y agoThe point about trying to stick with a single language build tooling really cannot be stressed enough. It is what prompted me to write a simplified version of Bazel, a generic "target determinator" with caching capabilities if you will. I call it "Grog", the monorepo build tool for the grug-brained developer. https://grog.build/why-grog/ https://grog.build/why-grog/
- bluGill 1y agoIf a single language is an option you are a small project that is not facing the problems people on large projects are facing. A monorepo will be easy for you without read the article and the lessons learned. Come back when you have millions of lines of code, written over decades by hundreds (or thousands) of full time developers.
- dezgeg 1y agoWhat a weird take, "millions of lines of code, written over decades" applies to quite many C (or C++) codebases where using a high-level language is not a possibility (and companies that do have such codebases are pretty conservative and don't even talk about Rust no matter how great fit it would be).
- bluGill 1y agoIn every case I've seen the vast majority might be C, but there are other other languages hidden in there that are hard to find. Many companies would use more languages if it wasn't such a pain. Rust for example would be really nice to use in new code, if only they can figure out how to mix it in.
- chrismatic 1y agoThere is a space between the two types of repositories you are describing. One where you just have enough tools/langs that a single-language setup does not cut it for you anymore, but investing all that effort into Bazel does not seem worth it yet. That is the gap that Grog is meant to fill.
- 1y ago
- boxed 1y agoThe article links to a site with this definition: > A monorepo is a single repository containing multiple distinct projects, with well-defined relationships. It would be better if there were terms that delineated "one repo for the company" from "one repo per project" from "many repos for a single project".
- bluGill 1y agoIdealy the term would indicate code and team size. Many commenting are working on tiny projects where they don't even see the problems that cause one to think of this debate
- wocram 1y agoI think most monorepo advocates are actually anti "one repo per project" at heart. That's the real anti-pattern imo.
- boxed 1y agoI don't get it. That's the only pattern that makes any sense imo. If you get problems because two or more projects needs to be updated simultaneously because of DRY or whatever, I would argue that you actually have ONE project and you're just fooling yourself. Hmm.. actually, I guess that's what you're saying? When I said "project" above, I meant something like "product" or "system".
- KaiserPro 1y agoOne of the things not covered here is how to deal with versioning. By default a monorepo will give you $current and nothing else. A monorepo is not a bad idea, but you should think about either preventing breaking changes in some dependency killing the build globally, or have some sort of artefact store that allows versioned libraries (both have problems, you'll need to work out which is better for you. )
- trollbridge 1y agoI have been approaching this by eventually breaking out a module into its own repo when the time comes for that (enough resources to dedicate to maintaining it independently, having tests, and so forth). When the folks working on the monorepo really need to slam through a change in the now-independent monorepo, we can use git submodules.
- DrScientist 1y agoI think a key idea often associated with the use of a monorepo is to encourage developer behaviour to do the integration/mitigation work at the point of change, rather than creating lots of integration debt in the form of versions ( however you do it ). You need to look at your development model as a whole and decide whether the happy path incentivises good or bad development practices. Do you want to incentivise the creation of technical debt with a myriad of versioned dependencies or do you want to incentivise designing code to be evolvable and resuable?
- KaiserPro 1y agoI worked at a startup with a "monorepo" (C++, cuda and python) it worked well and wasn't too hard to manage. Once someone bit the bullet and made some robust bazel spells it was brilliant to use and multi-platform too. Worked at a FAANG with a monorepo, and everything was partially broken most of the time. Its trivial to bring in dependencies, which is great, super fast re-use. The problem is, its trivial to add dependencies. That means that bringing in a library to manage messages also somehow requires a large amount of CUDA code as well. A basic python programme would endup having something like >10k build items to go through each build.
- countWSS 1y agoFrom viewpoint of security and separation of concerns giving unlimited access to everything by virtue of "everything" being stored in one giant repo sounds exceptionally short-sighted. A single rogue actor would be able to insert code to any component of choice instead of working on isolated repo with people who specifically know it and approve the code: the monorepo is a "big ball of mud" with vague shared responsibility that defers to people who worked on "specific parts" but they lack any authority or control, auditing the entire codebase doesn't scale.
- morbicer 1y agoCodeowners file + required review from the owner team solves like 90% of those worries
- jasminebelmont 1y ago[flagged]
- wh0knows 1y agoMonorepo != all devs having merge permissions to all directories. Every single large monorepo company will have granular permissions on who can approve PRs into which directories based on team ownership. This is orthogonal to monorepo vs polyrepo.
- deleted 1y ago[deleted]
- cousin_it 1y agoI've worked for a company with a large monorepo. At first I was a fan, but now I'm not so sure. The web of dependencies was too much. Now I think teams should be allowed to reuse other teams' code only as libraries or APIs with actual release cycles. There shouldn't be any "oh let's depend on the HEAD of this random build target somewhere else in the monorepo". There should be only "let's depend on a released version of such-and-such library or API". If you adopt this discipline, you basically don't need a monorepo. Every team can have its own repo and depend on other stuff as third party. This adds some friction, but removes some other kinds of friction, and overall I think it's a better compromise.
- ellisv 1y agoI haven't really worked with any large monorepos. I find your comment really interesting because having the capability to point to the HEAD (or realistically a commit SHA) is a feature I sometimes really enjoy about not using monorepos.
- eddd-ddde 1y agoThis just creates tons of fragmentation. The second you have multiple teams depending on multiple versions you are doomed. You are stuck maintaining multiple versions, with their own quirks and bugs. I think the one version rule is the most important part for a healthy monorepo.
- bluGill 1y agoI think one version is important for a healty polyrepo as well. You have to set lines where you say no new features unless you are all up to date. You can allow bug fix only releases to stay behind, but if you write a new feture it must be against the current latest of everything. Otherwise you are doomed because there are so many different versions of everything in use. Some day a zero-day issue will hit all your projects as the same time and you will need months to get each in use version fixed.
- senderista 1y ago
- calvinmorrison 1y agoOne to look at historically was KDE using SVN. all the downside of svn the partial checkout was great for a repo containing practically the entire K source tree
- bigbuppo 1y agoIt's kind of weird that both Microsoft and Google were both using Perforce. What does Perforce do that worked well at those companies for so long, and what caused them to dump it? Did they just get tired of the licensing cost? I think what I'm getting at is that maybe the real missing feature isn't whatever it is that allows you to make stupidly large monorepos, but that maybe we should add Perforce's client workspace model as a git extension?
- senderista 1y agoAt MSFT we used a Perforce fork (Source Depot), but the Windows codebase was still developed in separate repos ("depots"): kernel, shell, graphics, etc. We had custom tooling to coordinate cross-repo changes, so it was still far from a monorepo.
- WorldMaker 1y agoPerforce didn't do anything extraordinarily well, it was just dumb enough it didn't do anything particularly poorly. Perforce had a classic file locking model where a central server was in charge of file locks and a file was read-only until it was unlocked and the number of users that could unlock a file at the same time was often as low as 1. So even if most Perforce operations were O(n^2) or worse, they were often only n = unlocked files, not n = files in repo. git status checks the full worktree, so is n = files in (visible part of) repo. The "file is locked by another user" problem led to doing a lot of work outside Perforce itself. Often diff and patch tools and patch queues/changeset queue tools would proliferate around Perforce repos not provided by Perforce itself, but mini-VCSes built on top of Perforce. (Which is part of why Microsoft entirely forked Perforce early on. If you are already building a VCS toolkit on top of the VCS, might as well control that, too.) A big point about git and its support for offline work, is that it works nothing like Perforce and you mostly don't want it to. A big benefit to git's model is that we mostly aren't using git as a low-level VCS toolkit and using a diaspora of other tools on top of git. (Ironically so, given git's original intent was to be the low-level VCS toolkit and early devs expected more "porcelain" tools to be built on top of it as third-party projects.)
- jonthepirate 1y agoI'm on the build team at DoorDash. We're in year 1 of our Bazel monorepo journey. We are heavy into Go, already have remote execution and caching working, and are looking to add support for Python & C++ soon. If this sort of stuff happens to be something you might want to work on, our team has multiple openings... if you search for "bazel" on our careers page, you'll find them.
- ecoffey 1y agoMonorepo is one of few things I’ve drunk the koolaid on. I joke that the only thing worse than being in a monorepo, is not being in one.
- codethief 1y agoThanks, I'll steal that one! :-)
- v3ss0n 1y agoMonorepo in ai driven development world is a disaster. The context consumption gonna be so off the roof
- l5870uoo9y 1y agoSeparating out the database layer in a monorepo package was the best architectural decision I made this year. Now it is my default because at some point you either want to rebuild the existing app entirely or separate out services such as public API access that all need access to the same database.
- someone654 1y agoCan you elaborate on this? I’m facing a similar decision in my org and think sharing a common database store sounds smart. With rules of course, like clear ownership of data, only one writer, etc.
- marcosdumay 1y ago> in a monorepo package Hum... Does that phrase mean you don't use anything remotely similar to a monorepo?
- s17n 1y agoIf you've got less than 100 engineers, you aren't going to hit any of the scalability issues and there's literally no downside to a monorepo
- nc0 1y agoFor the people interested in a good VCS system to achieve such monorepos, have a look at Ark [0]. It works really well for huge codebases, it is really fast, faster than Perforce Helix, it has an ethical and respectful pricing scheme, with a self-hosting mentality. Also it's indie, which is typically better than greedy corporate. [0]: https://ark-vcs.com https://ark-vcs.com
- joaonmatos 1y agoAs an Amazon employee, this is the kind of discussion that makes me glad we have the Brazil build system.
- scrubs 1y agoThe OP had a point to make then made it. It's refreshing. And, moreover, I'm smarter for it. Well done. Thank you for posting it. As readers may see from my recent comments elsewhere, there's a ton of junk out there. But when the good stuff arrives, one likewise stops, and says so.