11 ms·
Josh: Get the advantages of a monorepo with multirepo setups
- creamytaco 5y agoI don’t see what problems this solves, but I do see plenty of problems it could introduce. There are tremendous benefits to sticking with the tried and tested approach supported by git rather than introducing yet more tooling.
- im_down_w_otp 5y agoI've come to learn that git monorepos are becoming quite popular. What I don't understand is why people are using git for this kind of workflow. It forces you to actively work against git's design goals and implementation. Which then compels the use of several odd workarounds and kludges to kind of seemingly reassemble a half-baked flavor of subversion. Why not just use a tool designed around monorepos and subtrees? I'm genuinely curious. I assume I'm missing something.
- endisneigh 5y agoWhat are these tools designed for monorepos and sub trees?
- fnord123 5y agoperforce helix core
- usrnm 5y agoPerforce?
- im_down_w_otp 5y agoProbably the most popular one would be the "new hotness" that everyone used, or aspired to use, before git became popular, Subversion (https://subversion.apache.org/ https://subversion.apache.org/). There are others that aren't free (e.g. Perforce) and some that aren't quite dedicated to the monorepos & subtree workflow, but which handle it better by design (e.g. Darcs, http://darcs.net/ http://darcs.net/). But, mostly I'm thinking of Subversion.
- 0xbadcafebee 5y agoThe reason I would use a Git monorepo over Subversion is that Subversion is really painful to use. I would rather workarounds and kludges than that tire fire.
- im_down_w_otp 5y agoThat was never my experience with Subversion, but YMMV.
- dalyons 5y agohaha yeah, noones going to go back to subversion. The sum of the pain of the SVN problems solved by git is much larger than the slight pain of using git for a mono-repo.
- triceratops 5y agoSubversion on the server used with git-svn is bearable. Most of the advantages of git (local commits, rebasing, easy branching and merging) manifest themselves on the client-side. The problem with using svn on the server is the lack of good tooling for things like code review. There's no SvnHub or SvnLab.
- deleted 5y ago[deleted]
- gedy 5y agoA big chunk of the industry joining in past decade has never seen/used anything but git, so it's the one hammer they have available.
- quotemstr 5y agoIs that even a bad thing? If people are familiar with git --- its command syntax, its commit model, its collaboration setup --- why not let them keep using this model? Isn't it better for everyone if we put a few hundred people on targeted scalability fixes for git instead of making a few million people drop productive work and learn a new tool? I mean, most of the people in industry today have known no character encoding except ASCII and its supersets --- and that's a good thing!
- Uberphallus 5y agoGit, while better now, it still allows people who are just familiar to shoot themselves in the foot.
- quotemstr 5y agoSo? It's the standard. Anything else starts with -1000 points. Something like hg might be better, but not better enough to be worth the cost of breaking uniformity in the industry.
- prionassembly 5y agoI never get the footgun argument. If you need a gun, you're liable to shoot yourself. Programming in general is more and more accessible to wider audiences, but it hasn't become intrinsically easier.
- deleted 5y ago[deleted]
- ot 5y agoThe problem is that the alternatives are much, much worse. There's nothing about the git/mercurial object models that makes them intrinsically inefficient with monorepos. What's inefficient is materializing the object database (when cloning) and the working copy (when checking out), when you're only going to need tiny portions of them. Subversion doesn't have the first problem (but comes with extremely slow history operations), but sparse checkouts don't really solve the second because you have to statically know what to filter. A better direction is instead to virtualize the filesystem, so you get the semantics of a real monorepo with git/mercurial, and you fetch only what's actually needed without any change to your tooling (it just needs to interact with the filesystem). It's also very easy to transparently implement caching and prefetching this way. This is the approach that Facebook and Microsoft took, and I believe Google too.
- WorldMaker 5y agoThough it is interesting that Microsoft's own efforts have been moving away from the virtual filesystem approach and back towards making sparse checkouts better/more reliable and better/more reliable support for partial checkouts (of history especially) and better memoization and caching of history reachability information (git commit-graph).
- jayd16 5y agoNo other DVCS is as popular. Git has a vast ecosystem of tools. No one wants to drop that. Its certainly not against the design goals as official tooling supports shallow and sparse checkouts, they're just experimental features. We should just strive forward and continue to make these features good.
- an_opabinia 5y ago> It forces you to actively work against git's design goals and implementation. It works against GitHub's design goals. You could work on multiple, logically distinct projects in the same git repository easily, if you so choose. git was designed for Linux's workflow.
- SquishyPanda23 5y ago> What I don't understand is why people are using git for this kind of workflow. People like to use distributed workflows even with monorepos. E.g. chains of commits, branches, rewriting local history, etc. It's clear that people want a mixture of monorepos with distributed workflows. There are two ways to get there: add distributed flows to a monorepos, or build a monorepo layer over a distributed tool. Both seem like valid approaches. The market will decide which approach it prefers.
- deleted 5y ago[deleted]
- TheNewAndy 5y agoIsn't git designed around the linux kernel repository, which is a monorepo? I get the impression that things like LFS, submodule and subtree are hacks added onto git to try to make it behave like something it isn't.
- jazzkingrt 5y agoI think this is quite exciting, as it solves a major unsolved problem for large git monorepos: enabling development or CI/CD inside a git monorepo without requiring a large checkout. As monorepos grow huge, this comes to be very costly or even prohibitive, and companies like Google simply don't use Git. Here are some problems with alternative approaches that have been mentioned: * VFS for Git: I believe abandonded by MSFT in favor of improved client-side tooling: https://github.com/microsoft/VFSForGit/blob/master/docs/faq.md#why-are-you-abandoning-vfs-for-git https://github.com/microsoft/VFSForGit/blob/master/docs/faq.... . * Sparse checkout: limits ability to use a build system to dynamically find any dependencies and rebuild them * Submodules: can't atomically update both the parent and the child repo, have to manually update the referenced commit of the child repo in the parent repo, and each collaborator must manually update their child repo when the commit changes
- bastardoperator 5y agoVFS is being replaced in favor of https://github.com/microsoft/scalar https://github.com/microsoft/scalar
- WorldMaker 5y agoWhich scalar is "mostly" just a config tool for git sparse checkout of git partial clones with git commit-graph support turned on. All of that is stuff contributed directly into the git client. Beyond that "mostly", it also configures git lfs, which likely will always be a git plugin and not directly in the client and the rest of it seems like stuff Microsoft is testing before upstreaming it directly into the git client.
- maratc 5y ago> CI/CD inside a git monorepo without requiring a large checkout. this can also be solved by using a git mirror.
- jazzkingrt 5y agoDoesn't this just provide another (perhaps more nearby) remote? You still end up doing some kind of large checkout.
- 0xbadcafebee 5y agoThey definitely need a long FAQ. I'm sure it's different than submodules somehow, but in what ways/circumstances/purposes, I have no idea.
- zactato 5y agoI might be reading this wrong, but it seems like this creates a second place where a project would need to express dependencies between modules. You would need to do it for JOSH and whatever your build tool is. You'd probably need some additional git commit hooks to ensure your build tool of choice config is in sync with josh.
- rwarmka 5y ago> ... this creates a second place where a project would need to express dependencies between modules. You're right, but from my experience with using josh it really hasn't been a pain point. This, however, is coming from the context of having a build system which (appropriately) checks out all relevant josh workspaces as part of the monorepo verification build, and builds all of those workspaces along with building the relevant parts of the monorepo. Thus, if someone adds a new dependency in a workspace project build, the main monorepo verifier build will stop them from committing that change without also making sure the dependency is added to the workspace file, because the associated workspace build part of the verifier will fail (due to the missing dependency). Additionally, if you're working in a workspace and you add a new dependency, your own local checkout build will fail due to that missing dependency until you add it to the workspace file (and then get that new dependency). As long as you have builds and tests covering your workspace, it's pretty easy to figure out if you forget to add the dependency to the workspace file. Lastly, at least from my experience, it's overall been a pretty inconsequential price to pay, as new dependencies aren't added frequently after a project finishes its initial start-up phase.
- derefr 5y agoIANAMU (I Am Not A Monorepo User), but as far as I understand monorepos and their advantages, I'm not sure what the use-case for this tooling is. Most of the time, when an org chooses to move to having a monorepo (rather than just being left with one by accident of history), the key advantage they're striving to attain, is the ability to make changes to cross-cutting concerns across many distinct applications/libraries, with single commits/PRs. To change an API, and all of its internal callers, atomically, without having to worry about symbolically binding the two together with dependency version constraint resolution. Which is to say, the key advantage of a monorepo comes from having the whole monorepo checked out locally.
- bognition 5y agoThey are very useful when you've got hundreds of engineers working on distinct projects that are loosely connected. Imagine you have a team working on core libraries, a few product platform teams, and finally a team working on a customer facing feature. The feature team commits to master and a build kicks off (build 1), that build fails for an unrelated issue (say the build node dies) and the build kicks off again (build 2). With multiple layers of libraries supporting the feature team, its entirely possible that a dependency could have changed and the end result of build 2 would be different than build 1. When each commit shares a common timeline it is really easy to rebuild build 1 with the exact same dependencies.
- derefr 5y agoGit submodules already solve that problem, though. As does publishing all the core libraries etc. as language-ecosystem packages on a private package namespace or internal corporate package repository, and then resolving/locking the language-package dependencies to specific tags/refs with a lockfile that gets committed to the downstream repo. These are the "obvious" solutions to this problem, the first ones the average software architect would reach for. What would lead them to ignore these options and choose a monorepo instead, if not for what I mentioned above — the ability to make atomic changes to cross-cutting concerns?
- 5y ago
- MakersF 5y agoLike some people, I was expecting to find a way to have the advantages of a monorepo while having projects in separate repos. This is something Bloomberg is doing, and it's very cool. Each project is a separate repo, but they have a central integrated "repo" with all the repos, which is the "source of truth", and were code is built and deployed from. You can commit changes in your repo, and then you "release" the code into the integrated repo, which will rebuild all the transitive dependencies and run their tests to make sure everything still works. If anything fails, your release of the code is not merged in the repo. I'm now working with a monorepo, and I much prefer the Bloomberg approach. Cross repo changes can be made atomically (you update the reference in the integration repo for multiple individual repositories at once), and that is usually the big sell point of monorepo. And it doesn't have the downsides of the monorepo. The only issue is that it's not very ergonomic, and there isn't a tool to make that easy. But building such a tool is definitely easier than implementing a virtual FS as it has been done in multiple companies. I'd love if someone still working there were to write a nice post about that system, it was the first of such a kind I saw.
- choeger 5y agoBut that only works in one direction, no? So it works if you can develop your single repo, but it doesn't help you when you depend on other projects. I think the best approach would be to have bidirectional links between the projects (if A needs B, then A has the stable version of B and vice versa). The point in that setup would be that "upstream" projects can notice when they are about to break tests in "downstream" repos and act accordingly.
- MakersF 5y agoThey do, the repositories define dependencies, so when you change something everything that depends on you is rebuilt and tested. This prevents breaking changes, both for your dependencies and your reverse dependencies, identically to a monorepo. It's a bit complicated to explain, but it works. That's why I hope they'll make a blog post :)
- throwaway315724 5y ago
- deleted 5y ago[deleted]
- runawaybottle 5y agoMost of you don’t need a monorepo, the same way most of you don’t need, well, half the shit peddled in the tech Instagram (conferences, meetups, mediums, blogs, hn). You just don’t need that stuff, there’s like 20 of you on a team and at best your app probably sucks and barely has users, and if it does have users, it’s probably some trivial bullshit. You’re all a bunch of ordinary folks, so stop fucking up the workplace with your identity crisis. No, you are not an elite engineer, you are Bob, the guy who goes home every day and watches Netflix/plays video games.
- detaro 5y agoThe nice thing is, you can replace "monorepo" with "multiple repos" in your comment and it's just as believable a statement arguing for avoiding trouble with coordination/packaging etc.
- runawaybottle 5y agoWell, it’s something that needs to be said about a lot of things. Life is a balance and now days in tech I see the pendulum swinging way too far to the other side.
- sergiomattei 5y agoSurprisingly, a monorepo is much easier for smaller teams and individuals to work with. As someone who has been maintaining a React and Django app solo for the past three years, two repositories or more is too much cognitive overhead to work with. Never doing that again. Monorepo are easier for small apps.
- runawaybottle 5y agoPutting stuff in the same repo is a pragmatic idea. Going into the monorepo isolated self contained publishable app/package is a whole ‘nother thing, along with all the tooling necessary to make it work seamlessly. You want to put stuff in the same repo, that’s fine. What’s with all the other bullshit?
- JohnHaugeland 5y agoI have yet to hear anyone give me a coherent explanation of why a monorepo is better.
- icythere 5y agoThere are a few (re)solutions in the wind. The latest one that I've known is `west` (part of Zephyr-RTOS project), but I haven't tried yet. There may have wrong description (FIXME) but a sort list is found here: https://github.com/icy/git_xy#why https://github.com/icy/git_xy#why
- klysm 5y agoA lot of the comments here are surprisingly dismissive. I think having the ability to project parts of your git repo (still with a normal git api!) is an incredibly useful feature. Take the example of DefinitelyTyped: the maintainers can just do all the things they want in one repository which _vastly_ reduces the development overhead, but consumers of that code can use it however they please. If I understand correctly, you could have a submodule reference work out of the box but to a subset of that repo that you care about - seems pretty damn cool to me!
- chrschilling 5y agoWhat you are describing is one of the main use cases at ESR Labs (where Josh was created): For developers it is very convenient to work in a single tree. For reviewers and CI it is useful to look at the changes in a larger context. For consumers/integrators however it is useful to only look at parts of the code that have to be shipped to particular customers, as submodules(or the like) in their repos. Plus a lot of package managers assume library == repo as a default, so it is also easy to integrate with those while keeping monorepo processes for development.
- Meleagris 5y agoThis seems to provide the same functionality as using Git Sub Modules[1]. Am I getting the correct impression? [1] https://git-scm.com/book/en/v2/Git-Tools-Submodules https://git-scm.com/book/en/v2/Git-Tools-Submodules
- Dobbs 5y agoAssuming I understand what this is doing correctly, it does the reverse of a submodule. git submodules let you tack a second git repo onto an existing one. For example repoA tracks `repoB@version1234` at path `/foo/bar/baz`. This on the other hand lets you take monorepo and checkout `/go/mysubservice` as a "repo" and treat it as its own repo. Then when you do git pushes etc, it translates the changes into the larger monorepo.
- andix 5y agoThat looks nice. It must have some major drawbacks? Sounds too good to be true…
- oofbey 5y agoIt’s definitely more complex than just using a monorepo. This tool that all your code runs through is young and not well supported.
- codetrotter 5y agoFrom the title I expected it to be a tool for treating multiple separate repos as though they were all just one single monorepo. But from the description in the README, it seems to be for treating subsets of a monorepo as though they were separate repositories. PS: The title, in case it is changed, is currently “josh: Get the advantages of a monorepo with multirepo setups”
- oftenwrong 5y agoI apologise for the title. HN has a short limit for title length, so I came up with my own title. I thought this title did a decent job of presenting, using short language, the main application that the authors gave top-billing in the README. I am not affiliated with the project. JOSH claims to be reversible, so it could be used in either direction, which is where the multiple use cases come in. Treating subsets of a repo as their own repo, or treating multiple repos as one. I would say there is some application overlap between this and git submodule/subtree/subrepo and also tools like copybara.
- tusharsadhwani 5y agoYou might want to check out https://github.com/asottile/all-repos https://github.com/asottile/all-repos :)
- lbhdc 5y agoThat was what I expected from the description as well. After reading the readme it's not clear to me what problem this is trying to solve, and why this is the solution.
- ganafagol 5y agoThe problem is that large codebases tend to have a huge footprint if you need to clone the whole repo. Git as-is does not allow you to only pull a subset, i.e. specific paths representing a sub project. That's what josh is trying to solve: a "virtual" repo that behaves like a real git repo but behind the scenes seemlessly integrates with the big monorepo.
- nemetroid 5y agoHow does this compare to partial clones and sparse-checkout? This question was raised as an issue in the project but was closed as "not really an issue", which I guess is true. https://github.com/esrlabs/josh/issues/23 https://github.com/esrlabs/josh/issues/23
- villasv 5y agoTechnically correct, but so unhelpful. No way I'm using a project that has this kind of community engagement.
- sjburt 5y agoGenerally, Github Issues are used as a bug tracker, not a community FAQ. Asking a maintainer to compare and contrast their project with another project or git feature seems a bit demanding.
- villasv 5y agoKeyword: generally. Plenty of projects do allow community questions, specially small ones or early stages. Is there anywhere else to ask that question? If there is, it isn't prominently signified, answering at least "ask this >here<" would be common sense. At a minimum, this issue evidences the need of documentation and should be addressed in some way with more than "I don't have to answer this".
- villasv 5y agoBefore the traditional HN comments "isn't this just ...?", the README already does that for you: > a blazingly-fast, incremental, and reversible implementation of git history filtering
- taeric 5y agoThe problem with trying to force one single history, is that it ignores deployments. And user onboarding/behavior change. All of which can be relevant when working on a project. With multi repository projects, this helps some thinking, as it is clear that the changes were not atomic between systems. They are literally separate at all layers, including the commit. I sympathize with wanting a simpler view. I'm just worried on an inflated value proposition.
- klodolph 5y ago> The problem with trying to force one single history, is that it ignores deployments. I'm not sure what, exactly, the problem is that you're talking about. When you deploy something, it's built from a specific commit. Deployment is not atomic, it's a gradual process that takes some amount of time. Between when you (or your automation) chooses to deploy a system and when the deployment finishes is some window of time. The system will often spend much of that time in a partially updated state. You may also choose to canary changes, so you will have a mix of different versions in production at any given time. At companies where I've worked, the time from deployment start to finish for backend systems ranges anywhere from hours to weeks. I don't understand how this relates to multi-repo or mono-repo concerns, however. The repo is a history of the source code (intentional changes by humans), it's not a history of the state of your production systems (which are the results of automation).
- taeric 5y agoMany of the mono repo guides I see are off the "all projects in one repository" kind. If the project doesn't build and deploy as an atomic unit, then having atomic code changes is a foot gun. It is amazing how many times I've seen folks think that just because they can get build time tests happy with changes in two projects, that they can safely send out the two changes. Does a multi repo "solve" this? Of course not. But it is easier to reason that two projects clearly need two deploys. Versus having to remember that one commit could be N project deployments.
- klodolph 5y ago
- Glavnokoman 5y agoHaving been to the both sides of it I can say there exist exactly 0 advantages of a monorepo setup.
- andix 5y agoSimplicity is always an advantage.
- indymike 5y agoYou can have your simplicity one of two ways: * Devops - one repo, one deploy, let the developers figure out which repo is which. * Developer - one codebase, let the devops people sort it out and write lots of tooling to make my monorepo work. If this stuff was easy, everyone would be a developer.
- zaphar 5y agoI have also been on both sides and I'll say that each side has different sets of advantages highly dependent on a host of factors a non exhaustive list of which is: * Culture * CI/CD tooling support * Codebase sizes
- geschwindner 5y agoAt our startup, we chose to start with a monorepo. Our team is small, but one of the big advantages we’ve had so far is avoiding the n*m (for n services and m tools) problem with dev tooling - which leads to a very smooth developer experience. For example, to run one or more services locally, we use a single script that sits at the repo base - ‘dev.sh service1,service2,...’. This avoids a lot of headaches for our developers, as we enforce compliance when adding a project to the repo. Lint config? One to rule them all. Test coverage thresholds? Single one. This consistency is the biggest win in my opinion. Similarly, our integration tests are very easy to write without commit skew. Finally, sharing libraries has been painless - since we have common/ and common/third_party/ directories at the monorepo root.
- mdtusz 5y ago
- DethNinja 5y agoThis can be kinda achieved with git submodules. I’m currently using a monorepo with submodules and it works really well. Dependency management is not a huge issue too, at least if you don’t have hundreds of submodules.
- tazjin 5y agoIt can not - josh can do many arbitrary kinds of (reversible) transformations on the repo, which allows you to have different external "projections" of your monorepo. Imagine a company that develops strongly interdependent software in a monorepo, but needs to publish different subsets of this software to external entities which also expect a coherent version history.
- mugsie 5y agogit filter-branch would fit this use case?
- chrschilling 5y agoOn a basic level, yes, both Josh and git filter-branch do essentially the same thing. The difference being that Josh is much faster not just compared to git filter-branch but also compared to all the other similar tools out there, especially when run repeatedly in the same repo. Also being a server it does not require any installation or resources on the developers machine. In addition to that over time more features where added that git-filter branch does not have, most notably "josh workspaces" which is a DSL for repo transformations.