5 ms·
And that way, you can't have atomic updates across the repositories and need to synchronize them all the time, great.
by Orphis 5y ago
And that way, you can't have atomic updates across the repositories and need to synchronize them all the time, great.
- swiley 5y agoYes you can, it happens when you bump the sub module reference. This is how reasonable people use git.
- Denvercoder9 5y agoSubmodules often provide a terrible user experience because they are locked to a single version. To propagate a single commit, you need to update every single dependent repository. In some contexts that can be helpful, but in my experience it's mostly an enormous hassle. Also it's awful that a simple git pull doesn't actually pull updated submodules, you need to run git submodule update (or sync or whatever it is) as well. I don't want to work with git submodules ever again. The idea is nice, but the user experience is really terrible.
- mdaniel 5y agoAnd woe unto junior developers who change into the submodule directory and do a git commit, then made infinitely worse if it's followed by git push because now there's a sha hanging out in the repo which works on one machine but that no one else's submodule update will see without surgery I'm not at my computer to see if modern git prohibits that behavior, but it is indicative of the "watch out" that comes with advanced git usage: it is a very sharp knife
- fpoling 5y agoLooking back I just do not understand why git came up with this awkward mess of submodules. Instead it should have a way to say that a particular directory is self-contained and any commit affecting it should be two objects. The first is the commit object for the directory using only relative paths. The second is commit for the rest of code with a reference to it. Then one can just pull any repository into the main repository without and use it normally. git subtree tries to emulate that, but it does not scale to huge repositories as it needs to change all commits in the subtree to use new nested paths.
- dylan-m 5y agoOr define your interfaces properly, version them, and publish libraries (precompiled, ideally) somewhere outside of your source repo. Your associated projects depend on those rather than random chunks of code that happen to be in the same file structure. It's more work, but it encourages better organization in general and saves an incredible amount of time later on for any complex project.
- throwaway894345 5y agoI don’t like this because it assumes that all of those repositories are accessible all of the time to everyone who might want to build something. If one repo for some core artifact becomes unreachable, everyone is dead in the water. Ideally “cached on the network” could be a sort of optional side effect, like with Nix, but you can still reproducibly build from source. That said, I can’t recommend Nix, not for philosophical reasons, but for lots of implementation details.
- pyrale 5y agoUsing the expression "that's how reasonable people do ..." is not a great conversation starter. I've always had a bad experience using submodules, they're imo the poor developer's versioning tool. It's useful when you use a language without a good build/packaging tool, but otherwise, I'm better off leaving the language-specific tool fetch the depended code.
- deleted 5y ago[deleted]
- slver 5y agoWe have repository systems built for centralized atomic updates, and giant monorepos, like SVN. Question is why are we trying to have Git do this, which was explicitly designed with the exact opposite goal? Is this an attempt to do SVN in Git so we get to keep the benefits of the former, and the cool buzzword-factor of the latter? I don't know. Also when I try to think about reasons to have atomic cross-project changes, my mind keeps drawing negative examples, such as another team changing the code on your project, is that a good practice? Not really. Well unless all projects are owned by the same team, it'll happen in a monorepo. Atomic updates not scaling beyond certain technical level is often a good thing, because they also don't scale on human and organizational level.
- alexhutcheson 5y ago1. You determine that a library used by a sizable fraction of the code in your entire org has a problem that’s critical to fix (maybe a security issue, or maybe the change could just save millions of dollars in compute resources, etc.), but the fix requires updating the use of that library in ~30 call sites spread across the codebases of ~10 different teams. 2. You create a PR that fixes the code and the problematic call sites in a single commit. It gets merged and you’re done. In the multi-repo world, you need to instead: 1. Add conditional branching in your library so that it supports both the old behavior and new behavior. This could be an experiment flag, a new method DoSomethingV2, a new constructor arg, etc. Depending on how you do this, you might dramatically increase the number of call sites that need to be modified. 2. Either wait for all the problematic clients to update to the new version of your library, or create PRs to manually bump their version. Whoops - turns out a couple of them were on a very old version, and the upgrade is non-trivial. Now that’s your problem to resolve before you proceed. 3. Create PRs to modify the calling code in every repo that includes problematic calls, and follow up with 10 different reviewers to get them merged. 4. If you still have the stamina, go through steps 1-3 again to clean up the conditional logic you added to your library in step 1. Basically, if code calls libraries that exist in different repos, then making backwards-incompatible changes to those libraries becomes extremely expensive. This is bad, because sometimes backwards-incompatible changes would have very high value. If the numbers from my example were higher (e.g. 1000 call sites across 100 teams), then the library maintainer in a monorepo would probably still want to use a feature flag or similar to avoid trying to merge a commit that affects 1000 files in one go. However, the library maintainer’s job is still dramatically easier, because they don’t have to deal with 100 individual repos, and they don’t need to do anything to ensure that everyone is using the latest version of their library.
- cryptica 5y agoIf the project has good separation of concerns, you don't need atomic updates. Good separation of concerns yields many benefits beyond ease of project management. It requires a bit more thought, but if done correctly, it's worth many times the effort. Good separation of concerns is like earning compound interest on your code. Just keep the dependencies generic and tailor the higher level logic to the business domain. Then you rarely need to update the dependencies. I've been doing this on commercial projects (to much success) for decades; before most of the down-voters on here even wrote their first hello world programs.
- iudqnolq 5y agoWhat do atomic source updates get you if you don't have atomic deploys? I'm just a student but my impression is that literally no one serious has atomic deploys, not even Google, because the only way to do it is scheduled downtime. If you need to handle different versions talking to each other in production it doesn't seem any harder to also deal with different versions in source, and I'd worry atomic updates to source would give a false sense of security in deployment.
- status_quo69 5y ago> If you need to handle different versions talking to each other in production it doesn't seem any harder to also deal with different versions in source It's much more annoying to deal with multi-repo setups and it can be a real productivity killer. Additionally, if you have a shared dependency, now you have to juggle managing that shared dep. For example, repo A needs shared lib Foo@1.2.0 and repo B needs Foo@1.3.4, because developers on team A didn't update their dependencies often enough to keep up with version bumps from the Foo team. Now there's a really weird situation going on in your company where not all teams are on the same page. A naiive monorepo forces that shared dep change to be applied across the board at once. Edit: In regards to your "old code talking to new version" problem, that's a culture problem IMO. At work we must always consider the fact that a deployment rollout takes time, so our changes in sensitive areas (controllers, jobs, etc) should be as backwards compatible as possible for that one deploy barring a rollback of some kind. We have linting rules and a very stupid bot that posts a message reminding us of that fact if we're trying to change something sensitive to version changes, but the main thing that keeps it all sane is we have it all collectively drilled in our heads from the first time that we deploy to production that we support N number of versions backwards. Since we're in a monorepo, the PR to rip out the backwards compat check is usually ripped out immediately after a deployment is verified as good. In a multi-repo setup, ripping that compat check out would require _another_ version bump and N number of PRs to make sure that everyone is on the same page. It really sucks.
- pyrale 5y agoIt's not just atomic updates, the critical part is to run all the relevant integration jobs. The aim is that when a change goes to the reference branch, you must understand what other code is going to get impacted. Without that, you can break builds depending on the repo you changed, and if the broken repos don't change often, the breakage could be discovered months later. If you're in that situation, you'll randomly discover piles of debt when you use a slow-moving repo. Atomic deploys are not as important, because you can still decide to version your APIs or releases even if you're using a monorepo. That being said, you can use multiple repos and still mostly avoid trouble by choosing how to cut your codebase (HR software is likely not going to depend heavily on presale, for instance). The metric to optimize is to minimize the required version bumps.