4 ms·
Can you help me understand what bit rot means as you are using it, and why it is different in a mono repo vs an equally large multi-repo system?
by wcdolphin 7y ago
Can you help me understand what bit rot means as you are using it, and why it is different in a mono repo vs an equally large multi-repo system?
- bunderbunder 7y agoJust cherry-picking one rather extreme example: There was a method in a core library that had an argument that was not used. Few people knew this, and it was one that was called at least once by most services. When I discovered this, I asked the module's maintainer about it. He explained that it used to be used, but requirements changed so that it wasn't needed anymore. He decided to leave it anyway, though, because the alternative would have been an immense effort to track down all the call sites, remove the argument, then trace all the resulting dead code, and remove that as well. It would have eaten up a week, I'm sure. Everyone had bigger fish to fry, and always would have bigger fish to fry. As I understand it, Google gets around that problem by having a group of people where frying fish like that is their primary responsibility.
- sa46 7y agoI don't understand how avoiding bitrot is better in a multi-repo setup. With a multi-repo, if you're changing a core-library you can make the change quickly but you can't upgrade clients until they upgrade on their own. You still have bitrot. The difference is that the bitrot isn't visible because it's compartmentalized into separate repos.
- andolanra 7y agoThe reason that a monorepo facilitates this kind of bit-rot is because of one of the advantages of the monorepo setup: that it also facilitates easily making connections between different parts of the codebase. In a multi-repo setup, you need to be intentional about when and how you pull in another repo, because the bar to doing so is so high. In a monorepo, every piece of every part of the code is available to you and no extra action is required to "pull in" that other repo, which means in practice that without tooling or discipline you end up with a codebase with lots of tight interconnections. So the reason it happens in a monorepo more is that those connections are easier to make and therefore a lot more common, not because those connections are impossible in a multi-repo setup.
- sa46 7y agoI see, the argument is that monorepo facilitates tight-coupling because it's easy to add dependencies. The multi-repo doesn't have the same problem because it's more difficult to have cross module dependencies. There's a number of solutions to prevent tight-coupling. 1. Language-level visibility modifiers. 2. Build system visibility. Bazel offers fine-grained visibility. I'm not sure about others. 3. Using RPCs as the primary interconnect between services. My view is that disadvantages of tight-coupling in a mono-repo is a much better problem to have than the disadvantages of having logic spread across different repos.
- bunderbunder 7y agoSo, to me, different things using different versions of the same library isn't inherently bitrot. For one, if you're actively maintaining your code, then everything's going to eventually end up upgraded. It's just a question of whether you want to boil the ocean all at once, in one big pot, or do it in a series of smaller batches. There are legitimate arguments for both options. The thing I'm trying to say is, you should beware the idea that going monorepo means the ocean's going to just boil itself somehow.
- joshuamorton 7y ago> As I understand it, Google gets around that problem by having a group of people where frying fish like that is their primary responsibility. I'd rephrase this: Google gets around this problem by having tooling (one such tool is the monorepo) that makes fixing issues like these fairly simple. A straightforward example like this one could be fixed with sed/clangtidy/refaster (depending on the language) + Rosie[1] in a few days, even if there were thousands (or 10s of thousands) of uses. But the work of upgrading is borne by the person doing the upgrade, not client teams who suddenly encounter compile errors when they try to version bump. This is much "fairer". [1]: https://cacm.acm.org/magazines/2016/7/204032-why-google-stores-billions-of-lines-of-code-in-a-single-repository/fulltext https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...
- closeparen 7y agoThis is exactly why we are adopting a monorepo: it makes tasks like this tractable. Across thousands of microrepos, they are not.
- user5994461 7y agoFunny you mention this as a drawback, while it's actually a strength of mono repo that is perfectly handled. For a widely used library. Leave the argument. This argument could be defaulted to NULL, the function can be marked @deprecated and it can log a warning message that the argument has no effect (depending on the language). That's how to handle updates with backward and forward compatibility. No need to break user code. That's a strength of a mono repo. It only lets you do the right thing. You can't ignore other developers in the company and break all their software. In a typical multi repo, there is the luxury to change anything anytime, as often as a developer feels like renaming a variable or a parameter. Creating a ton of unnecessary work for other developers who have the pain to keep up. That is, if they ever attempt to upgrade any library, there is no benefit to upgrading and it's impossible to keep up with all changes going on in all libraries thus it's rarely done in practice. The more it's behind, the more work there is to upgrade, the less likely it is to be done.
- tsss 7y agoI don't get how people can live like this. In any decent language with a primitive static type system this would take you 15 minutes at most and you could be sure that it works afterwards.