8 ms·
Scaling Mercurial at Facebook (2014)
- samfisher83 6y agoWhy not just have bunch of repos? You can use what google does with chrome and use depot tools to checkout from a bunch of repos. Edit: Why am I getting downvoted for asking a question? We use multiple repos and we have a pretty large codebase consuming a TB a day.
- q3k 6y agoBecause you lose the 'single head' concept of actually being able to always have a clear view of the current newest version and of linear version history. The ability to map a single revision number into the state of _all_ source code (including third party dependencies) is extremely powerful. You also lose the ability of performing sweeping changes across an entire codebase in lockstep (think: backwards compatibility breaking API changes). Not to mention that Google internally avoids multiple repositories (and instead runs of a single monorepo) - it's just the android/chrome/public stuff that's mostly split up.
- AlphaSite 6y agoAlso there isnt really a lot of great tooling for making sweeping changes across repos.
- erik_seaberg 6y agoDon't make breaking API changes, you can't deploy them atomically. You have to support the old API until everyone has safely migrated and confirmed no rollback will be needed.
- joshuamorton 6y agoThere are more apis than just rpc ones. If I change my library's api, I can update the callers atomically.
- fsociety 6y agoThen services would be stuck in a sea of technical debt. There are too many moving pieces to do that.
- erik_seaberg 6y agoTechnical debt is a tool to be used (with caution). Getting ludicrously far behind is a risk I'd avoid, but our stability is more important than staying on the bleeding edge and testing every single version of each of our dependencies.
- djohnston 6y agoyou can if it's a monorepo and all your callers are in the monorepo, which sort of answers OP's question
- erik_seaberg 6y agoThe commit can be atomic, but if prod is many machines, the deployment cannot.
- throwdbaaway 6y ago> performing sweeping changes across an entire codebase in lockstep This thinking right here, sounds like a reason why Google retires a lot more services at much higher frequency compared to other companies? In https://thehftguy.com/2019/12/10/why-products-are-shutdown-the-case-of-verizon-and-yahoo-groups/ https://thehftguy.com/2019/12/10/why-products-are-shutdown-t..., the HFT guy brought up the "5 years upgrade pain" as the reason why a service would be retired, at the time when the pain from upkeeping outweights the gain from revenue. By regularly making backward incompatible API change, as afforded by the monorepo, while also having engineers moving freely between projects, the upgrade pain cycle becomes way shorter for Google services. By the way, this line of thought also leaks into public facing source code, with guava being the poster child of breaking backward compatibility, althought it has learned its lesson starting with version 21.
- ardit33 6y agoWe tried at Spotify the multi-repo thing and it was a total nightmare. Once you have lots of people working on a product, it becomes very hard to do changes on a multi repo setup. The alternatives: 1. Completely de-coupled code. I.e. teams ship libraries, like a 3rd part api. This slows down development considerably, and makes re-use of code harder. 2. Keep everyone on one repo, and allow to changes stream in. 1. Slows down dev. (it suddenly becomes more beourocratic), 2. requires scaling your code versioning tool (gir/mercurial, etc) and processes associated with it. Also, single repos, make sense in one domain (e.g. server side, ios, android, etc...). You can have different repos for different 'domains' where code doesn't intersect with each other that much.
- dehrmann 6y ago> makes re-use of code harder I've found this to be a feature of polyrepos because monorepos can easily become a rat's nest of dependencies. Polyrepos make you think harder about what should really be exposed and shared.
- jonhohle 6y agoA phrase I like to use sums this up: “dependencies are easy to add, but hard to remove.”
- joshuamorton 6y agoBuild systems can enforce visibility, which solves this. Allowing fine grained visibility at the level of a file, package, or artifact is better than at the granularity of a repo.
- dehrmann 6y agoDepends how easy it is to change visibility. My point is polyrepos make things you need to be careful about like adding internal dependencies and API changes hard. Monorepos make those deceptively easy.
- 6y ago
- klodolph 6y agoMultiple repos... the most bureaucratic and painful way to scale your code base. Having seen "a bunch of repos" in action, I think people don't talk enough about just how awful the experience can be, and how much more work it is to manage multiple repos. As you increase the number of repos, the pain gets worse at a rate which is faster than linear. There are plenty of articles talking about how wonderful a monorepo is, just not many articles about how bad multirepo is. NPM hell is a close approximation of the multirepo experience. Try upgrading the dependencies in a large NPM project and you'll see all sorts of problems. You might find that upgrading X breaks Y, but you have to upgrade X in order to upgrade Z, and you need to upgrade Z for some reason. With monorepo, all the versions march forward in sync. If you fix trunk, you will probably ship it, eventually. With multirepo, you need to fight the tooling just to show (from the example above) that a patch to Y will let you upgrade your project to use a newer version of Z.
- an_opabinia 6y agoGiant companies consistently rather vendor the open source code they ingest rather than give back. Of course Google, Facebook etc have given back a lot to OSS. Just not extemporaneously to when they received the value. It may be years until the internal rebuild of something some guy copied from somewhere is rereleased. Git was designed for OSS. That includes its radically transparent form of development. Then again with submodules you can easily vendor your private stuff, rather than doing things the other way around.
- CJefferson 6y agoI have sympathy -- I "vendor" some libraries as I need to change something deep in the library. It's often not nice, and not something upstream would want.
- joshuamorton 6y agoThis isn't really true. The Google oss policies (opensource.google) require that you use a recent version. As a result upatreaming functional changes is encouraged and easier than maintaining a local patch set. Sometimes there are local changes but they're usually non functional stuff to make the library integrate with google infra. They do vendor, but that's for security and efficiency reasons mostly, not to avoid giving back.
- robocat 6y agoDan Luu gives a good summary of benefits of monorepo: http://danluu.com/monorepo/ http://danluu.com/monorepo/
- azangru 6y agoPOSTED ON JAN 7, 2014
- jinwoo68 6y ago2014
- deleted 6y ago[deleted]
- jeffbee 6y agoAnyone know if they're still on Mercurial?
- deleted 6y ago[deleted]
- Shish2k 6y ago“Yes”, but with custom server, custom virtual filesystem layer, and heavily-tweaked client — https://github.com/facebookexperimental/eden https://github.com/facebookexperimental/eden
- ngoldbaum 6y agoThey also stopped participating in upstream mercurial development, google is pretty much the only major corporate contributor and steward of the project at this point.
- sulam 6y agoI can't say for sure, but I'd be incredibly surprised if they weren't. What else would they be on? (Don't say Git, they tried to make Git scale for their monorepo and it would have required a lot of changes that the Git community wasn't interested in.)
- adamnemecek 6y agoPerforce is still popular.
- klodolph 6y agoI think FB and Google both surpassed Perforce's scale a long time ago. I've heard all these stories about build servers or source control servers where, for the longest time, the solution to scalability problems was to throw larger machines at it. At some point the machines have well into the TBs of RAM, each, and it becomes a major priority to make the system distributed.
- setheron 6y agoObjectively great contributions to OSS and mercurial but I wonder how many companies or projects will need to do something similar. I enjoy the patches the make mercurial usable without the crazy extensions that few of us will ever use.
- tikhonj 6y agoNote: this article is from 2014; it would be interesting to see how Mercurial has worked for Facebook since then.
- nuclear_eclipse 6y agoIt's a lot better than it was in 2014, but I still miss using git.
- ForHackernews 6y agoWhy? Git is such a worse, over-engineered experience. Github won, so we're all forced to use git now, but mercurial is really a simpler interface and a more straightforward mental model.
- nuclear_eclipse 6y agoBecause I really like the model/UX of the staging area. I'm not worried about the tool being simple; I use it every day, so it's worth learning inside and out. It's a much nicer experience IMO to be able to work on various different pieces, then iteratively run `git add -p` or similar to stage individual files or hunks, verify everything I'm adding with `git diff --cached`, then `git commit` when I'm actually ready and happy. It lets me fix multiple parts of my project at the same time, without interruption, while still letting me compose my batch of changes into multiple, self-contained changesets. `hg record` feels really limited by comparison, and rushes you immediately into committing your selected changes, and then further makes it difficult to selectively amend more (partial) changes to that commit afterwards.
- greggman3 6y agoI started on mercurial. I vastly prefer git. I admit it took a few months of asking teammates to help me get it but i'd never go back.
- throwdbaaway 6y agoSame here. I used to regard http://jordi.inversethought.com/blog/i-hate-git/ http://jordi.inversethought.com/blog/i-hate-git/ as my hero, and then I had to start using git in a new gig, and never touched mercurial since. If the goal of using a vcs is to have meaningful atomic commits, nothing beats git staging area and interactive rebase.
- based2 6y agohttps://github.com/facebookexperimental/eden https://github.com/facebookexperimental/eden https://news.ycombinator.com/item?id=23124095 https://news.ycombinator.com/item?id=23124095
- tehlike 6y agoAs an engineer that worked both at google and Facebook, I vastly prefer googles monorepo on perforce. Combining that with citc was pretty solid way to develop your project. Hg is a bit of a nightmare in the wfh situation. Really slow, hangs for a long time if you haven't synced in a few days. Yes, im sure there are ways to tweak, but not sure if you can tweak them enough!
- liveoneggs 6y agoI can only imagine an attempt at the initial clone.. the hg client is ridiculously slow, memory-hungry, and full of dumb behavior that seems to assume local connections
- est31 6y agohg has a mode to make initial clones faster by offering zipped compressed checkouts for download.
- liveoneggs 6y agoit does range requests to these bundles, attempts to extract the partials while keeping the http connection alive using up GB of memory in the process, and then totally fails pretty regularly when it need to go back to the network. It makes no attempt at recovering those chunks, re-establishing connections, or using any intelligence at all. It also overwhelms http header size norms on the regular and runs up memory when attempting to use those bundles. Chunks/bundles are, in fact, strictly worse unless you get your initial clones with a good piece of software like curl.
- Lt_Riza_Hawkeye 6y agoSomehow FB has made all the common operations work in <6 seconds with millions of files...
- gravypod 6y agoIt's an impressive feat but something I find even more impressive is the scale they need to be prepared for. The performance challenges of maintaining infrastructure for as many developers as Facebook/Google employ is constant and demanding. At my extremely small startup in the past ~6 months I've made about 1,000 files (@ HEAD, not including deletions). So, maybe, at the absolute worst case each engineer could produce ~150 files/month. Facebook likely has ~1,000 engineers. That's a worst case 1,800,000 files/year + an existing massive code base. Good luck to the engineers who are forever battling the scaling issues underlying these orgs! It's extremely impressive to even hear about overcoming these huge hurdles.
- qeternity 6y agoIn 2014, a large site but not the behemoth it is today, how did Facebook manage to clock in at 17 million LOC?? Genuinely curious for anyone with first party experience.
- dang 6y agoIf curious see also 2016 https://news.ycombinator.com/item?id=11789182 https://news.ycombinator.com/item?id=11789182 Discussed at the time: https://news.ycombinator.com/item?id=7019673 https://news.ycombinator.com/item?id=7019673
- rektide 6y agoDoes anyone have any good videos about Facebook's use of watchman to ship changes to central systems & be performing continuous test/analysis on code as it is developed? I feel like I saw a good video in 2017 or so but cannot find much these days.