10 ms·
Facebook hit git performance issue on large repository
- deleted 15y ago[deleted]
- courtewing 15y agoThis was actually pretty fascinating to me. On one hand, I am astonished at how long it takes to perform seemingly trivial git operations on repositories at this scale. On the other hand, I'm utterly mystified that a company like Facebook has such monolithic repositories. Even back when I was using SVN a lot, I relied on externals and such to break up large projects into their smaller service-level components. I'd be very interested to see some benchmarks on their current VCS solution for repositories of this scale.
- guan 15y agoFrom a followup post: “We already have some of the easily separable projects in separate repositories, like HPHP. If we could split our largest repos into multiple ones, that would help the scaling issue. However, the code in those repos is rather interdependent and we believe it’d hurt more than help to split it up, at least for the medium-term future. We derive a fair amount of benefit from the code sharing and keeping things together in a single repo, so it's not clear when it’d make sense to get more aggressive splitting things up.”
- ntkachov 15y agoRemember that this is a synthetic repository.
- vessenes 15y agoHe notes that these repositories are somewhat broken up already, and wants to keep them together. There are good reasons to keep code in one repository; particularly, git's submodule support has a number of nasty interface tradeoffs; I wouldn't say it breaks git, but you have to keep a clear understanding of all your submodules in your head when you have a lot of them. OK, it pretty much breaks git to have submodules that are interdependent. I know this because I am currently moving one of my organizations off this exact plan -- it's the opposite of useful and speedy to have to worry about versions across a large number of backend / frontend repositories. It is MUCH easier and therefore better for developers to put them together, and release together.
- chucknthem 15y agoNot sure if it's still the case, but Google hosts all their internal source code on a modified version of perforce, so they essentially have everything in one repo.
- wmf 15y agoGiven that Facebook is compiled into a single 1 GB executable, a git repo with 1.3 M files doesn't really surprise me.
- jlarocco 15y agoWhat? Do you have a reference for that?
- nbm 15y agoFor that and other info, check out the "Push" Tech Talk given last year by the Facebook release engineering team's leader, Chuck Rossi: https://www.facebook.com/video/video.php?v=10100259101684977 https://www.facebook.com/video/video.php?v=10100259101684977
- nbpoole 15y agohttps://github.com/facebook/hiphop-php/wiki/Running-HipHop https://github.com/facebook/hiphop-php/wiki/Running-HipHop It's how HPHP works.
- reid 15y agoFacebook Engineering posted a video about their build process in May 2011. Seek to 25:55 for the source: https://www.facebook.com/video/video.php?v=10100259101684977&oid=9445547199 https://www.facebook.com/video/video.php?v=10100259101684977...
- huytoan_pc 15y agoHere you go: http://www.facebook.com/note.php?note_id=10150121348198920 http://www.facebook.com/note.php?note_id=10150121348198920 "We can build a binary that is more than 1GB (after stripping debug information) in about 15 min, with the help of distcc. Although faster compilation does not directly contribute to run-time efficiency, it helps make the deployment process better."
- xxpor 15y agoOh God, it's like Amazon was 10 years ago.
- functionform 15y agoWhat do Facebook and the National Institutes of Health have in common? I'm pretty sure this will end with Facebook building their own versioning system from scratch and give it some kitchsy name like "Retro".
- djtriptych 15y agoI hope these guys do take the route of developing a large-scale performant patch. Git as so many interesting uses at scale as just a tool that navigates and tracks DAGs over time.
- djb_hackernews 15y agoThat's projected growth for two of their projects. Sounds like they have something brewing... Still amazed that breaking it up would do more harm than good when the code isn't even written yet...
- Judson 15y agoI read it as them being unable to break up a project, and the repo being a projection of future commits to a project that can't be split up.
- yuvadam 15y agoWhile I'd be interested in seeing this issue further unfold, just the prospect of a 1.3M-file repo gives me the creeps. I'm not sure what the exact situation at Facebook is with this repository, but I'm positive that if they had to start with a clean slate, this repo would easily find itself broken up into at least a dozen different repos. Not to mention the fact that if _git_ has issues dealing with 1.3M files, I wonder what other (D)VCS they're thinking of as an alternative that would be more performant.
- wbkang 15y agoPretty sure Perforce performs fine with that.
- yuvadam 15y agoWell, of course that at some specific scale, you're gonna start to have trouble with any DVCS maintaining a complete local copy of such a huge repository. It's even worse that just disk space and performance issues. I can totally imagine a huge, busy repository where by the time you've pulled and rebased/merged your stuff, the repo has already been committed to again, invalidating your fast-forward commit and forcing you to pull again and again before you have any chance of pushing back your changes. This is an inherent problem with DVCS that just can't be solved (trivially) when working on huge repositories that span millions of files and involve thousands of developers.
- AndrewDucker 15y agoIf you read further down the thread they say that that's already had the non-interlinked files split out. What they've got left isn't easily broken up.
- clord 15y agoBest argument I've ever seen for not wanting to work at Facebook... wow that's a lot intertwined spagetti code. Our source repo at work (a C++ compiler with full commit history going back to the early 90s...) is smaller and more componentized!
- dpcx 15y agoI don't want to imagine the actual kind of code that requires 1.3M files to run.
- deleted 15y ago[deleted]
- nbm 15y agoKeep in mind that your average repository doesn't only contain code that is compiled and executed (or interpreted), there is also documentation, static assets such as images (that may be processed), configuration, computed files (that may make sense to pre-compute once rather than compute on a hundred people's environments every build), and so forth. Also, it doesn't only include the current file set - they include files that have been deleted, been split into modular files, been merged, been wholesale rewritten, put into a new hierarchy (some VCS systems handle this better than others). (I work at Facebook, but not on the team looking into this stuff. I'm a happy user of their systems though. Keep in mind that the 1.3 million file repo is a synthetic test, not reality.)
- 0x0 15y agoAlso, it doesn't only include the current file set - they include files that have been deleted, been split into modular files, been merged, been wholesale rewritten, put into a new hierarchy (some VCS systems handle this better than others). The follow-up email still mentions a working directory of 9.5gb. I cannot fathom working on a code repository consisting of 9.5gb of text. There must be something else going on here, even considering any peripheral projects like the iOS and android apps, etc. (edit: if there are huge generated files intermingled with code, shouldn't those be hosted on a "pre-generated cache" web server instead of git, for example?)
- eropple 15y agoOur codebase at my employer currently hovers around 5GB in SVN. Binaries and other generated code are intermixed for historical reasons. Removing them is a non-starter due to the amount of time it'd take to do so; the best solution I've been able to come up with so far is to break out into multiple SVN repos (one for images, one for generated language files, etc.) and then, hopefully, get code into Github while externally using the SVN repos for stuff that shouldn't be versioned in a distributed manner (versioning that stuff is useful as a convenience - avoiding conflicts, etc.).
- sek 15y agohttp://thread.gmane.org/gmane.comp.version-control.git/189776 http://thread.gmane.org/gmane.comp.version-control.git/18977... They keep every project in a single repo, mystery solved. Edit: > We already have some of the easily separable projects in separate repositories, like HPHP. Yeah, because it makes no sense, it's C++. They probably use for everything PHP i assume then. Is there no good build management tool for it?
- masklinn 15y ago> They keep every project in a single repo, mystery solved. That's not true: > It is based on a growth model of two of our current repositories (I.e., it's not a perforce import). We already have some of the easily separable projects in separate repositories, like HPHP. If we could split our largest repos into multiple ones, that would help the scaling issue. However, the code in those repos is rather interdependent and we believe it'd hurt more than help to split it up, at least for the medium-term future. They already have multiple repositories, the stats they're doing there is based on "two of [their] current repositories" implying more than two.
- sek 15y agoWhy would he take HPHP as an example then? It should be obvious that there is not much interdependence with the other code. Sounds to me like this: http://thedailywtf.com/Articles/Enterprise-Dependency-Big-Ball-of-Yarn.aspx http://thedailywtf.com/Articles/Enterprise-Dependency-Big-Ba...
- dblock 15y agoOthers have tried and keep throwing more and more smart people at the problem they just shouldn't have. MSFT with Windows codebase that runs out of several labs. Crazy branching and merging infrastructure. They use source-depot, originally a clone of perforce. Google with all their source code in one Perforce repo. Facebook will be on perforce before we know it. The solution is an internal Github, not one giant project.
- sek 15y agoGoogle has everything in one Perforce repo? You mean the search engine, do you? I agree btw, the Github mindset is the best one. Create for every project a new repo and connect them with build tools. But why not hire 100 SOA-Consultants, they have enough money now.
- deleted 15y ago[deleted]
- mikeocool 15y agoNo, literally the entire codebase for all of their products is in one Perforce repo. Ashish Kumar, manager of the Engineering Tools team, mentions it in this presentation: http://www.infoq.com/presentations/Development-at-Google http://www.infoq.com/presentations/Development-at-Google
- sek 15y agoVery interesting, thank you.
- rachelbythebay 15y agoThe kernel? Android? Some other spooky stuff involving the pest control guy who's holding a big rubber mallet when you fail a unit test? Are you sure about that?
- nostrademons 15y ago
- deleted 15y ago[deleted]
- deleted 15y ago[deleted]
- akg 15y agoI don't think Git was designed to perform well with such a large repo. In this case, the best-practice is probably compartmentalizing the code and using Git submodules. The Git submodule interface is a little un-friendly, but I think it does work well for such large repos. I've been using submodules successfully for our development that tracks source files as well as binary assets.
- losvedir 15y agoHuh, fascinating. git was initially created for the Linux kernel development, and I haven't heard of any issues there. Offhand I would have said, as a codebase, the Linux kernel would be larger and more complex than facebook, but I don't have a great sense of everything involved in both cases. So what's the story here: kernel developers put up with longer git times, the kernel is better organized, the scope of facebook is more massive even than the linux kernel, or there's some inherent design in git that works better for kernel work than web work?
- cbs 15y agoFrom the sounds of it facebook has a really, really big ball of highly coupled code.
- joelhaasnoot 15y agoSounds like PHP. Oh wait, it is.
- jonursenbach 15y agoNot really relevant.
- joelhaasnoot 15y agoWhile it may not be, I do find it an every day battle to keep my PHP wel 'styled'. Sure, PHP is the first language I learnt, and I do use a framework, but sometimes a long method is easier in the short run than writing good model functions. And I have models, but mostly for the ORM and they're all completely interconnected. PHP makes me lazy, fast
- mise 15y agoIt's for this type of thinking that I've been looking at Rails recently.
- dustingetz 15y agothe obvious answer, repeatedly mentioned in comments: > factor into modules, one project per repo where i work we have a project with clear module boundaries, but all in the same repo. we have an "app" and some dependencies including our platform/web framework. none of these are stable, they're all growing together. Commits on the app require changes in the platform, and in code review it is helpful to see things all together. Porting commits across different branches requires porting both the application change and the dependent platform changes. Often a client-specific branch will require severe one-off changes so the platform may diverge -- it is not practical for us (right now) to continually rebase client branches onto the latest platform. this is just our experience, not facebook's, but lets face it: real life software isn't black and white, and discussion that doesn't acknowledge this isn't particularly helpful.
- sek 15y agoThis is what git submodules are for, but when they can't use them they don't have clear module boundaries.
- snprbob86 15y agoWe've experienced this. We've got a superproject with our server configs, and sub projects for our background processing, API, and web-frontend respectively. Often, each project can evolve and be versioned 100% independently. However, often you need to modify multiple projects and (especially with server config changes) coordinate changes via the super project. It's a little hairy sometimes and often feels like unnecessary overhead, but the mental boundary is extremely valuable on it's own. Being able to add a field to the API and check that commit into the superproject for deployment before the front end features are done is nice. The social impact on implementation modularity is valuable. We write better factored code by letting Git establish a more concrete boundary for us.
- iamleppert 15y agoI can believe this working with a former facebook employee. They do not believe in separating or distilling anything into separate repos. Why the fuck would you want to have a 15GB repo? Ideally they should have many small, manageable repositories that are well tested and owned by a specific group/person/whatever. At least something small enough a single dev or team can get their head around. Sheesh.
- SoftwareMaven 15y agoAnd then each of those dev teams can spend 1/2 their time writing code other people in the company have already written or every team can spend 1/2 their time publishing and reading documentation about what has been written. There is no simple answer. There is only optimization for a particular problem-set you are trying to minimize.
- brown9-2 15y ago> And then each of those dev teams can spend 1/2 their time writing code other people in the company have already written or every team can spend 1/2 their time publishing and reading documentation about what has been written. I don't see what this has to do with a discussion of one repo vs multiple repos. You think that in a multi repo world, the engineers aren't as aware of what code exists and where as they are in a single repo world? You think that code duplication and needing to read docs magically doesn't exist in a single repo world? The number of repositories is just an organizational construct. Communication still must take place no matter what.
- earino 15y agoIf this was the crazy size of your git repo, why wouldn't you make a tool that took your git repo and versioned it? Keep it in repos that can all be performant, since most of the time you are working with "time local" information?
- pwpwp 15y agoGit was designed for the Linux kernel, and it's simply not big: a couple thousand files, broken up into directories of dozens or hundreds of files. http://www.schoenitzer.de/lks/lks_en.html#new_files http://www.schoenitzer.de/lks/lks_en.html#new_files
- ctz 15y agoI'm surprised Facebook and all its peripheral development has that much source. I would expect something like 5-10 million lines of code, not ~100 million lines implied by the example.
- nbm 15y agoThe example is synthetic, so don't worry too much about the implications. It is useful to keep in mind that Facebook isn't just the front-end (and isn't just code, also images, configuration, and so forth). Just talking about open source stuff, Facebook also generates code like Cassandra, Hive (data warehousing application), Phabricator (a code review and lifecycle tool), HipHop for PHP (the translator/compiler, the interpreter, and the virtual machine), FlashCache (a kernel driver), Thrift, Scribe, and so forth. We also have had to build applications to support our operations, so think about what sort of effort goes into building scalable monitoring, configuration management, automatic remediation, logging infrastructure, and so forth. I don't know the actual lines of code across it all, and wouldn't mention it if I did, but people often underestimate the scale here.
- moe 15y agoAnd all of that must live in a single repository... because?
- nbm 15y agoIt doesn't live in a single repository. The commenter I was replying to mentioned "Facebook and all its peripheral development" and a number of lines of code. I wanted to give him a little insight into what sort of things all the peripheral development might include, since it isn't obvious.
- lbrandy 15y agoWow. I was expecting an interesting discussion. I was disappointed. Apparently the consensus on hacker news is that there exists a repository size N above which the benefits of splitting the repo _always_ outweigh the negatives. And, if that wasn't absurd enough, we've decided that git can already handle N and the repository in question is clearly above N. And I guess all along we'll ignore the many massive organizations who cannot and will not use git for precisely the same issue. So instead of (potentially very enlightening conversation) identifying and talking about limitations and possible solutions in git, we've decided that anyone who can't use git because of its perf issues is "doing it wrong".
- jamespo 15y agoI would have liked this comment better if it came up with some solutions itself, although maybe it's not easy to solve?
- djtriptych 15y agoThat's the problem. It's NOT an easy problem to solve. A lot of posts on hn describing some problem elicit "Why, that's no problem at all!" responses or "That's the wrong problem to think about" responses. Honestly that mindset is often really useful in programming, but when we get a problem that doesn't have a shortcut and is relevent, conversation goes to shit. Because I guess that's when programmers normally go into a hole and brute-force brain it out. How to use mass comms to talk about a difficult open problem is, I suppose, itself an open problem.
- sek 15y agoGit was not designed for this. It comes out of the Linux kernel, where you need a secure hash of a segment to prevent compromise. For big projects you have submodules, you can only get a level higher later. In a company, you trust the sources. With Perforce you check out files and work with the part you want. It is a design decision and they could have known before.
- lincolnq 15y ago
- loeg 15y agoI'm curious what their performance numbers look like if they host the .git repo on tmpfs -- 15GB isn't unreasonable on a beefy (24-32GB of ram) machine.
- wmf 15y agoProbably the same as the warm cache results, since that's basically what tmpfs is. I wonder if git does all that stat()ing serially or in parallel, though...
- durin42 15y agoI don't have the link handy, but IIRC we did some experiments with that for Mercurial and found that stat() in parallel didn't really help much until you were using NFS or similarly slow network-type latency filesystems.
- patangay 15y agoFor what it's worth, Josh has been doing all his tests on machine with SSDs and 72GB of RAM.
- jrockway 15y agoYes, it's well known that big companies with big continuously integrated codebases don't manage the entire codebase with Git. It's slow, and splitting repositories means you can't have company-wide atomic commits. It's convenient to have a bunch of separate projects that share no state or code, but also wasteful. So often, the tool used to manage the central repository, which needs to cleanly handle a large codebase, is different from the tool developers use for day-to-day work, which only needs to handle a small subset. At Google, everything is in Perforce, but since I personally need only four or five projects from Perforce for my work, I mirror that to git and interact with git on a day-to-day basis. This model seems to scale fairly well; Google has a big codebase with a lot of reuse, but all my git operations execute instantaneously. Many projects can "shard" their code across repositories, but this is usually an unhappy compromise. People always use the Linux kernel as an example of a big project, but even as open source projects go, it's pretty tiny. Compare the entire CPAN to Linux, for example. It's nice that I can update CPAN modules one at a time, but it would be nicer if I could fix a bug in my module and all modules that depend on it in one commit. But I can't, because CPAN is sharded across many different developers and repositories. This makes working on one module fast but working on a subset of modules impossible. So really, Facebook is not being ridiculous here. Many companies have the same problem and decide not to handle it at all. Facebook realizes they want great developer tools and continuous integration across all their projects. And Git just doesn't work for that.
- Splines 15y agoAt Google, everything is in Perforce, but since I personally need only four or five projects from Perforce for my work, I mirror that to git and interact with git on a day-to-day basis. At MS we also use Perforce (aka Source Depot), and I've toyed with the idea of doing something similar. Have you found any guides for "gotchas" or care to share what you've learned going this route?
- jrockway 15y agoI used git-p4 at my last job, and the only thing that ever got weird was p4 branches. At Google we have an internal tool that's similar to git-p4, and it always works perfectly for me. Enough developers are using it such that most of the internal tools understand that a working copy could be a git repository instead of a p4 client. So if you're planning on doing this at your own company, my advice is to write your own scripts that make whatever conventions you have automatic, and to move everyone over at the same time. That way, you won't be the weird one whose stuff is always broken. I think most people got burned by cvs2svn and git-svn and think that using two version control systems at once is intrinsically broken. It's not. svn was just too weird to translate to or from. (People that skipped svn and went right from cvs to git had almost no problems, I'm told.)
- jpdoctor 15y ago$100B company, maybe they can afford to put some people onto solving this for the open software community (and put the solution into the open), especially since nobody else in the community seems to have this problem.
- rogerbinns 15y agoIf you proposed a good solution I'm sure they'd be happy to provide time and money and open source the result. But most of the responses aren't even that there is a solution - they say to split the repository into smaller pieces and spend time and money internally having their internal developers deal with that. A good solution will benefit everyone who uses git. Codebases get larger over time. There is more forking and experimentation. More spoken languages can be supported. More computer languages can be interfaced. The O(n) operations becoming less than that will benefit you in the future as your code grows.
- jpdoctor 15y ago> If you proposed a good solution I'm sure they'd be happy to provide time and money and open source the result. If they provided money, I'd provide the time in order to produce a good result. See the problem? More to the point: FB is all take and no give, as near as I can tell.
- karlshea 15y agohttps://developers.facebook.com/opensource/ https://developers.facebook.com/opensource/ Looks like a fairly long list to me.
- groby_b 15y agoThey are putting people on this. Thankfully, those people are smart enough to ask for help before blindly going off and doing their own thing. Now, if it's going to end up OSS, that's a different question. (I'm not implying it's not - I'm saying that's a decision that could go either way)
- gokhan 15y agoLarge repos bring their own problems, and results in some design decisions accordingly. For example, Visual Studio itself is 5M+ files and this affected some of the the initial design decisions (Server side workspaces, for this example) when developing TFS 2005 (the first version) [1]. That decision suits MS but not the small to medium clients well. So they're now alternating that design with client side workspaces. It's not wise to offer Facebook to split the repository. Looks like it's time to improve the tool. [1] http://blogs.msdn.com/b/bharry/archive/2011/08/02/version-control-model-enhancements-in-tfs-11.aspx http://blogs.msdn.com/b/bharry/archive/2011/08/02/version-co...
- julian37 15y agoSomewhat off-topic, could somebody explain why echo 3 | tee /proc/sys/vm/drop_caches rather than just echo 3 > /proc/sys/vm/drop_caches Is it because the output to stdout lets you be extra sure that the right data was sent to the kernel? I'm just wondering if this is an idiom with a deeper meaning that I'm not aware of. EDIT: I'm guessing that when you run it in a script (without set -x), rather than on the command line, you can see in the log what it is you sent?
- pdw 15y agoBecause you can echo 3 | sudo tee /proc/sys/vm/drop_caches but sudo echo 3 > /proc/sys/vm/drop_caches won't work.
- jochu 15y agoAside from reasons you mentioned, I can imagine it being because it easily allows one to add a sudo or being habit because of it. For example: echo 3 | sudo tee /proc/sys/vm/drop_caches Will allow you to write as root and sudo echo 3 > /proc/sys/vm/drop_caches Will be a permission error. It executes the echo as root and the write as the user
- slashclee 15y agoThese times are for spinning-platter hard drives. I wonder what the numbers look like on a modern SSD?
- xxiao 15y agogit is not memory efficient by design, i used to push about 1G commit to the server and it hangs forever, i had to abort it and push in as small chunks instead
- lnguyen 15y agoThere's two issues: the width of the repository (number of files) and the depth (the number of commits). Since "status" and "commit" perform fairly well after the OS file cache has been warmed up, that probably can be resolved by having background processes that keep it warm. (Also, how long would it take to just simply stat that number of files? ) The issue of "blame" still taking over 10 minutes: We need to know far back in the repository they're searching. What happen if there's one line that hasn't been changed since the initial commit? Are you being forced to go back to through the whole commit history? How old is the repository? Years? Months? I'm probably guessing in the at least years range based on the number of commits (unless the developers are extremely commit-happy). At a certain point, you're going to be better off taking the tip revisions off a branch and starting a fresh repository. It doesn't matter what SCM/VCS tool you're using (I've been the architect and admin on the implementation of a number of commercial tools). Keep the old repository live for a while and then archive it. You'll find that while everyone wants to say that they absolutely need the full revision history of every project, you rarely go back very far (aka the last major release or two). And if you do need that history, you can pull it from the archives.
- bos 15y agoFacebook engineer here, working on this problem with Joshua. What this comes down to is that git uses a lot of essentially O(n) data structures, and when n gets big, that can be painful. A few examples: * There's no secondary index from file or path name to commit hash. This is what slows down operations like "git blame": they have to search every commit to see if it touched a file. * Since git uses lstat to see if files have been changed, the sheer number of system calls on a large filesystem becomes an issue. If the dentry and inode caches aren't warm, you spend a ton of time waiting on disk I/O. An inotify daemon could help, but it's not perfect: it needs a long time to warm up in the case of a reboot or crash. Also, inotify is an incredibly tricky interface to use efficiently and reliably. (I wrote the inotify support in Mercurial, FWIW.) * The index is also a performance problem. On a big repo, it's 100MB+ in size (hence expensive to read), and the whole thing is rewritten from scratch any time it needs to be touched (e.g. a single file's stat entry goes stale). None of these problems is insurmountable, but neither is any of them amenable to an easy solution. (And no, "split up the tree" is not an easy solution.)
- tonfa 15y agoNice to see you're still in the DVCS business ;)
- deleted 15y ago[deleted]
- Terretta 15y ago> waiting on disk I/O Out of curiosity, why are these benchmarks using regular disk and flash disk? At only 15 GB, what happens using ram disk? Sure SSD is fast, but for these things it's still really slow.
- groby_b 15y agoAn inotify daemon could help, but it's not perfect: it needs a long time to warm up in the case of a reboot or crash So does, presumably, the cache when you use lstat. (Let's scratch presumably. It does. Bonus points if you can't use Linux and use an OS that seems to chill its caches down as soon as possible. ) I hope I'm wrong, but the proper solution to this seems to be a custom file system - not only will it allow you to more easily obtain a "modified since" list of files, it also allows you to only get local files "on demand". (E.g. http://google-engtools.blogspot.com/2011/06/build-in-cloud-accessing-source-code.html http://google-engtools.blogspot.com/2011/06/build-in-cloud-a...) That still doesn't solve the data structure issues in git, but at least it takes some of the insane amount of I/O off the table. I'm looking forward to see what you guys cook up :)
- ramanujan 15y agoThis looks like it could be of assistance: http://source.android.com/source/version-control.html http://source.android.com/source/version-control.html Repo is a repository management tool that we built on top of Git. Repo unifies the many Git repositories when necessary, does the uploads to our revision control system, and automates parts of the Android development workflow. Repo is not meant to replace Git, only to make it easier to work with Git in the context of Android. The repo command is an executable Python script that you can put anywhere in your path. In working with the Android source files, you will use Repo for across-network operations. For example, with a single Repo command you can download files from multiple repositories into your local working directory. http://google-opensource.blogspot.com/2008/11/gerrit-and-repo-android-source.html http://google-opensource.blogspot.com/2008/11/gerrit-and-rep... With approximately 8.5 million lines of code (not including things like the Linux Kernel!), keeping this all in one git tree would've been problematic for a few reasons: * We want to delineate access control based on location in the tree. * We want to be able to make some components replaceable at a later date. * We needed trivial overlays for OEMs and other projects who either aren't ready or aren't able to embrace open source. * We don't want our most technical people to spend their time as patch monkeys. The repo tool uses an XML-based manifest file describing where the upstream repositories are, and how to merge them into a single working checkout. repo will recurse across all the git subtrees and handle uploads, pulls, and other needed items. repo has built-in knowledge of topic branches and makes working with them an essential part of the workflow. Looks like it's worth taking a serious look at this repo script, as it's been used in production for Android. Might allow splitting into multiple git repositories for performance while still retaining some of the benefits of a single repository.
- jmccaffrey 15y agoHaving worked with repo professionally, I'm not a fan. You lose simple ability to track dependencies across repositories or even revert to a previous consistent point in time without diligent tagging. Even with good tags, restructuring your project setup and changing your repo manifest can still break your ability to go back in time.
- cdibona 15y ago
- DannoHung 15y agoMultiple people in this conversation section have asserted that code sharing is way easier when all the code is in a single repo, but from my understanding of sub-modules, it would be a fairly simple matter of setting up your pre/post-commit hooks to update submodules to a branch automatically and get useful company wide change atomicity (after all, changes should only propagate between teams/projects once they have some stability). Putting aside the question of whether or not an enormous singular repo can be broken up intelligently into modular projects, is there something about the submodule approach that makes it a uniquely unsuitable way for sharing changes amongst projects?
- groby_b 15y agoIf you limit change propagation, your changes won't propagate as fast. That goes for bugs and bug fixes. I can certainly see why you would have the latter propagated instantaneously, or close to it. There's also the point that if you don't propagate change to everybody at the same time, you'll have dozens of slightly different versions of those projects across your company. The question of submodules vs. large repo is not as easily decided as you think - there are large upsides (and downsides) to both approaches.
- DannoHung 15y agoNo, you're misunderstanding what I'm saying with regards to publishing stable changes. Say you have two branches: master and next. Stable work goes in master, unstable work goes in next. When the code is ready for consumption, you merge it into master. Anyone who is using a project has it setup as a submodule. They add post/pre-commit hooks to update all project submodules. These submodules pull from master. This way, everyone will get all stable changes on all submodule projects at the time of the next change to their own project.
- groby_b 15y agoI do understand what you mean just fine. Except it doesn't work that way. If I have a critical bug fix in libA, I need it rolled out now. What's more, I need it rolled out across all other projects that use libA, immediately. And no, I don't want to work until all projects committed a change of their own. Even more, I (or my team) are not the only ones working on libA. Others are too. So keeping changes in 'next' and pushing to master only on occasion doesn't help much. Yes, it keeps non-working patches out - but that's what local branches on your machine are for. (I'm not even going to mention the issue of merge conflicts. If you work on a massive scale, the longer you stay in a branch, the more likely you are to get a merge conflict. There's easily the chance to go into a several weeks long merge hell. Pull from master, resolve conflicts, run local tests - oh wait, master is already updated by somebody else)
- alok-g 15y agoIs anything known for scalability to such sizes for Subversion, Mercurial, Bazaar, and others?
- charlieok 15y agoI think it's a bad practice to keep a giant code base in one repo. Split the code base into purpose-specific modules, just as you would split any project into purpose-specific modules. In fact, those two things might well line up 1:1. If a project depends on other projects, have it reference the other projects. Where appropriate, include exact version numbers and/or commit hashes. Gemfiles are good examples of this good practice at work. Yes, git has submodules for this sort of thing, but after investigating that route, I decided against using git submodules. Use something independent of the VCS instead. Then git won't do weird or unexpected things when you switch branches. Also, you might want to mix in projects that use other version control systems. And really, why unnecessarily couple a project to its version control system? If (when?), even after splitting a megaproject into manageable subprojects, these performance issues creep in, I'd certainly be interested in whatever improvements people are coming up with...
- teyc 15y agoThis is an interesting social AND a technical problem. The problem for FB is that it is all too easy for them to just fork git, create the necessary interfaces and then hope the git maintainers would accept it (they mightn't) or release it into the wild (and incur bad karma and wrath of OS developers who'd see this has schism or even heresy). They've reached out to the developers on git, and I guess that's a first step.
- r15habh 15y agoDoes Facebook really believe that because they have the most users, they should also have the biggest git repo? Amazon and Google have already solved this problem, and the solution is to reorganize things into smaller manageable packages.
- r15habh 15y agoWhat I mean is, even if they manage to solve this problem with some tweaks, they will again hit the bottleneck in an year or two, so they should rethink their source code management
- redstone 15y agoThis is Joshua (who posted the original email). I'm glad to see so much interest in source control scalability. If there are others who have ever contemplated investing a bit of time to improving git, it'd be great to coordinate and see what makes sense to do - even if it turns out that the right answer is just to make the tools that manage multiple repos so good that it feels as easy as a single repo.
- railsmax 15y agoHey do you know a lot of sites with such needs? Facebook is the first, and probably all sites with such needs I can count with fingers on my one hand. I don't think it's git issue - everyone use this system and all are happy using it. This is like a new feature, but not issue.
- MikeOnFire 15y agoMy first thought, as suggested by some on the list, was modularization. Redstone's response (that the 1.3 million files are essentially all interdependent) terrifies me.