13 ms·
Facebook's git repo is 54GB
- ianphughes 12y agoI wonder what their branching strategy is like and how merges are gated with a single codebase of that size?
- indygreg2 12y agoThey aim for a completely linear history. They may even have a policy of not allowing merge commits. It is described in various places on the internet. I like https://secure.phabricator.com/book/phabflavor/article/recommendations_on_branching/ https://secure.phabricator.com/book/phabflavor/article/recom... because it and its sister articles on code review and revision control are terrific reads.
- natrius 12y agoDear everyone: you should be using Phabricator. It is Facebook's collected wisdom about software development expressed in software. It has improved my life substantially. The code review is better than Github's, and their linear, squashed commit philosophy has worked out much better than the way I used to do things.
- ianphughes 12y agoIt looks pretty great. How does it compare to Atlassian products (if you have used any)?
- natrius 12y agoI've used JIRA, and I slightly prefer Phabricator's tasks. They're mostly the same.
- andrewjshults 12y agoMixed bag. The code review part is much better than stash and significantly better than crucible. Namely, diff of diffs makes reviewing changes based on comments infinitely easier (especially on large reviews). We installed phabricator just for the code review piece initially. Repo browsing is about on par with stash, but it doesn't seem to experience the horrific slow downs that our stash server does. We don't use the tasks because a number of non engineering roles also use JIRA and the tasks functions in phabricator don't have nearly the depth of security and workflow options we need.
- andrewjshults 12y agoMixed bag. The code review part is much better than stash and significantly better than crucible. Namely, diff of diffs makes reviewing changes based on comments infinitely easier (especially on large reviews). We installed phabricator just for the code review piece initially. Repo browsing is about on par with stash, but it doesn't seem to experience the horrific slow downs that our stash server does. We don't use the tasks because a number of non engineering roles also use JIRA and the tasks functions in phabricator don't have nearly the depth of security and workflow options we need.
- Timothee 12y agoThat might be a silly question but it's not useful for just an individual in a team, correct? We use GitHub at work, but I wouldn't be able to try Phabricator on my own, right?
- natrius 12y agoMostly correct. You can use the code review tools without needing your repository set up in Phabricator or anyone else with Phabricator accounts. You could comment on a diff and point your coworker to it, but it's unlikely to feel better than Github pull requests in that scenario.
- nvarsj 12y agoOh yeah, it's great when one person can break the build and stop all active development. It scales so well. Oh I know, to prevent any issues, let's protect ourselves with feature toggles. Oh and let's build a set of database driven rules to manage those toggles. Oh what about dependencies? Let's build a manager to manage the DAG of feature toggle dependencies. Need I go on? :-) You've replaced a relatively simple system of merge requests with some pseudo in-code versioning system controlled through boolean global variables. I'll take feature branches any day of the week over that mess. The github model is far superior IMO.
- natrius 12y agoMy team uses Phabricator without any feature toggles in our code. You land code onto master when it's ready to ship. Until then, you have features developed and reviewed on separate branches. I don't get how that's more or less fragile than merges.
- nvarsj 12y agoI was addressing the idea of committing directly to master, protecting your code with feature toggles so it doesn't break things. Maybe I misunderstood the OP. I think feature toggles can be extremely useful, but still develop in a branch and merge after review/qa.
- danudey 12y agoWe're using Phabricator at my company, and I'm getting to the point of encouraging people to use it and start seeing the benefits. I work on infrastructure, so when people come to me with issues, I find out what problem they're having and then get them to submit a ticket to me. I've also started creating tickets for issues and assigning them to other people to get them to take a look at it, and people seem pretty receptive. It hasn't become part of everyone's workflow yet, but it's pretty useful.
- wtetzner 12y agoSeems like you lose most of the value of git if you can't do merges.
- natrius 12y agoYou can still merge from master (or whatever your main development branch is) into your feature branches. You just never merge into your main branch. You squash the commits and rebase. The history ends up looking much better, and `git blame` becomes much more useful.
- SnakeDoc 12y agoLinear codebase history? Why even use Git then... that's SVN stuff... we use Git now-a-days for a reason...
- donaldguy 12y agoLast I knew, the FB mainline codebase was in fact still in SVN with git-svn dcommit (possibly hidden under arc) being how things land in "master" (and the revisions being merged quasi-manually to a release branch immediately before HH compile) FB doesn't need to branch ... Gatekeeper (their A/B, feature flag system) really takes care of that concern logically
- djur 12y agoThere's a big difference between a linear history produced by actual linear development and one produced by `git rebase -i`. They both have the advantage of being easier to understand later, however.
- ulisesrmzroche 12y agoThat's not true dude. If you're following actual linear development you're likely to see a lot of 'poke build' and 'change css' and stuff like that. Git rebase -i gives you a change to rename your commit and organize it in a readable way. So git rebase -i will be more readable, while actual linear history is always gibberish.
- VikingCoder 12y agoPay attention to the footnote: *8 GB plus 46 GB .git directory
- danbruc 12y ago8 GB is still a lot. Would be interesting to know how much of it is actual code and how much is just images and so on.
- cliveowen 12y agoExactly, there's no way they wrote 8GB of code.
- justincormack 12y agoI think they did, have heard similar figures from them before, all deployed as one file. Macroservices...
- SnakeDoc 12y agounlikely. they probably have a very dirty repo with tons of binaries, images, blah blah blah. It's highly unlikely they actually wrote 8GB of code, and the 46GB .git directory will be littered with binary blob changes, etc. This is really just to "impress" two people: 1) People who love Facebook 2) People who don't know anything about version control and/or how to do proper version control (no binaries in the scm).
- justincormack 12y agoIn 2012 the Facebook binary was 1.5GB http://arstechnica.com/business/2012/04/exclusive-a-behind-the-scenes-look-at-facebook-release-engineering/ http://arstechnica.com/business/2012/04/exclusive-a-behind-t...
- vmarsy 12y agoI cited this interesting paper from Facebook engineers(2013) in another comment : https://news.ycombinator.com/item?id=7648802 https://news.ycombinator.com/item?id=7648802 In 2011 they had 10 millions LoC, up to 500 commits a day, but if we asssume the plots keeps going up like this, now in 2014 it can be pretty big. Their binary was 1.5GB when the paper was written.
- deleted 12y ago[deleted]
- _kushagra 12y agoWhy's it bad to store source code on Git?
- samcasas 12y agoIts not bad, is really nice, but Git has one problem, when you codebase is big, the process takes a long time, imagine git scanning those 8GB every time you do a commit, that is why Facebook was looking to port all their code to another VCS
- aeturnum 12y agoFunny you should mention Facebook maybe having performance problems with Git: http://git.661346.n2.nabble.com/Git-performance-results-on-a-large-repository-td7250867.html http://git.661346.n2.nabble.com/Git-performance-results-on-a... (2012)
- pionar 12y agoThis is when a centralized VCS with checkin-checkout concepts really shines. The client only has to check the checked-out files on commit.
- chrismonsanto 12y agoI think it's worth making a distinction between the Git plumbing and the Git porcelain when talking about performance. The core functionality (the plumbing) is very fast regardless of repository size. The slowdowns people describe are almost always related to the porcelain commands, which are poorly optimized. Almost every porcelain-level command will cause Git to lstat() every file in your tree, as well as check for the presence of .gitignore files in all of the directories. It's very wasteful. The fix for this is pretty simple: use filesystem watch hooks like inotify to update an lstat cache. I wrote something like this for an internal project and the speed difference was night and day. I remember reading that there had been progress on the inotify front on the git dev mailing list a few years ago, don't know what the current status is.
- jasonnutter 12y agoRelevant: https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc...
- antimatter 12y agoDidn't they switch to Mercurial? https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc...
- joshstrange 12y agoSomebody asked that on Twitter and the OP responded with: >> At least according to the presentation by a Facebook engineer that I just watched, they're still on git. [0] [0] https://twitter.com/feross/status/459335105853804544 https://twitter.com/feross/status/459335105853804544
- ableal 12y agoYou can check out, but you can never leave?
- rickr 12y agoI thought I had read an article about facebook switching to perforce due to their really large git repo. Were they at least thinking about it? A quick google comes up with nothing but I could have SWORN I read that.
- ianphughes 12y agoNot like they dont have the money for it, but that would be a very expensive for them and sort of anti-Open Source, no?
- delroth 12y agoYou are thinking about Mercurial, not Perforce: https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc...
- TheCoreh 12y agoI bet most of that size is made up from the various dependencies Facebook probably has, though I'm still surprised it's that large. I expected the background worker things, like the facial recognition system for tagging people, and the video re-encoding libs, to be housed on separate repositories. I also wonder if that size includes a snapshot of a subset of Facebook's Graph, so that each developer has a "mini-facebook" to work on that's large enough to be representative of the actual site (so that feed generation and other functionalities take somewhat the same time to execute.)
- fletchowns 12y agoThere's gotta be a ton of binaries in there
- pavel_lishin 12y agoSurely a mini-graph would be something that would be loaded by a script, instead of being stored in the repo itself.
- weavie 12y agoReminds me of a recent project we had. The code we wrote came in at about 150k after an npm and bower install it was at 192mb.
- uaygsfdbzf 12y agoout of curiosity, why were (and if so, why?) you committing node_modules and the bower destination dir.?
- indygreg2 12y agoHaving all code in a single repository increases developer productivity by lowering the barrier to change. You can make a single atomic commit in one repository as opposed to N commits in M repositories. This is much, much easier than dealing with subrepos, repo sync, etc. Unified repos scales well up to a certain point before troubles arise. e.g. fully distributed VCS starts to break down when you have hundreds of MB and people with slow internet connections. Large projects like the Linux kernel and Firefox are beyond this point. You also have implementation details such as Git's repacks and garbage collection that introduce performance issues. Facebook is a magnitude past where troubles begin. The fact they control the workstations and can throw fast disks, CPU, memory, and 1 gbps+ links at the problem has bought them time. Facebook made the determination that preserving a unified repository (and thus preserving developer productivity) was more important than dealing with the limitation of existing tools. So, they set out to improve one VCS system: Mercurial (https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc...). They are effectively leveraging the extensibility of Mercurial to turn it from a fully distributed VCS to one that supports shallow clones (remotefilelog extension) and can leverage filesystem watching primitives to make I/O operations fast (hgwatchman) and more. Unlike compiled tools (like Git), Facebook doesn't have to wait for upstream to accept possibly-controversial and difficult-to-land enhancements or maintain a forked Git distribution. They can write Mercurial extensions and monkeypatch the core of Mercurial (written in Python) to prove out ideas and they can upstream patches and extensions to benefit everybody. Mercurial is happily accepting their patches and every Mercurial user is better off because of Facebook. Furthermore, Mercurial's extensibility makes it a perfect complement to a tailored and well-oiled development workflow. You can write Mercurial extensions that provide deep integration with existing tools and systems. See http://gregoryszorc.com/blog/2013/11/08/using-mercurial-to-query-mozilla-metadata/ http://gregoryszorc.com/blog/2013/11/08/using-mercurial-to-q.... There are many compelling reasons why you would want to choose Mercurial over other solutions. Those reasons are even more compelling in corporate environments (such as Facebook) where the network effect of Git + GitHub (IMO the foremost reason to use Git) doesn't significantly factor into your decision.
- gavinpc 12y agoIn terms of engineering tradeoffs, this reminds me of a recent talk by Alan Kay where he says that to build the software of the future, you have to pay extra to get the hardware of the future today. [1] Joel Spolsky called it "throwing money at the problem" when, five years ago he got SSD's for everybody at Fog Creek just to deal with a slow build. [2] I don't use Facebook, and I'm not suggesting that they're building the software of the future. But surely someone there is smart enough to know that, for this decision, time is on their side. [1] https://news.ycombinator.com/item?id=7538063 https://news.ycombinator.com/item?id=7538063 [2] http://www.joelonsoftware.com/items/2009/03/27.html http://www.joelonsoftware.com/items/2009/03/27.html
- corysama 12y agoYou should also watch his talk where he questions if it is really necessary to build large software products from such large codebases that, if printed out, would stack as high as a skyscraper. And, then talks about what he is doing to demonstrate that it is not :) https://news.ycombinator.com/item?id=7538073 https://news.ycombinator.com/item?id=7538073
- lnanek2 12y agoFacebook tends to throw engineer time at the problem, though. I know one Facebook DevCon I went to they presented how they completely wrote their own build system because Ant was too slow for them.
- georgemcbay 12y agoHaving used ant as a build system for Android projects, I don't blame them. In my admittedly limited experience (Windows 7 x64, ant, Android SDK) ant is terribly slow to build projects with multiple source library dependencies and throwing hardware at the problem doesn't speed it up that much.
- diek 12y agoI don't see how this is an Ant-specific issue. Ant is just calling into javac with a classpath parameter. The actual execution time spent in Ant should be minimal.
- hk__2 12y agoIs there a reason why they keep everything in the same repo? Can’t you just split the code across multiple smaller repos?
- rcxdude 12y agoIt becoms a lot harder to keep everything in sync, especially if internal interfaces change frequently. At facebook scale though it's probably a good idea to defined boundaries between areas in the application better.
- ulisesrmzroche 12y agoNo, it doesn't. It's actually the complete opposite because you know, 'though shall separate those things that change frequently from those that don't'
- Oompa 12y agoYou end up with less developers having to pull & merge/rebase if you have things in separate repos. Individual libraries/dependencies get worked on by themselves, with an API that other applications use. Then the other apps just bump a version number and get newer code.
- taeric 12y agoThe problem with this, is that you are assuming the APIs change in some sort of odd isolation to the parts that use them. That is, the reason an API changes is because a use site has need of a change. So, at a minimum, you need to make that change and test it against that site in a somewhat atomic commit. Then, if the change has any affect on other uses, you need a good way to test that change on them at the same time. Otherwise, they will resist pulling this change until it is fixed. Add in more than a handful of such use sites, and suddenly things are just unmanageable in this "manageable" situation. Not that this is "easy" in a central repo. But at least with the source dependency, you can get a compiler flag at every place an API change breaks something. And, true, you can do this with multiple repos, too. But every attempt I have seen to do that just uses a frighteningly complicated tool to "recreate" what looks like a single source tree out of many separate ones. (jhbuild, and friends) So, if there is a good tool for doing that, I'd certainly love to hear about it.
- dnlserrano 12y agoSomeone recently told me that Facebook had a torrent file that went around the company that people could use to download the entire codebase using a BitTorrent client. Is there any truth in this? I mean, the same guy that told me this, also said that the codebase size was about 50 times less than the one reported in this slide, so it may all be pure speculation.
- ionforce 12y agoFacebook uses BitTorrent to deploy their binaries to their many servers. So using torrents isn't foreign to them.
- vmarsy 12y agoIf you're interested in the deployment process at Facebook, look at the link of a Facebook engineers paper I submitted in my other comment in this thread : https://news.ycombinator.com/item?id=7648802 https://news.ycombinator.com/item?id=7648802 "The deployed executable size is around 1.5 Gbytes, including the Web server and compiled Facebook application. The code and data propagate to all servers via BitTorrent, which is configured to minimize global traffic by exploiting cluster and rack affinity. The time needed to propagate to all the servers is roughly 20 minutes."
- threedaymonk 12y agoThat sounds entirely believable. At $PREVIOUS_JOB, the Puppet git repository was large enough (several GB) that cloning it was painful, and new starters were handed a pruned repository via the local network so that they could get something done today.
- deleted 12y ago[deleted]
- general_failure 12y agoThe worrying point here is the checkout of 8GB as opposed to the history size itself (46GB). If git is fast enough with SSD, this is hardly anything to worry about. I actually prefer monolithic repos (I realize that the slide posted might be in jest). I have seen projects struggle with submodules and splitting up modules into separate repos. People change something in their module. They don't test any upstream modules because it's not their problem anymore. Software in fast moving companies doesn't work like that. There are always subtle behavior dependancies (re: one module depends on a bug in another module either by mistake or intentionally). I just prefer having all code and tests of all modules in one place.
- Touche 12y agoHow does monolithic repos solve that. Surely people who fix bugs in a library aren't testing the entirety of Facebook every time (how long would that even take? Assuming they've even set such a thing up.)
- taeric 12y agoIt is at least easier to correlate the changes. When you have X+ modules, you have potentially X+ histories you have to look at to know when a change was seen with another change.
- GregorStocks 12y agoI used to work at Facebook. They have servers that automatically run a lot of their test cases on every commit.
- deleted 12y ago[deleted]
- thathonkey 12y agoWe use separate repos and it works out well. It's nice having separate Git histories that pertain to different areas of the codebase. Our workflow covers all the potential problems you named (eg. scripts to keep everything up to date, tests that get run at build or push time after everything is already checked out from the individual repos, etc.). We've been running this way for over a year with literally zero issues.
- slig 12y agoAm I missing something or this means a new intern working on a small feature, for instance, would have access to entire codebase?
- dman 12y agoThat is a feature not a bug! Discoverability of code helps improve code quality and makes things less fragile.
- ulisesrmzroche 12y agoDiscoverability decreases as LOC increase, so that's not true.
- TD-Linux 12y agoYes, and an intern would be subject to a code review before pushing to master - no way would they have write access to the "master" repo.
- slig 12y agoYes, but the new intern would be able to read all the source and "secret sauces". I doubt that an intern on Google would've access to the search codebase. I'd wager that only a handful of trusted employees have access to that codebase.
- kevinsf90 12y agoI thought they used Mercurial
- SnakeDoc 12y agoI hope everyone realizes this is not 54GB of code, but in fact, is more likely a very public showing of very poor SCM management. They likely have tons of binaries in there, many many full codebase changes (whitespace, tabs, line endings, etc). Also not to mention how much dead code lives in there?
- lnanek2 12y agoHonestly, I prefer check-in everything shops. It's way too often otherwise some different Java version, or IDE version, or Maven central being down screws something up, or you have to wait a long time for a Chef recipe or disk image to give you the reference version. Half my day today was dealing with someone updating Java on half our continuous build system slave computers and breaking everything because it didn't have JAVA_HOME and unlimited strength encryption all setup properly.
- SnakeDoc 12y agoThat sounds like either poor Sys Admin'ing and/or poor documentation... SCM should not have "clutter" in it, otherwise you wind up with an all-day download of 54GB of dead or useless garbage. The Kernel's repo is only a few GB's and it has MANY more changes and much more history than FB does...
- kyberias 12y agoYou should realize that the kernel doesn't have image and other assets that FB might have. And the repo obviously is the right place for them.
- pekk 12y ago"The" repo. As if they had no choice but to put everything for every aspect of the business into ONE giant repo. Hopefully they actually have some separate sites, separate tools and separate libraries. Or could understand how to use submodules or something rather than literally putting everything in one huge repository. Whether to put images and other assets into git repos is a separate decision.
- korzun 12y agoThat's actually not that bad for a engineering shop of their size. I would start archiving metadata at some point.
- lightblade 12y ago> @readyState would massively enjoy that first clone @feross The first clone does not have to go over the wire. Part of git's distributed nature is that you can copy the .git to any hard drive and pass it on to someone else. Then... > git checkout .
- pearjuice 12y agoSomeone must have forgotten a .gitignore or two.
- rl3 12y agoAlthough this is large for a company that deals mostly in web-based projects, it's nothing compared to repository sizes in game development. Usually game assets are in one repository (including compiled binaries) and code in another. The repository containing the game itself can grow to hundreds of gigabytes in size due to tracking revision history on art assets (models, movies, textures, animation data, etc). I wouldn't doubt there's some larger commercial game projects that have repository sizes exceeding 1TB.
- deleted 12y ago[deleted]
- kahoon 12y agoBut they surely don't use git for that, right? In scenarios like this a versioning system that does not track all history locally would be a better fit.
- stonemetal 12y agoPerforce tends to be big in game dev because it does better with repos full of giant blobs, has locking for working with them etc.
- eropple 12y agoPerforce is king in game development. It's also the only place I personally still use Subversion.
- efraim 12y agoI think most use Perforce.
- rl3 12y agoAs the other replies say, Perforce is dominant in commercial game development. However, Perforce does have Git integration now, allowing for either a centralized or distributed version control model. Considering the popularity of Git, I wouldn't doubt smaller Perforce-based game projects are going the DVCS route. Also, hypothetically speaking, consider if you had a game project that would eventually grow to 1-2TB in repository size. If you spent $100 per developer to augment each of their workstations with a dedicated 3TB hard drive, you would have an awesome level of redundancy using DVCS (plus all the other advantages). I know it's no replacement for cold, off-site backups, but it would still be nice.
- Dorian-Marie 12y agoThey must storing a lot of images and binary files I guess.
- alayne 12y agoFB had previous scaling problems with git which they discussed in 2012 http://comments.gmane.org/gmane.comp.version-control.git/189776 http://comments.gmane.org/gmane.comp.version-control.git/189... It appears they are now using Mercurial and working on scaling that (also noted by several others in this discussion): https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc...
- com2kid 12y agoMeh. I'm working on a comparably small project (~40 developers), and we're over 16GB. Mostly because we want a 100% reproducible build environment, so a complete build environment (compilers + IDE + build system) is all checked into the repro.
- math0ne 12y agoIDE checked into the repo eh? For some reason I kinda like that idea. So portable... if it works.
- bananas 12y agoCompany I did a contract for last year has 8MB of (Java) source code and a 52MB SVN repo and make £40 million a year out of it... We're doing something wrong.
- coherentpony 12y agoSo what? This probably means they're versioning data files they shouldn't be. I feel like this just exists here as a pissing contest.
- pekk 12y agoIt is supposed to justify the engineering effort they put into switching to Mercurial, then trying to make it "scale." (Rather than just using separate repositories to begin with, according to the design of the tool and best practices)
- SnakeDoc 12y agoaka. a pissing match to show "they are too big for any standard industry tools". Really speaks to the level of (non)expertise employed at FB.
- ausjke 12y agogosh, last time I had trouble with 8GB data checking in, it's very memory hungry when the data set is big and then you need check them in all at once, how much memory on the server side you need when you want to 'git add .' all the repo of 54GB? what about a re-index or something, will that take forever? I worry at such size the speed will suffer, I feel git is comfortable with probably a few GBs only? anyway it's good to know that 54GB still is usable!
- pekk 12y agoWhy don't people use multiple git repos for multiple internal projects? It seems totally nonsensical and undesirable.
- _ak 12y agowell, git gc --aggressive --prune=now, duh. (jk)
- deleted 12y ago[deleted]
- NAFV_P 12y agoNAFV_P@DEC-PDP9000:~$ python Python 2.7.3 (default, Feb 27 2014, 19:58:35) [GCC 4.6.3] on linux2 Type "help", "copyright", "credits" or "license" for more information >>> t=54*2**30 >>> t 57982058496 # let's assume a char is 2mm wide, 500 chars per meter >>> t/500.0 115964116.992 #meters of code # assume 80 chars per line, a char is 5mm high, 200 lines per meter >>> u=80*200.0 >>> v=t/u >>> v 3623878.656 # height of code in meters # 1000 meters per km >>> v/1000.0 3623.878656 # km of code, it's about 385,000 km from the Earth to the Moon >>> from sys import stdout >>> stdout.write("that's a hella lotta code\n")
- SnakeDoc 12y agoHey Facebook! You're doing it wrong!!!!!
- negativity 12y ago...but 8GB for the actual current version. How much of it is static resources, like CSS sprite images?