9 ms·
Wow. I was expecting an interesting discussion. I was disappointed. Apparently the consensus on hacker news is that there exists a repository size N above which
by lbrandy 15y ago
Wow. I was expecting an interesting discussion. I was disappointed. Apparently the consensus on hacker news is that there exists a repository size N above which the benefits of splitting the repo _always_ outweigh the negatives. And, if that wasn't absurd enough, we've decided that git can already handle N and the repository in question is clearly above N. And I guess all along we'll ignore the many massive organizations who cannot and will not use git for precisely the same issue.
So instead of (potentially very enlightening conversation) identifying and talking about limitations and possible solutions in git, we've decided that anyone who can't use git because of its perf issues is "doing it wrong".
- jamespo 15y agoI would have liked this comment better if it came up with some solutions itself, although maybe it's not easy to solve?
- djtriptych 15y agoThat's the problem. It's NOT an easy problem to solve. A lot of posts on hn describing some problem elicit "Why, that's no problem at all!" responses or "That's the wrong problem to think about" responses. Honestly that mindset is often really useful in programming, but when we get a problem that doesn't have a shortcut and is relevent, conversation goes to shit. Because I guess that's when programmers normally go into a hole and brute-force brain it out. How to use mass comms to talk about a difficult open problem is, I suppose, itself an open problem.
- sek 15y agoGit was not designed for this. It comes out of the Linux kernel, where you need a secure hash of a segment to prevent compromise. For big projects you have submodules, you can only get a level higher later. In a company, you trust the sources. With Perforce you check out files and work with the part you want. It is a design decision and they could have known before.
- lincolnq 15y agoI had the same reaction as you. </meta> Stat'ing a million files is going to take a long time. Perforce doesn't have this problem because you explicitly check out files (p4 edit). (Perforce marks the whole tree read-only, as a reminder to edit the file before you save.) It seems like large-repo git could implement the same feature. You would just disable (or warn) for operations which require stat'ing the whole tree. Then the question is how to make the rest of the operations perform well -- git add taking 5-10 seconds seems indicative of an interesting problem, doesn't it?
- ori_b 15y agoIt seems to me that you could have a daemon that uses inotify to make operations O(changed) vs O(size).
- groby_b 15y agoWhich would also be tremendously useful for e.g. make.
- shabble 15y agoThere already exists tup: http://gittup.org/tup/ http://gittup.org/tup/ which does that sort of thing.
- groby_b 15y agoIt seems eminently obvious to me that having basically a "change log" for a (part of a) filesystem is something that's valuable independent of your build system, revision control system, whatnot. At least that's what I'd like to see - it's functionality that's orthogonal to those tools.
- cheez 15y agoOh my god, that would be awesome at the FS level.
- sek 15y agoYou have a point. It is just surprising when git was designed for the Linux kernel and we all here have a Github mindset.
- kenrik 15y agoWith that in mind it seems like there is a market for a git replacement for these huge repos.
- to3m 15y agoI fear the market would be small. It is my guess (though I have no proof) that most places with particularly large repositories have lots of binary files in them. It's hard to get a 15GB repository if you just have text. This sort of thing suggests a centralized check-in/check-out model, because binary files are difficult to merge sensibly, and nobody wants to spend terabytes of hard drive space storing the repository locally. And your centralized check-in/check-out needs, whatever scale they might be, are probably tolerably well served by one of the existing solutions.
- alttag 15y agoYes, but why is that a show stopper? It's a small market filled only with people who typically have large fist-fulls of cash and are dependent on version control. It's a small market, but companies in it have the resources for a good solution.
- gchpaco 15y agoBecause those companies generally already pay the $$$ for Perforce (which has any number of deeply terrifying, shiny red candylike self destruct buttons and makes git's user interface look kind) which for all its other faults handles this specific user case extremely well.
- joeyh 15y agoThe naive solution to binary files is a centralized model. But here's an alternate, fully distributed implementation: http://git-annex.branchable.com/ http://git-annex.branchable.com/ 15 gb is a tiny, tiny repo. I have a 7 tb repo "here" (really, spread amoung various drives, servers, S3, etc). :)
- forrestthewoods 15y ago
- akkartik 15y agoYou're right. I found the original email equally disappointing, though. It boils down to "We pushed the envelope on size, it's too slow, we'd like to speed it up." Well, duh. He uses the word 'scalability' early in the email, but shows no indication that he knows what it means. I'd love to hear if different operations slow down at different rates as the repo accumulates commits. Do they scale linearly, sublinearly, or superlinearly as the repo grows? Are there step functions at which there's a sudden dramatic slowdown (ran out of RAM, etc.)?
- lnguyen 15y agoIt's intentionally vague but with enough details that if you're actually in a position to help, you'll recognize what's going on and actually directly contact to get more information. You don't spill internal processes and configurations without some kind of disclosure agreements and certainly not in a public forum.
- zobzu 15y agoif you were working for a truly open company, you could :)
- lnguyen 15y agoI don't think Facebook is claiming to be. And as much as I'd like being truly open as an ideal, it falls apart when you're dealing with competition (not cooperation) and money. At best you try to keep things open enough.
- Retric 15y agoI don't see how Facebook's build needs to be kept secret. It's a purely internal process and while they might lose something by giving details they can also gain if someone suggests improvements. That said, there are plenty of tings they need to keep secret. EX: Letting anyone export FB's full social graph would be really stupid.
- 15y ago
- kinofcain 15y agoYour comment was at the top so I continued to read expecting to find a bunch of ignorant group think about how git is awesome and Facebook is dumb, but that's not really what's going on down below. I don't know what facebook's use case is, so I have no idea if their repositories are optimally structured. However, I've used git on a very large repository and ran into some of the same performance issues that they did (30+ seconds to run git status), so I don't think it's terribly hard to imagine they're in a similar situation. What we did to solve it is exactly what you're excoriating the people below for suggesting: we split the repos and used other tools to manage multiple git repos, 'Repo' in some situations, git submodules in others. However, we moved to that workflow mainly because it had a number of other advantages, not just because it made day-to-day git operations faster. I hope git gets faster, some of the performance problems described are things we saw too, but things are always more complicated and I see nothing below that looks like the knee-jerk ignorant consensus you're describing. Sometimes the answer to "it hurts when I do this" is "don't do that... because there's other ways to solve the same issue that work better for a number of other reasons and we haven't bothered fixing that particular one because most of the time the other way works better anyway."
- wisty 15y agoOn a simliar note, I've heard there of people who would hit the limit on fortran files, so they put every variable into a function call to the next file, which itself contained one function and a function call to the next file after that (if necessary). Making stuff modular is often a good idea.
- LearnYouALisp 15y agoCellular, modular, and interactive-odular!
- ay 15y agoIt is intuitively obvious that it is better to be rich and healthy than poor and ill. Sadly, the reality choices are neither. And you can not just split a repo. Hearing some ideas for that case would have been interesting. Solving a scaling problem by splitting it is, well, obvious. And, yes, I also ran github on a couple of projects at $work and the issues are real, seen them. So, if it hurts when I try to use git - the answer will be don't use git... But the conveniences are so tempting...
- robfig 15y agoIn violent agreement here. Git and HG: 1. Require you to be sync'ed to tip before pushing. 2. Cannot selectively check out files. The former means that in any reasonably sized team, you will be forced to sync 30 times a day, even if you are the only one editing your section of the source tree. The latter means that Joe who is checking in the libraries for (huge open source project) for some testing increases everyones repo by that much, forever, even if it's deleted later. Needless to say, the universal response is that I'm doing it wrong. Perforce 4 life! But seriously, it says that Google adopted Git for their repo --- does anyone know how they use it? I would expect them to want a linear history, but their teams are way too big to be able to have everyone sync'ed to tip to push...
- caf 15y agoRequire you to be sync'ed to tip before pushing. That's not the case. In fact, in the context of Linux kernel development, there's many emails on LKML where Linus is telling someone that they shouldn't be merging random-kernel-of-the-day into their development branch.
- robfig 15y agoIt's required to maintain a linear history, no?
- ori_b 15y agoGit is not used for their main repo. Git is used as a local cache for perforce where a branch roughly corresponds to a CL. Only subtrees of interest are checked out.
- adgar 15y ago> Git is used as a local cache for perforce where a branch roughly corresponds to a CL. Only subtrees of interest are checked out. That's a common use for git at Google, but not the only one (I'm a SWE). When I do use perforce I've got enough rhythm that it doesn't get in my way, but I really like git at Google for local branches on rapidly-changing subtrees. A lot of time I'll work on a branch to submit as a CL, but then realize I should do something else that depends on it. Perforce is a mess at this situation if the tree is changing much, and git is perfect if you just make a new branch.