19 ms·
The largest Git repo
- vtbassmatt 9y agoA handful of us from the product team are around for a few hours to discuss if you're interested.
- CiPHPerCoder 9y agoSure, a couple of questions: 1. How do you measure "largeness" of a git repo? 2. How are you confident that you have the largest? 3. How much technical debt does that translate to?
- vtbassmatt 9y ago1. Saeed is writing a really nice series of articles starting here: https://www.visualstudio.com/learn/git-at-scale/ https://www.visualstudio.com/learn/git-at-scale/ In the first one, he lays out how we think about small/medium/large repos. Summary: size at tip, size of history, file count at tip, number of refs, and number of developers. 2. Fairly confident, at least as far as usable repos go. Given how unusable the Windows repo is without GVFS and the other things we've built, it seems pretty unlikely anyone's out there using a bigger one. If you know of something bigger, we'd love to hear about it and learn how they solved the same problems! 3. Windows is a 30 year old codebase. There's a lot of stuff in there supporting a lot of scenarios.
- deleted 9y ago[deleted]
- MikusR 9y agoIs it possible to checkout (if not build) something like Windows 3.11 or NT 4?
- hyperrail 9y agoAs far as I can recall, this is not possible using Windows source control, as its history only goes back to the lifecycle of Windows XP (when the source control tool prior to GVFS was adopted). Microsoft does have an internal Source Code Archive, which does the moral equivalent of storing source code and binary artifacts for released software in a underground bunker. I used to have a bit of fun searching the NT 3.5 sources as taken from the Source Code Archive...
- qznc 9y agoI recently heard a story that someone tried to push a 1TB repo to our university Gitlab which then ran out of disk space. Sure, that might have been not be a usable repo but only an experiment. Still, I would bet against the claim that 300GB is the largest one.
- tormeh 9y ago1 TB of code?
- BenjiWiebe 9y agoI'd sure like to run that as my operating system, browser, virtual assistant, car automation system, and overall do-everything-for-me system...
- ethomson 9y ago300 GB is not the size of _the repository_. It's the size of the code base - the checked out tree of source, tests, build tools, etc - without history. It's certainly possible that somebody created a 1 TB source tree in Git, but what we've never heard of is somebody actually _using_ such a source tree, with 4000 or more developers, for their daily work in producing a product. I say this with some certainty because if somebody had succeeded, they would have needed to make similar changes to Git to be successful, though of course they could have kept such changes secret.
- rajathagasthya 9y agoVery cool blog! As I understand, you dynamically fetch a file from the remote git server once for the first time I open the file. Do you do any sort of pre-fetching of files? For example, if a file has an import and uses a few symbols from that file, do you also fetch the imported file beforehand or just fetch it when you access it first time?
- deleted 9y ago[deleted]
- saeednoursalehi 9y agoWe don't currently do that sort of predictive prefetching, but it's a feature we've thought a lot about. For now, users can explicitly call "gvfs prefetch" if they want to, or just allow files to be downloaded on demand.
- vtbassmatt 9y agoFor now, we're not that smart and simply fetch what's opened by the filesystem. With the cache servers in place, it's plenty fast. We do also have an optional prefetch to grab all the contents (at tip) for a folder or set of folders.
- muglug 9y agoWhat's the PR review UI built in?
- vtbassmatt 9y agoCustom JQuery-based framework, transitioning to React.
- taylorlafrinere 9y agoActually, I think we finished the conversion to React :). So, React.
- vtbassmatt 9y agoTaylor is the dev manager for that area so I'm inclined to believe his correction :)
- kingbirdy 9y agoThe article mentions relying on a windows filesystem driver. Two questions about that: 1) Why include it in default windows? It seems that 99.99% of users would never even know it existed, let alone use it 2) Does that mean GVFS isn't useable on *nix systems? Any plans to make it useable, if so?
- hirsin 9y agoIt's not included in Windows, which is why they have a signed drop of the driver for you to install. Even internally we have to install the driver. E- oops, missed that line. Cool, wonder if it will show up in more than Git
- saeednoursalehi 9y ago1) The file system driver is called GvFlt. If it does get included in Windows by default, it'll be to make it easier for products like GVFS, but GvFlt on its own is not usable by end users directly. 2) GVFS is currently only available on Windows, but we are very interested in porting to other platforms.
- ATsch 9y agoWhat are your thoughts on implementing something more general like linux's FUSE[1] instead? A general Virtual Filesystem driver in Windows could be used for a wide range of things and means you don't just have a single-purpose driver sitting around. [1] https://en.m.wikipedia.org/wiki/Filesystem_in_Userspace https://en.m.wikipedia.org/wiki/Filesystem_in_Userspace
- bluejekyll 9y agoI'm currently investigating using GitLFS for a large repo that has many binary and other large artifacts. I'm curious, did you experiment with LFS for prior to building GitVFS? Also, I know that there is an (somewhat) active effort to port GitVFS to Linux, do you know if any of the Git vendors (GitLab and/or GitHub) are planning to support GitVFS in their enterprise products?
- saeednoursalehi 9y agoYes we did evaluate LFS. The thing about LFS is that while it does help reduce the clone size, it doesn't reduce the number of files in the repo at all. The biggest bottleneck when working with a repo of this size is that so many of your local git operations are linear on the number of files. One of the main values of GVFS is that it allows Git to only consider the files you're actually working with, not all 3M+ files in the repo.
- bluejekyll 9y ago> One of the main values of GVFS is that it allows Git to only consider the files you're actually working with, not all 3M+ files in the repo. That is an excellent point. Thanks!
- sytse 9y agoWe at GitLab are looking at GitVFS but have not made a decision yet https://gitlab.com/gitlab-org/gitlab-ce/issues/27895 https://gitlab.com/gitlab-org/gitlab-ce/issues/27895
- sigil 9y agoThis is amazing, congrats. I worked on Windows briefly in 2005 (the same year git was released!) and was surprised at how well Source Depot worked, especially given the sheer size of the codebase and the other SCM tools at the time. Is there anything people particularly miss about Source Depot? Something SD was good at, but git is not?
- vtbassmatt 9y agoI just got a request today for an API equivalent to `sd files`, which is not something Git is natively great at without a local copy of the repo.
- taylorlafrinere 9y agoAnother interesting complaint is one that we hear from a lot of people who move from CVCS to DVCS: there are too many steps to perform each action. For example, "why do I have to do so many steps to update my topic branch". While we find that people get better with these things over time, I do think it would be interesting to build a suite of wrapper commands that roll a bunch of these actions up.
- hart_russell 9y agoWhat was the impetus for switching to git?
- vtbassmatt 9y agoMore or less: - Availability of tools - Familiarity of developers (both current and potential)
- chinhodado 9y agoWhy do you name it "GVFS" instead of something more descriptive like "GitVFS"?
- saeednoursalehi 9y agoThis was discussed a bit in the comments in Brian Harry's last post on GVFS: https://blogs.msdn.microsoft.com/bharry/2017/02/03/scaling-git-and-some-back-story/ https://blogs.msdn.microsoft.com/bharry/2017/02/03/scaling-g... We're building a VFS (Virtual File System) for Git (G) so GVFS was a very natural name and it just kind of stuck once we came up with it.
- Kostic 9y agoAre you aware that the name was already taken[0] for something which also has to do with file systems? [0] https://wiki.gnome.org/Projects/gvfs https://wiki.gnome.org/Projects/gvfs
- alkonaut 9y agoDoes the virtualization work equally well for lots of history as it does for large working copies? I have a 100k commit svn repo I have been trying to migrate but the result is just too large. Partly this is due to tons of revisions of binary files that must be in the repo. Does the virtualization also help provide a shallow set of recent commits locally but keep all history at the server (which is hundreds of gigs that is rarely used)?
- vtbassmatt 9y agoGVFS helps in both dimensions, working copy and lots of history. For the lots of history case, the win is simply not downloading all the old content. A GVFS clone will contain all of the commits and all of the trees but none of the blobs. This lets you operate on history as normal, so long as you don't need the file content. As soon as you touch file content, GVFS will download those blobs on demand.
- alkonaut 9y agoThanks - that sounds perfect for lots of binary history as you never view history on the binaries, only the source files.
- mschuster91 9y agoHow do you prevent data exfiltration? I mean, in theory you could restrict the visibility of repos to the user based on team membership/roles and so prevent a single person from unauditably exfiltrating the whole Windows source code tree. In contrast with a monorepo there likely won't be any alerts triggered if someone does do a full git clone, except for someone saturating his switch port...
- asdfgadsfgasfdg 9y agoWhat would someone do with the source code for Windows? No one in open source would want to touch it. No large company would want to touch it. Grey/black hats are probably happier with their decompilers. Surely it would be easier to pirate than build (assuming their build system scales with most build systems I've observed in the wild). No small company would want to touch it. Anyway MS share source with various third parties (governments at least and I believe large customs in general) so any of these are a potential leak source.
- vtbassmatt 9y agoThis is all correct. Also, we'd notice someone grabbing the whole 300GB whether it's in 40 SD depots or a single Git repo.
- e40 9y agoMind if I ask how you'd notice?
- DoofusOfDeath 9y ago"A handful of us from the product team are around for a few hours to discuss if you're interested." Thanks! This is a little off-topic, but why can't Windows 10 users conclusively disable all telemetry? (I consider the question only a little off-topic, because I have the impression that this story is part of an ongoing Microsoft charm-offensive.)
- gear54rus 9y agoHaha, no answer, as expected. HN got butthurt as well lol.
- DoofusOfDeath 9y agoI knew there was a risk of getting downvoted, but I was surprised that it went to "-4".
- colejohnson66 9y agoWhy not TFS?
- wilatmsft 9y agohttps://news.ycombinator.com/item?id=14411588 https://news.ycombinator.com/item?id=14411588
- e40 9y agoAny plans to port GVFS to Linux or macOS?
- tobyhinloopen 9y agoI wonder why Windows is a single repository - Why not split it in separate modules? I can imagine tools like Explorer, Internet Explorer/Edge, Notepad, Wordpad, Paint, etc. all can stay in its own repository. I can imagine you can even further split things up, like a kernel, a group of standard drivers, etc. If that is not already the case (separate repos, that is), are the plans to separate it in the future?
- howinator 9y agoSo, this is actually pretty common. I know that both Google and Facebook use a huge mono-repo for literally everything (except I think Facebook split out their Android code into a separate repo?). So, all of Facebook's and Google's code for front-end, back-end, tools, infrastructure, literally everything, lives in one repo. It's news to me that Windows decided to go that route too. Personally, I think submodules and git sub-trees suck, so I'm all for putting things in a monorepo.
- Jyaif 9y agopros to big repo: -dont have to spend time to think about defining interfaces cons: -history is full of crap you dont care about -tests take forever to run -tooling breaks down completely, though thanks to MS the limit was increased seriously
- bpicolo 9y agoAre the big monorepo companies actually waiting for global test suite completion for every change? I'd doubt that, I'm sure they're using intelligent tools to figure out what tests to actually run. Compute for testing is massively expensive at that scale so it's an obvious place to optimize
- puzzle 9y agoGoogle's build and testing system is smart in which tests to run, as you suspect, but it still has a very, very large footprint.
- breck 9y agoThis is so awesome. Brilliant move MS! In addition to enabling Windows engineers to be significantly more productive (eventually), it will go a long way to enabling engineers in other departments to contribute to Windows. For example, I used to work in the Azure org and once noticed a relatively simple missing feature in Windows. I filed a bug and was in contact with a PM who suggested if I wanted I could work on adding it myself. I dipped my toe in, but the onboarding costs were just too high and I quickly decided against it. With Windows on git, much more likely to have dived in.
- hyperrail 9y agoI'm not so sure moving to Git alone would have helped your case. Getting an enlistment is only a small part of contributing to Windows.
- vtbassmatt 9y agoTrue, but the move to Git is part of our larger "1ES" (One Engineering System) effort across the company. The idea is, if you know how to enlist/build/edit/submit in any team, you know how to do the same in any team.
- breck 9y agoAgreed, but probably one of the top 5 road blocks.
- vmasto 9y ago19 seconds for a commit (add + commit) might be long but the new improvements look promising (down to ~10s). (Please correct me if the COMMIT column in the perf table includes the staging operations.) This looks awesome. I just wish Facebook would also share some perf and time statistics on their own extensions for Mercurial, last time I checked their graphs were unitless.
- saeednoursalehi 9y agoIndeed, while 19 seconds for commit is far better than 30 minutes we would have seen without GVFS, it's way too slow to actually feel responsive while you're coding. And in fact, it was sometimes worse than 19 seconds because commands like status and add would generally get slower as you access and hydrate more files in the repo. With the big O(modified) update that we just made to GVFS, git commands no longer slow down as you access more files, so now our devs see a consistent commit time of around 10 seconds, and consistent and faster times for most other commands too.
- slededit 9y agoYou have to put this into perspective with what they are replacing. You'd never get a submit done in less than 19 seconds using the old source depot tools anyways. When you work on projects this big that takes hours to compile and minutes to incremental compile - responsiveness just isn't something you get to have at scale.
- vmasto 9y agoThis is mainly why I'm asking for data from Facebook. I've seen claims that their vcs operations (at least the most common ones) are near instant, but nothing official. It would appear that FB have solved the responsiveness problem with Mercurial but, again, no official data to back it up.
- sp332 9y agoArchive Team is making a distributed backup of the Internet Archive. http://archiveteam.org/index.php?title=INTERNETARCHIVE.BAK http://archiveteam.org/index.php?title=INTERNETARCHIVE.BAK Currently the method getting the most attention is to put the data into git-annex repos, and then have clients just download as many files as they have storage space for. But because of limitations with git, each repo can only handle about 100,000 files even if they are not "hydrated". http://git-annex.branchable.com/design/iabackup/ http://git-annex.branchable.com/design/iabackup/ If git performance were improved for files that have not been modified, this restriction could be lifted and the manual work of dividing collections up into repos could be a lot lower. Edit: If you're interested in helping out, e.g. porting the client to Windows, stop by the IRC channel #internetarchive.bak on efnet.
- gcb0 9y agointernet archive sounds like the best ever use case for IPFS
- sp332 9y agoIt was considered but it just didn't get enough attention from anyone to get it done. http://archiveteam.org/index.php?title=INTERNETARCHIVE.BAK/ipfs_implementation http://archiveteam.org/index.php?title=INTERNETARCHIVE.BAK/i...
- StavrosK 9y agoWouldn't IPFS be much, much more suitable for this purpose?
- Willamin 9y agoI don't know much about Windows development, but I'm sure the system is modularized in some way. Why wouldn't you want to break up the project into multiple repos for different parts of the system? That would let you work on and test each part independent of the rest. Each part should be able to function on its own, right? Of course some engineers would need to build and test the entire OS as a whole, but I'd wager that (for example) the team working on visual design of the settings app doesn't need to have the source code of how the login screen verifies passwords. Clearly Microsoft's process works well enough for them, so I wonder what benefits there are to using the monolithic repo choice over many smaller repos.
- domoritz 9y agoGoogle has a single repo. The advantages are that you don't need to version anything because you always build against head. It's awesome but requires some discipline and good infrastructure.
- sseveran 9y agoIf you have many small repos for a large interconnected project you simply move the complexity of managing a commit that requires changes into another tool that can manage cross repo changes and dependencies. With a single repo you can change something and build it, fix any breaks and then commit it with just source source control and build system. The many small repos has in my experience been driven by either poor processes or tooling limitations.
- bousaid 9y agoWindows was developed a long time ago, and I'm guessing components were never fully separated as the codebase grew larger everyday.
- JimA 9y agoNon-dev here, but does this replace/overlap with TFS? What was the driver to adopt Git?
- taylorlafrinere 9y agoSo, TFS/VSTS is a suite of developer services. They fully support and integrate with git. In other words, git is a first-class citizen in TFS/VSTS. The centralized version control system in TFS/VSTS is called "Team Foundation Version Control" or TFVC. There were a bunch of drivers to move to git: 1. DVCS has some great workflows. Local branching, transitive merging, offline commit, etc. 2. Git is becoming the industry standard and using that for our VC is both a recruiting and productivity advantage for us. 3. Git (and it's workflow) helps foster a better sense of sharing which is something we want to promote within the company. There are more but those are the major ones.
- vtbassmatt 9y agoGood questions. TFS is a whole suite of services: 2 version control systems (TFVC and Git), work item tracking, build orchestration, package management, and more. VSTS is the roughly-analogous cloud-hosted version. I'd have to dig up the link: a few years ago our VP had a good blog post on why we chose to add a Git server to our offering. TFVC is a classic centralized version control system. When we wanted to add a distributed version control, we looked at rolling our own but ultimately concluded that it was better to adopt the de facto standard.
- FLGMwt 9y agoultimately concluded that it was better to adopt the de facto standard Thank you for that : )
- criddell 9y agoAt one point, I believe Microsoft was using a modified Perforce server for source code. Is that completely gone now?
- vtbassmatt 9y agoMost of the large Source Depot users have moved to Git or, like Windows, are in the process of moving to Git. Legacy stuff will probably live on in SD for a long time, possibly forever, for maintenance work.
- fmihaila 9y agoYou are thinking of Google.
- fmihaila 9y agoI meant that Google was the company who used Perforce in the past, and not Microsoft. Google isn't using it anymore either; they switched to their own thing named Piper. https://www.wired.com/2015/09/google-2-billion-lines-codeand-one-place/ https://www.wired.com/2015/09/google-2-billion-lines-codeand...
- vtbassmatt 9y agoSD was based on a very old Perforce as well.
- Kenji 9y agoI am kinda surprised that Microsoft doesn't use tfs - after all, it's their own version control system. But then again, we use tfs at work and not a day goes by on which I do not long for git.
- UK-AL 9y agoAll the modern development on VSTS is focused on git as well.
- taylorlafrinere 9y agoAnd, just to be clear, git is a first-class citizen in VSTS/TFS. We've fully embraced git as THE DVCS solution within VSTS/TFS. It is seen as a companion to TFVC.
- maxxxxx 9y agoHere is my cynical view: From what I know they have a history of not using their own tools. They didn't use SourceSafe, but Perforce. Then they made an effort to switch to TFS, realized that it sucks and moved on to git. You can see the same pattern in Windows desktop apps. They didn't use MFC for themselves, didn't use Winforms, used WPF only a little.
- smilekzs 9y agoUWP everywhere now though.
- maxxxxx 9y agoIs that true? Which larger app is written in UWP? Something like Office, Skype or Visual Studio.
- contextfree 9y agoWindows itself is gradually rewriting its shell components in UWP XAML.
- 9y ago
- falsedan 9y agoLooks like all of the charts were made in Excel… that's some dedication to staying on-brand!
- FLGMwt 9y agoHah, they made sure there was just enough styling left for you to notice.
- bhauer 9y agoAs opposed to what?
- niklasrde 9y agoHTML and co. It is a website..
- vxNsr 9y agoprobably also about familiarity. while they probably could have mashed together something using html and css, excel is all about nice looking tables and charts. If only they offered a way to embed them instead of needing to take low res screenshots...
- garyclarke27 9y agoLinus must be very proud - his favourite software Windows - now depends on GIT.
- geodel 9y agoI hope this largest repository has enough space for Clippy as Linus loves it.
- geodel 9y agohttps://plus.google.com/+LinusTorvalds/posts/5s9ZLWQDjwn https://plus.google.com/+LinusTorvalds/posts/5s9ZLWQDjwn
- sriram_sun 9y agoWell how the tables have turned! Only about 3 yrs back I was having a conversation with a Microsoft engineer about them evaluating a closed source Hadoop clone because Microsoft policy prohibited them from using open source.
- tonmoy 9y agoUsing open source in their product is different from using it for development. I wonder which one you talking about with that employee
- sriram_sun 9y agoI'm pretty sure it was for a product (server side), not an internal dev tool. Back in 2000 one of my roommates was an intern with the VS team and he said a lot of the devs were using emacs.
- vtbassmatt 9y agoDifferent divisions have had different stances on open source code for a long time. Somewhere I still have the t-shirt from our first "Open Source Day" event back in 2008 (and it's not like that was the first time any MS employee had ever considered using open source). Things are a lot more standardized now, with a big push from both the top and the bottom to use open source wherever it makes sense. Why reinvent the wheel?
- kk1274 9y ago300GB of code WOW! Just for comparison the entire English Wikipedia dump including all media is about 50-60GB. What are you guys doing there and how large do you see this growing?
- mschuster91 9y agoHmm. I believe it's likely there's also lots of binary assets - the WAV sound files, BMP images, the "hello world" videos, for example - and possibly also the raw versions of the assets. And if it's really the whole history of Windows in there, that's a LOT of binary assets in LOTS of versions.
- taylorlafrinere 9y agoMost of that 300GB isn't text. There are test assets, images, videos, built binaries, vhd's, etc. Also, I should be clear that that 300GB is just at tip (no history). We can debate about whether or not those things should be checked into the repo but they are there now.
- mschuster91 9y agoWhoa. 300GB with a shallow clone?! What size does the whole repo use on the server side?
- wilatmsft 9y agoThe pack file size for a full clone is 187GB. The 300GB is the working directory. We did not import the history of the code base, so the current repo only has about 5 months of history. As others have called out, there are a lot of assets in the repo that don't compress.
- bokchoi 9y agoWhy only 5 months? Will more of the history be added to the git repository eventually?
- MS_Buys_Upvotes 9y agoWow the majority of posts here are from Microsoft employees.
- scrollaway 9y agoSo? Why is that surprising?
- adoggman 9y agoHeh, the place I work at might have the single largest monolithic SVN repo. It works surprisingly well.
- deleted 9y ago[deleted]
- systems 9y agois GVFS portable to other OSes?
- vtbassmatt 9y agoBy design, yes. There are not (yet) implementations on other OSes.
- steve_avery 9y agoWhat was performance like for the Source Depot system? It would be interesting to note the comparison between the old SDX system and GVFS.
- wilatmsft 9y agoQuoting Brian Harry from a comment response at https://blogs.msdn.microsoft.com/bharry/2017/05/24/the-largest-git-repo-on-the-planet/ https://blogs.msdn.microsoft.com/bharry/2017/05/24/the-large... "It depends a great deal on the operation. SourceDepot was much faster at some things – like “sd opened”, the equivalent of “git status”. sd opened was < .5s. git status is at 2.6s now. But SD was much slower at some other things – like branching. Creating a branch in SD would take hours. In Git, it's less than a minute. I saw a mail from one of our engineers at one point saying they'd been putting off doing a big refactoring for 9 months because the branch mechanics in SD would have been so cumbersome and after the switch to Git they were able to get the whole refactoring done in a topic branch in no time. On an operation, by operation basis, SD is still much faster than our Git/GVFS solution. We're still working on it to close the gap but I'm not sure it will ever get as fast at everything. The broader question, though is about overall developer productivity and we think we are on a path to winning that."
- yeukhon 9y agoIf you have a large repo, it actually helps you to save storage if you don't need the full history on clone by using shallow clone. hg is still behind this, AFAI can tell from search. FB has this as an extension. https://bitbucket.org/facebook/hg-experimental/ https://bitbucket.org/facebook/hg-experimental/ Is FB fully using hg internally, or both Git and hg, because obviously FB has public repo on Github.
- isignal 9y agoMy understanding is that fb is actively promoting hg for internal repositories. Not sure how they sync between public git and internal hg.
- iamNumber4 9y agoThis just seems like the exact opposite of "Do one thing; and do it well".
- svanwaa 9y agoCan you go into any more detail of the breakdown of your repo structure? Thanks!
- vtbassmatt 9y agoedit: forgot, no Markdown here Do you mean across all of Microsoft? Different teams have different structures. Speaking only for TFS and VSTS, we have a single repo containing the code for both, a handful of "adjunct" repos containing tools like GVFS, a repo for the documentation [1], and a bunch of open source repos for the build and release agent [2], agent tasks [3], API samples [4], and probably more I don't know about. [1] https://www.visualstudio.com/docs https://www.visualstudio.com/docs [2] https://github.com/microsoft/vsts-agent https://github.com/microsoft/vsts-agent [3] https://github.com/Microsoft/vsts-tasks https://github.com/Microsoft/vsts-tasks [4] https://github.com/Microsoft/vsts-dotnet-samples https://github.com/Microsoft/vsts-dotnet-samples
- lloeki 9y agoComing from the days of CVS and SVN, git was a freaking miracle in terms of performance, so I have to just put things into perspective here when the topmost issue of git is performance. It's just a testament how huge are the codebases we're dealing with (Windows over there, but also Android, and surely countless others), the staggering amount of code we're wrangling around these days and the level of collaboration is incredible and I'm quite sure we would not have been able to do that (or at least not that nimbly and with such confidence) were it not for tools like git (and hg). There's a sense of scale regarding that growth across multiple dimensions that just puts me in awe.
- emodendroket 9y agoAt the risk of sounding like a downer, this was a migration of an existing codebase.
- mdekkers 9y agoI agree. I really think Linux needs a Nobel Price
- adrianN 9y agoFor what? Peace?
- comex 9y agoBroadly speaking this is true, but note that in some ways CVS and SVN are better at scaling than Git. - They support checking out a subdirectory without downloading the rest of the repo, as well as omitting directories in a checkout. Indeed, in SVN, branches are just subdirectories, so almost all checkouts are of subdirectories. You can't really do this in Git; you can do sparse checkouts (i.e. omitting things when copying a working tree out of .git), but .git itself has to contain the entire repo, making them mostly useless. - They don't require downloading the entire history of a repo, so the download size doesn't increase over time. Indeed, they don't support downloading history: svn log and co. are always requests to the server. Unfortunately, Git is the opposite, and only supports accessing previously downloaded history, with no option to offload to a server. Git does have the option to make shallow clones with a limited amount of (or no) history, and unlike sparse checkouts, shallow clones truly avoid downloading the stuff you don't want. But if you have a shallow clone, git log, git blame, etc. just stop at the earliest commit you have history for, making it hard to perform common development tasks. I don't miss SVN, but there's a reason big companies still use gnarly old systems like Perforce, and not just because legacy: they're genuinely much better at scaling to huge repos (as well as large files). Maybe GVFS fixes this; I haven't looked at its architecture. But as a separate codebase bolted on to near-stock Git, I bet it's a hack; in particular, I bet it doesn't work well if you're offline. I suspect the notion of "maybe present locally, maybe on a server" needs to be baked into the data model and all the tools, rather than using a virtual file system to just pretend remote data is local.
- deleted 9y ago[deleted]
- isignal 9y agoThose of us working on smaller codebases may wonder what the big deal is. Facebook had a similar problem leading them to switch out to mercurial. https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc... It is awesome that the problems could be solved in git itself. Also, kudos to the writer of the blog. It is a really high quality blog post. The percentile measures of performance, survey responses from users etc are very typical of solid incremental approaches to challenges faced by startups except these are internal customers.
- js2 9y agoWindows, because of the size of the team and the nature of the work, often has VERY large merges across branches (10,000’s of changes with 1,000’s of conflicts). At a former startup, our product was built on Chromium. As the build/release engineer, one of my daily responsibilities was merging Chromium's changes with ours. Just performing the merge and conflict resolution was anywhere from 5 minutes to an hour of my time. Ensuring the code compiled was another 5 minutes to an hour. If someone on the Chromium team had significantly refactored a component, which typically occurred every couple weeks, I knew half my day was going to be spent dealing with the refactor. The Chromium team at the time was many dozens of engineers, landing on the order of a hundred commits per day. Our team was a dozen engineers landing maybe a couple dozen commits daily. A large merge might have on the order of 100 conflicts, but typically it was just a dozen or so conflicts. Which is to say: I don't understand how it's possible to deal with a merge that has 1k conflicts across 10k changes. How often does this occur? How many people are responsible for handling the merge? Do you have a way to distribute the conflict resolution across multiple engineers, and if so, how? And why don't you aim for more frequent merges so that the conflicts aren't so large? (And also, your merge tool must be incredible. I assume it displays a three-way diff and provides an easy way to look at the history of both the left and right sides from the merge base up to the merge, along with showing which engineer(s) performed the change(s) on both sides. I found this essential many times for dealing with conflicts, and used a mix of the git CLI and Xcode's opendiff, which was one of the few at the time that would display a proper three-way diff.)
- draw_down 9y agoGod, that sounds hellish.
- malnourish 9y agoFor 3 way merging, I've had good luck with beyondcompare
- Peaker 9y agoWhen you have that many conflicts, it's often due to massive renames, or just code moves. If you use git-mediate[1], you can re-apply those massive changes on the conflicted state, run git-mediate - and the conflicts get resolved. For example: if you have 300 conflicts due to some massive rename, you can type in: git-search-replace.py[2] -f oldGlobalName///newGlobalName git-mediate -d Succcessfully resolved 377 conflicts and failed resolving 1 conflict. <1 remaining conflict shown as 2 diffs here representing the 2 changes> [1] https://medium.com/@yairchu/how-git-mediate-made-me-stop-fearing-merge-conflicts-and-start-treating-them-like-an-easy-game-of-a2c71b919984 https://medium.com/@yairchu/how-git-mediate-made-me-stop-fea... [2] https://github.com/da-x/git-search-replace https://github.com/da-x/git-search-replace
- deleted 9y ago[deleted]
- wodencafe 9y agoSo THIS is why they developed GVFS.
- quotemstr 9y agoI have tremendous respect for Microsoft pulling itself together over the past few years.
- ericfrederich 9y agoThis may be the thing that gets Google to switch. They like having every piece of code in a single repository which Git cannot handle. Now that it is somewhat proven, maybe Google will leverage GVFS on Windows and create a FUSE solution for Linux.
- twinge 9y agoGoogle already has a FUSE layer for source control: http://google-engtools.blogspot.com/2011/06/build-in-cloud-accessing-source-code.html http://google-engtools.blogspot.com/2011/06/build-in-cloud-a...
- farresito 9y agoThey use mercurial (or were), which is as good as git. In fact, I bet a lot of people at Google are happy to use mercurial instead of git, given git's bad reputation with its command line interface.
- Thaxll 9y agoThey don't use mercurial.
- sbuttgereit 9y agoYou're thinking of Facebook if I'm not mistaken.
- farresito 9y agoI had seen several sources that affirmed that Google used Mercurial, but I'm not sure to what extend, so I will retract it :-)
- Dunedan 9y ago> You also see the 80th percentile result for the past 7 days […] What'd be even more interesting to see is something like the 95th or 99th percentile, as showing that 80% of all operations finish in acceptable time is nice, but probably not what's necessary to have satisfied customers.
- bischofs 9y agoAny word on open sourcing parts of the windows OS now that MS is seeing the light? The head guys have to see the benefits by now. It says something that MS chose Git over anything proprietary that they developed.
- midnitewarrior 9y agoI'm guessing that would be a licensing nightmare. They must pay many companies for licensed technologies inside Windows, and many of those licenses likely wouldn't be compatible with open source licensing. All of their source code would have to go through legal review, some with each check in. I don't see that happening for legacy code.
- cyphar 9y agoIf you look at how long it took for Sun to make Solaris free software (and even then it wasn't truly free in some cases) I doubt Microsoft would ever consider spending that much time doing it.
- YeGoblynQueenne 9y agoSo, if windows engineers are using git now, who is using TFVC? That's Team Foundation Version Control- the original TFS version control engine.
- vtbassmatt 9y agoLots and lots of external customers, and a handful of internal folks. FWIW Windows was never on TFVC (at least not the main development group).
- a_imho 9y agoWhat is the opposite of dogfooding?
- YeGoblynQueenne 9y agoThat wouldn't be dogfooding - the windows team is not responsible for TFVC, innit.
- cryptonector 9y agoAt Sun Microsystems, Inc., (RIP) we have many "gates" (repos) that made up Solaris. Cross-gate development was somewhat more involved, but still not bad. Basically: you installed the latest build of all of Solaris, then updated the bits from your clones of the gates in question. Still, a single repo is great if it can scale, and GVFS sounds great! But that's not what I came in to say. I came in to describe the rebase (not merge!) workflow we used at Sun, which I recommend to anyone running a project the size of Solaris (or larger, in the case of Windows), or, really, even to much smaller projects. For single-developer projects, you just rebased onto the latest upstream periodically (and finally just before pushing). For larger projects, the project would run their own upstream that developers would use. The project would periodically rebase onto the latest upstream. Developers would periodically rebase onto their upstream: the project's repo. The result was clean, linear history in the master repository. By and large one never cared about intra-project history, though project repos were archived anyways so that where one needed to dig through project-internal history ("did they try a different alternative and found it didn't work well?"), one could. I strongly recommend rebase workflows over merge workflows. In particular, I recommend it to Microsoft.
- holtalanm 9y agowe use a rebase workflow in git at my current employer, and it is amazing. previous employer used a merge workflow (primarily because we didnt understand git very well at the time), and there were merge conflicts all the time when pulling new changes down or merging new changes in. It was a headache to say the least. As the integration manager for one project, I usually spent the better part of an hour just going through the pull requests and merge conflicts from the previous day. I managed a team that was on the other side of the world, so there were always new changes when I started working in the morning.
- cryptonector 9y agoYes! One of the most important advantages of a rebase workflow is that you can see immediately what upstream commits your conflict with, as opposed to some massive merge you have to go chasing branch history to figure out the semantics of the change in question. "Amazing" is right. Sun was doing rebases in the 90s, and it never looked back.
- manyoso 9y agoI would pay money to see a video camera of Linus' face reading this article. I think we'd probably get impossible new shades of the color red heretofore unknown to humanity.
- rbanffy 9y agoThe pain increases with the square of the number of files and to the fourth power of the dependencies between specific versions of them. I'm not sure a big repo is a wise thing, even though I understand it may make sense for multiple reasons for a company, understanding it may damage brains far more sophisticated than mammalian ones.
- cubano 9y agoIsn't this perhaps the greatest validation of Linus's design genius that what was initially a weekend project[0] has successfully scaled to this? They could no longer use their revision control system BitKeeper and no other Source Control Management (SCMs) met their needs for a distributed system. Linus Torvalds, the creator of Linux, took the challenge into his own hands and disappeared over the weekend to emerge the following week with Git. [0] https://www.linux.com/blog/10-years-git-interview-git-creator-linus-torvalds https://www.linux.com/blog/10-years-git-interview-git-creato...
- Sharlin 9y agoAnd Linux was a hobby project as well... now it runs most of the internet AND most of the personal computing devices on the planet.
- wfunction 9y ago> Isn't this perhaps the greatest validation of Linus's design genius that what was initially a weekend project has successfully scaled to this? I thought the entire point of the article was to show how git didn't scale, and how they're basically rewriting the project and changing it as much as necessary to make it scale. It's not like Linus designed git to scale as O(modified).
- happycube 9y agoOn the other hand, becoming something bigger out of his hands is validation in it's own right...
- skybrian 9y agoNo, not really. When it scaled enough for the Linux kernel he lost interest in scaling any further, and that's nowhere near what's needed for a monorepo (or a Linux distribution). And I recall that there are was fairly infamous tech talk at Google where he basically dismissed Google's concerns in an arrogant way. So I'd say it's mostly a validation of git and Github's popularity with developers, that other people were willing to put in so much work to improve it.
- nthcolumn 9y agoLinus Torvalds rocks. Windows sucks. Subversion was crap so he just made git instead. git beat out svn, TFS and all that other crap legions of overpaid engineers came up with (or what they didn't get source control???) because unix design philosophy and therein lies the lesson still unlearned for they hath loaded all their bloat into one repo. Windows. It sucks and it will forever suck because it sucks by design. Bill say 'Thank you Linus - I owe you sooo much because git is way better than the best I could do' I mean has there ever been worse software ever written than the stuff being loaded into git right now? Awful, awful garbage, creaking and reeking of dirty hacks, different for the sake of it designs, misshapen, bolted together, bloated, willfully annoying, antisocial, phone home, locking-in, full of resolutely, defiant ancient unfixed bugs, butt ugly, horrible UI, full of errors and meaningless error messages, incessant nagging and weird quirks, wtf folders, command line from hell and urgh... note pad ... and oh dear god I almost forgot mmc consoles and visual studio and inconsistent flows, viral load by the galactic shit tonne, complete and utter drivel makes me want to vomit every time I hear that sickening jingle and after all those gazillions of engineering hours an absolute world wonder of fail? Two guys working out of a garage could do better. :P (Windows sucks btw)
- BenjiWiebe 9y agoSadly we will have to quit blaming Bill Gates. I doubt he makes very many design decisions any more. :)
- nthcolumn 9y agoIf he had only listened to me and re-released Xenix open source with a decent WM we could have avoided all this unpleasantness but no, he had to listen to Monkeyboy. :/
- a_imho 9y agoAnd then there's Windows Subsystem for Linux as well.
- Myrmornis 9y agoRather than worry about getting `status` under 10 seconds, just focus on `diff --stat` and `diff --cached --stat`. Those two replace most uses for `status`.
- drawkbox 9y agoGame development also has very large files and codebases, Git LFS is sometimes not enough. This is great for everyone really but very nice for game development and larger codebases that might have lots of assets along with it. Microsoft is doing great work here and hope it makes it to bitbucket, github etc.
- faragon 9y agoI have Git repositories much larger than 300GB, for binary data. The title should be "the largest Git repo for source code", in my opinion. BTW, it is a nice thing Windows development being moved to Git SCM. What's Linus opinion on that? What a victory :-)
- graycat 9y agoSounds like a lot of good work. But, in "Git repo", what the heck is a "repo"? A repossession as in repossessing a car? In the OP with "Everything you want to know about Visual Studio ALM and Farming", what is ALM -- air launched missile? What do air launched missiles and "farming" have to do with Visual Studio? To Bill Gates and Microsoft: For my startup, I downloaded, read, indexed, and abstracted 5000+ Web pages from the Microsoft Web site MSDN. That took many months. Then I typed in the software for my startup, 24,000 programming language statements in Visual Basic .NET 4 and ADO.NET (Active Data Objects, for getting to the relational data base management system SQL Server) and ASP.NET (Active Server Pages, for building Web pages) in 100,000 lines of typing. For that work, all of it that was unique to me and my startup was fast, fun, and easy. Far and away the worst problem in my startup, that delayed my work for YEARS, was the poor quality of the technical writing in the Microsoft documentation. Some of the worst of the documentation was for SQL Server: Gee, I read the J. Ullman book on data base quickly and easily while eating dinner at the Mount Kisco Diner. But the Microsoft documentation was clear as mud. Just installing SQL Server ruined my boot partition: SQL Server would not run, repair, reinstall, or uninstall, and I had to reinstall all of Windows and all my applications and try again, more than once. Quickly I discovered that documentation of logins, users, etc. were a mess: Basically the ideas seemed to be old capabilities, attributes, authentication, and access control lists, but nothing from Microsoft was any help at all. Eventually via Google searches I discovered some simple SQL statements, I could type into a simple file and run with the SQL Server utility SQLCMD.EXE; that way I got some commands that worked for much of what I needed. Now those little files are well documented and what I use. For getting a connection string that worked, again the documentation was useless, and I tried over and over with every variation I could think of until, for no good reason, I got a connection string to work. Once I tried to get a new installation of SQL Server to recognize, connect to, and use a SQL Server database from the previous installation of that version of SQL Server, but the result just killed the installation of SQL Server. Again, once again, over again, yet again, one more time, far and away the worst problem in my startup is making sense out of Microsoft's documentation. I found W. Rudin, Real and Complex Analysis fast, fun, and easy reading; Microsoft's documentation was an unanesthetized root canal procedure -- OUCH! So, again, once again, over again, yet again, one more time, please, Please, PLEASE, for the sake of my work, Microsoft, and computing, PLEASE get rid of undefined terms and acronyms in your technical writing. Get them out. Drive them out. Out. Out of your writing. Out of your company. Out of computing. No more undefined terms and acronyms, none, no more. I can't do it. You have to do it. Then, DO IT.
- cryptonector 9y agoI see that Windows engineering uses a merge workflow. I wonder why. See other comments in this thread about rebasing.
- DigitalJack 9y agoMaybe I missed it, but I wish they would have compared times to what they were using before (Source Depot). I guess I would more specifically like to know what pain points drove Microsoft to even try such a massive change.
- red023 9y agoStill cant get over it hat M$ is now working with linux and git extensively. I have got this feeling they don't deserve to use it. But something good for others may come out of it for sure.
- nathan_f77 9y agoThis is pretty crazy. It's very hard to imagine working on a single codebase with 4,000 other engineers. > Another key performance area that I didn’t talk about in my last post is distributed teams. Windows has engineers scattered all over the globe – the US, Europe, the Middle East, India, China, etc. Pulling large amounts of data across very long distances, often over less than ideal bandwidth is a big problem. To tackle this problem, we invested in building a Git proxy solution for GVFS that allows us to cache Git data “at the edge”. We have also used proxies to offload very high volume traffic (like build servers) from the main Visual Studio Team Services service to avoid compromising end user’s experiences during peak loads. Overall, we have 20 Git proxies (which, BTW, we’ve just incorporated into the existing Team Foundation Server Proxy) scattered around the world. If I was a hacker, this paragraph would probably encourage me to study the GVFS source code and see if I can find some of these Git proxies. I have no idea how you would find them, but there might be some public DNS records. This sounds like some very new technology and some huge infrastructure changes, which are pretty good conditions for security vulnerabilities. What kind of bounty would Microsoft pay if you could get access to the complete source code for Windows? $100,000? [1] [1] https://technet.microsoft.com/en-us/library/dn425036.aspx https://technet.microsoft.com/en-us/library/dn425036.aspx
- DominikD 9y agoSorry to see SourceDepot (slowly) decommissioned. I loved it and since it was a Perforce fork, what I've learned was directly applicable when I started using P4 in my subsequent job. Perhaps I'm old fashioned but I really see little appeal in DVCSes. I liked Hg but in the long run it's going to be completely run over by Git so I'd rather not invest in it. I'm rambling, sorry.
- TorKlingberg 9y agoThis is both very cool and eerily reminiscent of MVFS and ClearCase. It's a huge change in Git, going from de-centralized to hyper-centralized. If I read it right, git status, git commit and even running a build or cat'ing a file may not work if your network or the central server is down. I hope they have thought hard about how to get the "Git proxy" for remote sites working well. If they end up with remote sites working on a 15 min - 1 hour old tree that will be very annoying.
- microcolonel 9y agoThe day has come that Microsoft employees are celebrating how good they're getting at running Linus Torvalds' source code management tool. Cats and dogs, flying pigs. Might be good to start work on a compatible client and server for FUSE-based systems (Linux, OpenBSD, macOS [with a FUSE kernel module]).
- debamitro 9y agoI am not sure, but doesn't this look like an open source version of what ClearCase provides?