19 ms·
Google Is 2B Lines of Code, All in One Place
- rbinv 11y agoThose are mind-boggling numbers. Although I kind of doubt that "almost every" engineer has access to the entire repo, especially when it comes to the search ranking stuff.
- kcorbitt 11y agoFrom the article: "There are limitations this system. Potvin says certain highly sensitive code—stuff akin to the Google’s PageRank search algorithm—resides in separate repositories only available to specific employees."
- dfc 11y agoDid you read the article? "There are limitations this system. Potvin says certain highly sensitive code—stuff akin to the Google’s PageRank search algorithm—resides in separate repositories only available to specific employees. "
- afandian 11y ago> Potvin says certain highly sensitive code—stuff akin to the Google’s PageRank search algorithm—resides in separate repositories only available to specific employees.
- foobar2020 11y ago> There are limitations this system. Potvin says certain highly sensitive code—stuff akin to the Google’s PageRank search algorithm—resides in separate repositories only available to specific employees. > (...) all 2 billion lines sit in a single code repository available to all 25,000 Google engineers.
- Lewisham 11y agoFWIW, apart from the previously mentioned sensitive stuff, we give engineering interns the same level of access we give full-time engineers. We keep things open because it makes things faster; we have an excellent code search tool that's great for navigating through the Piper repo (e.g. finding subclasses, finding uses of an API) which really speeds up dev time. When we're not talking about the sensitive stuff, there's not much magic to what many engineers write every day, it's the same "glue technology X to technology Y" stuff you see everywhere, so I don't think there's any value to hiding that in the name of secrecy.
- petra 11y agoI though Google's search advantage really speeds up software development. I wonder if they use the same principle in alphabet across other engineering disciplines - creating some unique knowledge tools, and how do they look.
- maximilianburke 11y agoHow are changes that affect sensitive code handled? Are the owners of that code on the hook for making any API updates that the person pushing the change can't make?
- Lewisham 11y agoHaving never worked on the secret sauce, I honestly don't know. There is a small team of people who tend to do many of the global refactors, I might expect that they are given special permission.
- packetslave 11y agoI'm a random engineer and I have access to probably 99% of those 2B lines of code. The restricted stuff is a pretty tiny subset of the whole.
- qq66 11y agoWhy wouldn't almost every engineer have access to almost the entire repo? Most of Google's code is only relevant to another company working at Google scale, such as Facebook, Microsoft, Apple, etc. These are the companies with deep pockets that are willing to spend lots of money to acquire technology that will help them compete with Google. But none of these companies will buy code that's been stolen from Google. The most famous corporate trade secret, the Coke formula, was stolen by two employees who attempted to sell it to Pepsi. Pepsi alerted Coke, the companies worked together to bring in the FBI, and both employees went to prison: http://www.cnn.com/2007/LAW/05/23/coca.cola.sentencing/ http://www.cnn.com/2007/LAW/05/23/coca.cola.sentencing/
- hellbanner 11y ago"LGTM is google speak for Looks good to me" - actually common outside of Google.
- malkia 11y agoSGTM
- emanuelsaringan 11y agoSounds SGTM to me.
- k33n 11y agoComparing "Google" to Windows isn't really a fair comparison. I'm sure all of the code that represents products that Microsoft has in the wild far exceeds 2B lines.
- guelo 11y agoAgree especially since Google's repo contains their version of almost the entire Microsoft Office suite.
- DannyBee 11y agoNote that this is just the monolithic repository. Google also has other non-piper repositories containing hundreds of millions of lines too :P For example, android and chrome are git based. Note also that when codesearch used to crawl and index the world's code, it was not actually that large. It used to download and index tarballs, svn and cvs repositories, etc. All told, the amount of code in the world that it could find on the internet a few years ago was < 10b lines, after deduplication/etc. So while you may be right or wrong, i don't think it's as obvious you are right as you do.
- ocdtrekkie 11y agoI'm still trying to figure out why having everything dumped in one big pile is something worth bragging about. I'd far rather have code sorted well into proper repositories.
- BooneJS 11y agoSo much is shared, though, right? Which is why Android is sorted into proper repositories but still has the 'repo' front-end wrapper to make sure you're getting the right versions of everything you need. If I wanted to change something fundamental, like I found a 10% speedup in Protobuf wire decode by changing the message slightly, there are likely very many services that all need it. Everyone at Google operates on HEAD. You're not allowed to break HEAD, and pre-submit/post-submit bots ensure you don't and will block your submit.
- lighthawk 11y ago"The two internet giants (Google and Facebook) are working on an open source version control system that anyone can use to juggle code on a massive scale. It’s based on an existing system called Mercurial. “We’re attempting to see if we can scale Mercurial to the size of the Google repository,” Potvin says, indicating that Google is working hand-in-hand with programming guru Bryan O’Sullivan and others who help oversee coding work at Facebook." Why Mercurial instead of Git?
- urda 11y agoBecause Google and Facebook are using Mercurial over Git internally. Edit: And for those that are just shocked that git isn't the answer. Facebook: https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc... Google: http://www.primordia.com/blog/2010/01/23/why-google-uses-mercurial-over-git/ http://www.primordia.com/blog/2010/01/23/why-google-uses-mer...
- foobar2020 11y agoHow do you know that?
- skj 11y agoHe/she doesn't. It's false.
- Lewisham 11y agoWell, Piper conforms to the Perforce API-ish, and Android and Chrome are both on Git. Mercurial was pushed internally as being the "better" (for some dimension of better) between it and Git back in 2010, but I think even the most hardline Mercurial fans have realized that in order to meet developers in the middle in 2015, we need to use Git for our open-source releases. We have a large investment in Gerrit [1] and Github [2] now. So the Mercurial comment is probably entirely based on scaling and replacement for the Piper Perforce API, rather than anything externally facing. [1] https://www.gerritcodereview.com/ https://www.gerritcodereview.com/ [2] https://github.com/google https://github.com/google
- sa2015 11y agoI wonder how close the "piper" system is to the code.google.com project.
- a1k0n 11y agoIIRC, Piper is a reimplementation of the perforce backend, in order to handle the code size and the sheer number of "changelists" submitted per second. Nothing to do with code.google.com.
- spectral321 11y agoThey are unrelated. :)
- DannyBee 11y agoI worked on code.google.com, i can tell you the are 100% unrelated. piper grew out of a need to scale the source control system the initial internal repositories were using code.google.com was a completely separate thing supporting completely different version control models, and a very different scale (very large number of small repositories, vs very small number of very large repositories)
- kuschku 11y agoThis explains quite some things. Still, this is not a very forward-thinking solution. Building and combining microservices – effectively UNIX philosophy applied to the web – is the most effective way to make progress. EDIT: Seems like I misunderstood the article – from the way I read it, it sounded like Google has a monolithic codebase, with heavily dependent products, deployed monolithically. As zaphar mentioned, it turns out this is just bad phrasing in the article and me misunderstanding that phrasing. I take everything back I said and claim the opposite.
- thomashabets2 11y agoThat's why Google is so unsuccessful at scaling technical solutions, unlike you they're not forward-thinking.
- kuschku 11y agoNo, it’s not that they are unsuccessful, it’s that they are unable to maintain it properly. Already today they have tons of open security issues. Or think about April 1st, when they set a Access-Control-Location: * header on google.com because someone wrote the com.google easteregg. Read the post from the SoundCloud dude from yesterday to find out how to do software management properly (hint: modularization is everything)
- t0mk 11y agolink to the post? Neither HN search nor Google search show anything to "modularization is everything".
- DanBC 11y agoPerhaps this? http://philcalcado.com/2015/09/08/how_we_ended_up_with_microservices.html http://philcalcado.com/2015/09/08/how_we_ended_up_with_micro...? (That question mark is part of the URL)
- 11y ago
- aikah 11y agolol git clone http://urlto.google.codebase.git http://urlto.google.codebase.git ... I wonder how much time it takes to clone the repo, provided they use git.
- robertk 11y agoIt's 80TB. You don't clone, just ask for views.
- ChuckMcM 11y agoI will say that I saw and experienced many things that changed my definition of 'large' at Google, but the most amazing was the source code control / code review / build system that kept it all together. The bad news was that it allowed people to say "I've just changed the API to <x> to support the <y> initiative, code released after this commit will need to be updated." and have that effect hundreds of projects, but at the same time, the project teams could do the adaptation very quickly and adapt. With the orb on their desk telling them at that their integration and unit tests were passing. I thought to myself, if there is ever a distributed world wide operating system / environment, it is going to look something like that.
- devit 11y agoThe solution to the excessive API change problem is to force whoever changes the API to fix all the consumers himself before the change is accepted. The Linux kernel generally uses this policy for internal APIs for example.
- Lewisham 11y agoWe do, mostly. Because Piper is a global repository, we have systems to do global safe refactors, and do so often. If the API changes drastically, there's usually a lengthy deprecation period before the API is switched over.
- revelation 11y agoI guess that works for the Linux kernel, but I would presume that for a large distributed operation like Google it would be much better to simply deprecate/version APIs and have the project teams update to a deadline. I mean, it's presumably impossible to have a single computer running a single OS build all of the Google software and run the testing.
- teraflop 11y agoGoogle is rumored to have an extremely powerful distributed compilation farm. It wouldn't surprise me if a single developer could make a change that affects the entire codebase and test it themselves.
- low_battery 11y agoDirect link to talk (The Motivation for a Monolithic Codebase ): https://www.youtube.com/watch?v=W71BTkUbdqE https://www.youtube.com/watch?v=W71BTkUbdqE
- Walkman 11y agoThis is crazy :D I have never heard tools and workflows like this.
- jfkw 11y agoHow do the monolithic repository companies handle dependencies on external source code? Are libraries and large projects e.g. RDBMS generally vendored/forked into the monolithic repositories, regardless of whether the initial intent is to make significant changes?
- jpollock 11y agoThere's typically a subdirectory called third_party, with subdirectories for each vendor, product and version. If the team is smart, they will also enact a rule saying "only one version". If you're really, really smart, local changes are kept as a set of patches, keeping them separate from the imported tar file. So, for source deliveries: third_party/apache/httpd/2.4/release.tgz /patch.tgz /Makefile (or other config) third_party/apache/httpd/2.2/release.tgz /patch.tgz ...
- cpeterso 11y agoFor example, here is Chromium's third_party directory: https://chromium.googlesource.com/chromium/src.git/+/master/third_party/ https://chromium.googlesource.com/chromium/src.git/+/master/...
- a3n 11y agoIn the spirit of "You didn't build that," I wonder how many lines of code comprise the binaries that Google binaries run on? Windows, Linux, network stacks, Mercurial, etc, etc. I also wonder if there's a circular relationship anywhere in there.
- Splines 11y agoIt's turtles all the way down, and also includes all the hardware and people.
- melling 11y agoI imagine that there's a lot of Java and C++. I do like Go but it makes you wonder if a more expressive language that requires a fraction of the code would be helpful. Maybe Steve Yegge will see Lisp at Google after all.
- astrange 11y agoHe claims to have stopped using it (#5): https://sites.google.com/site/steveyegge2/ten-predictions https://sites.google.com/site/steveyegge2/ten-predictions
- sytse 11y agoSo a monolithic codebase makes it easier to make an organization wide change. Microservices make it easier to have people work and ship in independent teams. The interesting thing is that your can have have microservices with a monolithic codebase (as Google and Facebook are comprised of many services). But you can also have a monolithic service with many codebases (like our GitLab that uses 800+ gems that live in separate codebases). And of course you can have a monolithic codebase with a monolithic service (a simple php app). And you can have microservices with diverse codebases (like all the hipsters are doing). I'm wondering if microservices force you to coordinate via the codebase just like using many codebases force you to coordinate via the monolithic service. Does the coordination has to happen somewhere? I wonder if early adopters of microservices in many codebases (SoundCloud) are experiencing coordination problems trying to change services.
- scrollaway 11y agoI'd be interested on an analysis of the meta-differences between those, but I think one of the main ones is a monolithic codebase makes it massively harder to open source components. I'm sure that, working on GitLab, you can agree with that - if GitLab were a massive, monolithic codebase and you wanted to open source specific parts of it, it'd be a huge pain.
- sytse 11y agoIndeed, if you have a monolithic codebase releasing your code for other people use is much harder. In git it is possible to create a separate repository for it (using subtree and filter-branch). But the harder part is versioning and release management. At Google they avoid the cost of versioning their services and the release management around that, but this makes it really hard to have the outside world use their code. Maybe this is hard anyway since their code probably depends on a lot of services common to Google (GFS, BigTable). This is probably why they can't release Borg/Omega but have to make Kubernetes instead.
- sytse 11y agoI think that advantages mentioned in the presentation about the monolithic codebase can be achieved if you have one source code server that is open to everyone (for example GitLab with most projects set to internal). Some of the tools will be easier to write for one repository than iterating over many, but that seems solvable. The biggest advantage seems to be that when you are an author of a dependency you can propose upgrades to all services that use your application. It is not clear to me but it seems that for small changes you can just force that change on the code owners. This ensures that the dependency author incurs the cost of a change (as is done for API changes in the Linux kernel) and that you do not need to version the API of the dependency. Interestingly Google recently started marking API's private by default. So they are moving in the direction of explicit API management. As soon as you work with people that are outside your control (as is common in open source) you would need to version the API as well in my opinion.
- sytse 11y agoSo a monolithic codebase makes it easier to make an organization wide change. Microservices make it easier to have people work and ship in independent teams. The interesting thing is that your can have have microservices with a monolithic codebase (as Google and Facebook are comprised of many services). But you can also have a monolithic service with many codebases (like our GitLab that uses 800+ gems that live in separate codebases). And of course you can have a monolithic codebase with a monolithic service (a simple php app). And you can have microservices with diverse codebases (like all the hipsters are doing). I'm wondering if microservices force you to coordinate via the codebase just like using many codebases force you to coordinate via the monolithic service. Does the coordination has to happen somewhere? I wonder if early adopters of microservices in many codebases (SoundCloud) are experiencing coordination problems trying to change services.
- sytse 11y agoThe CitC filesystem is very interesting. This is local changes overlaid on top of the full Piper repository. Commits are similar to snapshots of the filesystem. Sounds similar to https://github.com/presslabs/gitfs https://github.com/presslabs/gitfs
- sandGorgon 11y agoWhat are the best practices to follow in a single-repo-multiple-projecrs world? Some people recommend git submodule, others recommend subtree. How do you guys manage alerts and messages - does every developer get a commit notification,or is there a way to filter out messages based upon submodule. How does branching and merging work? I'm wondering what processes are used by non-Google/FB teams to help them be more productive in a monolithic repo world.
- ajross 11y agoFWIW: git submodules are not a single repo by definition. It's just a way to automate the checkout of specifically-versioned external projects without requiring hackery like packing tarballs into the project source. It has its uses, but it's definitely not what they're talking about here.
- luckydude 11y agoAgree 100%. Git submodules are for tracking other stuff, not for doing dev on that other stuff. If you would like to see how things would work with submodules that behaved just like files behave (full distributed workflow) we've got a (unfortunately commercial) solution here: http://www.bitkeeper.com/nested http://www.bitkeeper.com/nested
- cmrdporcupine 11y agoGenerally branching isn't really a thing at Google. Work is done at the code review level per change list ("CL"). Most changes happen through incremental submission of reviewed CLs, not by merging in feature branches. Every CL must run the gauntlet of code review, as well as can not usually be submitted without passing tests. There are rare cases where branching is used, but not commonly. As for notifications, the CL has a list of reviewers and subscribers. If you want to see code changing, you watch those CLs. Most projects have a list where all submitted CLs go.
- sandGorgon 11y agoCan you explain this a little more - what is a CL vs a changeset...and what do you mean by watching changelists. It sounds like you're subscribing to specific commits...but I'm talking about more at a project/directory level within the monolithic repo.
- 727374 11y agoReally? This article sounds very over simplified, but I haven't worked at google so I wouldn't know. I'm assuming if you want to change some much depended on library, there's a way to up the version number so you don't hose all your downstream users. That's the way it worked at Amazon at least. Also, I wonder why the people in the story think Google's codebase is larger than that of other tech giants, not that it really matters.
- rictic 11y agoIt's incumbent upon the person updating the library to get all users migrated to the new one. There are a few strategies for doing this though, including temporarily having two versions of the library. There are also tools for making large scale changes safely and quickly.
- zBard 11y agoLast I heard Google is still on Java 7 precisely because of this, although that might have changed. It's fun seeing the different theologies at Amazon and Google - I remember Yegge's famous platform rant, and he highlighted the Amazon versioned-library system as something which it did better than Google.
- jsolson 11y agoGoogle mostly works at HEAD. Very little is versioned, and branches are almost unheard of. In general you change the much depended on library and all of its consumers (probably over time in multiple changes, but you can do it in one go if it really needs to be a single giant change).
- devinj 11y agoThe whole point of one big repository is being able to avoid versioning and always work at head.
- yongjik 11y agoOne humorous side-effect of having all that code viewable (and searchable!) by everyone was that the codebase will contain whatever typo, error, or mistake you can think of (and convert into a regular expression). I remember seeing an internal page with dozens of links for humorous searches like "interger", "funciton", or "([A-Z][a-z]+){7,} lang:java"...
- wetmore 11y ago> "([A-Z][a-z]+){7,} lang:java Yeah this one was my favorite of the code search examples, there are some really good ones in there.
- cag_ii 11y agoCan you explain this? It looks to me like a regexp that searches Java source for words 7+ characters that start with a capital letter?
- yongjik 11y agoIt searches for CamelCase identifiers that are made of seven or more "terms", where each term is a capital letter followed by one or more lowercase letters. E.g., ProjectPotatoLoginPageBuilderFactoryObserver. (Disclaimer: I just made it up. Not an actual Google project name.) "lang: java" is not a part of regexp; just a Google code search extension that searches for Java.
- robryk 11y agoThis searches for camelcase identifiers with at least 7 words.
- nandhp 11y agoAnd then you killed off the public version and you keep that fun (and useful) toy to yourself. (But as great as Google Code Search was, my grudge is because of Reader.)
- bubersson 11y ago
- Apocryphon 11y agoLooks like someone's going to have to update this: http://www.informationisbeautiful.net/visualizations/million-lines-of-code/ http://www.informationisbeautiful.net/visualizations/million...
- buro9 11y agoThis hurts just thinking about what the build, test and deploy systems must look like.
- jsolson 11y agoWell, for build take a look at bazel, although attach it to a cluster of machines that can all read from Piper.
- dekhn 11y agoI'm a google software engineer and it's nice to see this public article about our software control system. I think it has plusses and minuses, but one thing I'll say is that when you're in the coding flow, working on a single code base with thousands of engineers can be an intensely awesome experience. Part of my job- although it's not listed as a responsibility- is updating a few key scientific python packages. When I do this, I get immediate feedback on which tests get broken and I fix those problems for other teams along side my upgrades. This sort of continuous integration has completely changed how I view modern software development and testing.
- nevir 11y agoBeing able to make sweeping changes to a shared piece of code, and ensure that everyone's up to date (Hi, Rosie!) and not broken by your change (yay TAP train!) is phenomenal as well.
- Touche 11y agoHow is this a side effect of it being in the same repository/
- nevir 11y agoSay I make a change to a commonly used library (let's say deprecating a function, and replacing it with another): * I can see literally _every use_ of the old function. * I can run the tests for everyone who uses that function. * * this is automated; the build/test tooling can figure out the transitive set of build/test targets that are affected by such a change. * I can (relatively) easily update _every use_ of the deprecated call with the new hotness * I can do that all within the same commit (or set of commits, realistically) --- None of this is impossible with multiple repos, it's just a lot more difficult to coordinate
- dekhn 11y agoTechnically, you can't see every use of a function because dynamic dispatch mechanisms aren't available at source code time. That's why you would normally run tests of all your library's dependents (having this dependency graph is one of the most important parts of the piper/blaze system).
- kazinator 11y agoI am unable to believe that Google has 2B lines of original code written from scratch at Google. Maybe they are counting everything they use. Somewhere among those 2B lines is all the source code for Emacs, Bash, the Linux kernel, every single third-party lib used for any purpose, whether patched with Google modifications or not, every utility, and so on. Maybe this is a "Google Search two billion" rather than a conventional, arithmetic two billion. You know, like when the Google engine tells you "there about 10,500,000 results (0.135 seconds)", but when you go through the entire list, it's confirmed to be just a few hundred.
- sp332 11y agoYes, that is counting everything. It's "the software needed to run all of Google’s Internet services" (so probably not Emacs, but the other stuff). But it's all in the repo, and it all has to be maintained.
- roxmon 11y agoGoogle has been around for 17 years and employees roughly 10,000+ software developers. I think it's reasonable to assume that the 2B LOC metric is accurate...
- hk__2 11y agoWindows has been around for 35 years and Microsoft had 61,000+ employees (ok, that’s not only software developers and they don’t work only on Windows) in 2005; and it’s only ~50M LOC. I don’t think the number of years + developpers really show something; you don’t write new code everyday.
- scott_s 11y agoYou pointed it out yourself, but I think you underestimated its importance: Microsoft works on many other things. Office, XBox, Windows Phone, Exchange, SQL Server, .Net, etc. I suspect Microsoft's total line count is similar to Google's. The difference, however, is that it's not one codebase.
- 11y ago
- ksk 11y agoIts interesting that they compare LoC with Windows. I suppose that this article wants us to be amazed at those numbers. However, my experience with Google's products indicates a gradual decline in performance and a simultaneous gradual increase in memory bloat (Maps, Gmail, Chrome, Android). Which ironically, FWIW, hasn't been the case with Windows. I have noticed zero difference in performance going from Windows 7 to 8 to 10.
- branchless 11y agoI'd have to disagree with this. First the baseline: windows is very slow. Second I found later versions slower. Third (and most maddening) every version of windows I've ever used has gotten slower over time (including not installing new s/w and defragmenting).
- ocdtrekkie 11y ago8.1 and 10 run incredibly well even on very old hardware. I will agree a given Windows install may feel slower over time, and it makes sense to rebuild the PC occasionally, though that may, again, be less so with 8.1 and 10.
- branchless 11y agoOr if you are a non-techie it means forced upgrade is built into your product.
- ocdtrekkie 11y agoI think Microsoft has tried to address this with 8/10, which have a "refresh" feature, which tries to clean out everything besides Modern apps and your personal data.
- branchless 11y agoGlad they are addressing a fundamental flaw in windows albeit not in the root cause (the need for a "refresh").
- dblock 11y agoA giant repo works for Google, and works for Facebook, and Microsoft, but it's bad for the development community at large. If you start centralizing your development you’re killing any type of collaboration with the outside world and discouraging such collaboration between your own teams. http://code.dblock.org/2014/04/28/why-one-giant-source-control-repository-is-bad-for-you-and-facebook.html http://code.dblock.org/2014/04/28/why-one-giant-source-contr...
- wgpshashank 11y agoCool , How much front and back end each ?
- antics 11y agoJust because people are talking about it: I work at MSFT, and the numbers Wired quotes for the lines of code in Windows are not even close to being correct. Not even in the same order of magnitude. Their source claims that Windows XP has ~45 million lines of code. But that was 14 years ago. The last time Windows was even in the same order of magnitude as 50 million LOC was in the Windows Vista timeframe. EDIT: And, remember: that's for _one_ product, not multiple products. So an all-flavors build of Windows churns through a _lot_ of data to get something working. (Parenthetical: the Windows build system is correspondingly complex, too. I'll save the story for another day, but to give you an idea of how intense it is, in a typical _day_, the amount of data that gets sent over the network in the Windows build system is a single-digit _multiple_ of the entire Netflix movie catalog. Hats off to those engineers, Windows is really hard work.)
- 4ad 11y agoSo, how many lines of code does it have?
- antics 11y agoI can't say. I work here, but I don't speak for the company.
- deleted 11y ago[deleted]
- fishtoaster 11y agoThat figure is common knowledge, not new information that antics is sharing with us. https://en.wikipedia.org/wiki/Source_lines_of_code#Example https://en.wikipedia.org/wiki/Source_lines_of_code#Example
- deleted 11y ago[deleted]
- antics 11y ago
- deleted 11y ago[deleted]
- therealmarv 11y agoWhat? This is surpassing the mouse genom complexity. See this charts for comparison: http://www.informationisbeautiful.net/visualizations/million-lines-of-code/ http://www.informationisbeautiful.net/visualizations/million...
- MrBra 11y agoAm I the only one who initially read 28 instead of 2B ? :)
- MrBra 11y agodownvoter: laughter is good for your health.
- therealmarv 11y agoSo they do not suffer on git submodules I guess
- h1fra 11y agoThe comparison with windows really is just here to provide a something to compare for casual reader, it's not really that good. An OS is a huge project. But google has hundred of different project, apis, library, framework... Even unix with an "unlimited" source of developpers does not reach that point.
- jakub_g 11y agoSome questions that immediately come to my mind: - What is the disk size of a shallow clone of a repo (without history)? - Can each developer actually clone the whole thing, or you do partial checkout? - Does the VCS support a checkout of a subfolder (AFAIK mercurial, same as git, does not support it)? - How long does it take to clone the repo / update the repo in the morning? Since people are talking about huge across-repo refactorings, I guess it must be possible to clone the whole thing. Facebook faces similar issues as Google with scaling so they wrote some mercurial extensions, e.g. for cloning only metadata instead of whole contents of each commit [1]. Would be interesting to know what Google exactly modified in hg. [1] https://code.facebook.com/posts/218678814984400/scaling-mercurial-at-facebook/ https://code.facebook.com/posts/218678814984400/scaling-merc...
- lrem 11y agoIn practice: none of these operations take long enough to tempt you into alt-tabbing to cat videos.
- thrownaway2424 11y agoMost of your questions don't apply to the system described in this article. You do not clone the repository, you merely chdir into a vfs that is backed by a consistent view of the repository at a point in time, which view is served from a large distributed service that lives in Google datacenters alongside other Google services like Search, Maps, and Gmail. Because it is enormous and nobody clones it, it is also true that nobody partially clones it. You do not "checkout a subfolder" either. Your last point is the only one that applies. If you want your view to advance from revision 123 to revision 125 it takes about a second to do so. If you have pending (not yet submitted) changes in your client, they might have to be merged with other changes, which can take a bit longer. If you have a really huge pending change, and your client is way behind HEAD, it might take a few tens of seconds to merge everything.
- bruckie 11y agoMost of these questions are answered in the talk. The tl;dr is that you don't clone or check out anything at all: instead, you use CitC to create a workspace, and the entire repository is magically available to you to view or edit. This model precludes offline work, of course. But that's not much of a problem in practice.
- amelius 11y agoIs this article saying that all developer employees have access to the "holy" search algorithm internals? I can hardly believe that to be true, given the fact that SEO is a complete industry.
- deleted 11y ago[deleted]
- jsolson 11y agoIt is not saying that. FTA: > There are limitations this system. Potvin says certain highly sensitive code—stuff akin to the Google’s PageRank search algorithm—resides in separate repositories only available to specific employees. The vast majority of code is visible to everyone, though.
- shampine 11y agoNo, it specifically says the opposite: "Potvin says certain highly sensitive code—stuff akin to the Google’s PageRank search algorithm—resides in separate repositories only available to specific employees."
- enf 11y agoOnce upon a time it was all in one repository. Shortly after I started there in late 2005, the "HIP" source code (high-value intellectual property, I think it stood for) was moved to its own source tree, with only precompiled binaries available to the rest of the company. Looks like there is a Quora question that mentions this too: https://www.quora.com/How-many-Google-employees-can-read-access-all-of-the-source-code-for-Googles-search-engine https://www.quora.com/How-many-Google-employees-can-read-acc...
- deleted 11y ago[deleted]
- Strikingwolf 11y agoReally interesting article. Sounds like a great solution to the problem in git of submodules. Definitely worth looking at. Thanks for posting OP. IMO this system would best be suited for large companies, but I could see the VCS that they are developing being used by anyone if it gets a github-esque website.
- Sven7 11y agoNow I know why my google plus page takes half a day to load.
- temuze 11y agoAssuming these numbers are right... (15 million lines of code changed a week) / (25,000 engineers) = 600 LOC per engineer per week Is ~120 LOC per engineer per workday normal at other companies?
- ajg360 11y agoI write between 4-600 lines of code a day where I work... I feel that 120 LOC is a day is on the smaller side (of what I'm used to anyway).
- xur17 11y agoIt really depends on what you're writing. Lower level c / c++, doubtful. Python, javascript, java, etc, yeah, it's believable.
- _delirium 11y agoElsewhere in this thread it's mentioned that Google makes use of large-scale, automated refactoring tools: http://research.google.com/pubs/pub41342.html http://research.google.com/pubs/pub41342.html Would be interesting to know what percentage of the total LoC touched are typically from that kind of automated refactor. Depending on the codebase, you can touch a ton of lines of code in a very small amount of time with those tools.
- dchichkov 11y agoI remember somebody wise had said once: "Every line of code is a constraint working against you."
- linkydinkandyou 11y agoIf they had done it in LISP, it would have only been 200K lines.
- sshumaker 11y agoXoogler here. There were tons of benefits to Google's approach, but they were only viable with crazy amounts of tooling (code search, our own version control system, the aforementioned CitC, distributed builds that reused intermediate build objects, our own BUILD language, specialized code review tools, etc). I'd say the major downside was that this approach basically required a 'work only in HEAD' model, since the tooling around branches was pretty subpar (more like the Perforce model, where branches are second-class citizens). You could deploy from a branch but they were basically just cut from HEAD immediately prior to a release. This approach works pretty well for backend services that can be pushed frequently and often, but is a bit of a mismatch for mobile apps, where you want to have more carefully controlled, manually tested releases given the turnaround time if you screw something up (especially since UI is really inefficient to write useful automated tests around). It's also hard to collaborate on long-term features within a shipping codebase, which hurts exploration and prototyping.
- nulltype 11y agoCould you elaborate how the single repo model causes that thing you said in the last sentence?
- nootropicdesign 11y agoOMG it's all in one file? OMG OMG it's all on ONE LINE????!!!
- makecheck 11y agoI really wish there was a tendency to track all change/activity and not just total size; maybe like the graphs on GitHub. Removing things is key for maintenance and frankly if they haven't removed a few million lines in the process of adding millions more, they have a problem. Having a massive code base isn't a badge of honor. Unfortunately in many organizations, people are so sidetracked on the next thing that they almost never receive license to trim some fat from the repository (and this applies to all things: code, tests, documentation and more). It also means almost nothing as a measurement. Even if you believe for a moment that a "line" is reasonably accurate (and it's tricky to come up with other measures), we have no way of knowing if they're measuring lots of copy/pasted duplicate code, massive comments, poorly-designed algorithms or other bloat.
- dekhn 11y agoAlthough I agree that line counting is a silly exercise much of the time, the talk did cover change activity as well as total size. With regard to copy/pasted duplicate code and massive comments, we do have ways of knowing that as both of those are easily computable. Duplicate code can be matched using hashes and comments are delimited, making their measurement easy.
- nhaehnle 11y agoThe article claims 2 billion lines of code across 25000 engineers, which boils down to 80k lines of code per engineer. I'm not sure what to think about that. It seems to be in a reasonable order of magnitude for C++/Java-type languages compared to projects that I have seen, but it does imply a significant chunk of code that is not actively being worked on for a long time (which is not necessarily a bad thing - don't change a running system and all that).
- brozak 11y agoThe comparison of Windows to all of Google's services is pointless and misleading. It's like comparing the weight of a monster truck and the total weight of all the cars at a dealership...
- nemesisrobot 11y agoThe comparison bewteen the total LOC across all of Google's products against just one of Microsoft's is a bit unfair.
- Locke1689 11y agoWhat I'd like to know and no one seems to mention: What's the experience like for teams not running a Google service and instead interacting with external users and contributors, e.g. the Go compiler or Chrome.
- bruckie 11y agoMany larger external projects are hosted in other repositories (Chrome and Android are well-known examples). Smaller stuff (like, say, tcmalloc or protocol buffers) is usually hosted in Piper and then mirrored (sometimes bidirectionally) to an external repository (usually GitHub these days).
- Locke1689 11y agoThanks, but I guess I was asking more about how this affects the other development characteristics described. You still have to deal with the massive repository and infrastructure, but if you're Go, for example, and you want to change an API 1) you can't see the consumers because many or most won't be Google-internal, and 2) even if you could see them, you can't change them. Even the build/test/deploy systems are somewhat compromised because you can't rely on all builders of your components being Google employees and having access to those resources. So in these scenarios, what does Google's infrastructure buy you, if anything? And if it doesn't buy you anything, how does that influence Google culture? Are teams less willing to do real open development due to infrastructure blockage?
- skybrian 11y agoWorking with multiple source control systems, multiple issue trackers, and multiple build systems has its challenges. It's true that you don't know about all callers if you're working on open source software. There's no magic there; you need to think about backward compatibility. (On the other hand, if it's a library, your open source users can usually choose to delay upgrading until they're ready, so you can deprecate things.) The main advantage for an open source project is that, though you don't know about all callers, you still have a pretty large (though biased) sample of them. If you want to know how people typically use your API's, it's pretty useful. Running all the internal tests (not just your own, but other people's apps and libraries) will find bugs that you wouldn't find otherwise. There were changes I wouldn't have been confident making to GWT without those tests, and bugs that open source users never saw in stable releases because of them. On the other hand, there were also changes I didn't make at all because I couldn't figure out how to safely upgrade Google, or it didn't seem worth it.
- wellsjohnston 11y agoWhat is a "line of code"? out of the 2b lines of code google has, how much of it was auto-generated? how many of those lines are config files? This is a very silly article that has little to no value.
- aMoniker 11y agoIs anyone else a little amazed that they have so many employees with access, yet their single codebase has not yet been leaked? It almost seems inevitable. They must have some contingency planning for such an event.
- breatheoften 11y agoAre the source of piper and the build tools also in the mono repo and also developed/deployed off the head branch? Seems like a random engineer could royally fubar things if they broke a service which the build system depends on ...
- QuercusMax 11y agoEverything has to go through pre-submit checks before it makes it to HEAD. And if you get it past those and it starts breaking stuff, there are robots that will automatically roll back your change if it breaks enough stuff.
- thrownaway2424 11y agoYou said "developed/deployed" as if it were the same thing. Even if you somehow checked in the giant flaw, bypassing all code review and automated testing, it's not like that would suddenly appear in production. Google isn't some PHP hack where you just copy a tarball to The Server. Binaries of even slightly important systems typically go through many stages of deployment, first into unimportant test systems, then usually very, very slowly into production with lots of instrumentation and of course, quick and easy methods of rolling back to the previous release.
- breatheoften 11y agoI see - it was something of a half baked thought but in my defense I wasn't trying to suggest that I thought the head was automatically deployed to production ... Deployed to testing round 1 ... N is still a "deployment" isn't it ...? The shared boilerplate for how that magic works in a scaleable way for so many different projects must be quite complex and itself hard to test ...
- rbanffy 11y agoWhat I find most distressing is that their Python code indents with two spaces... This is so wrong, Google.
- dblotsky 11y agoEven if the numbers are off, the assumption that 40M lines of code take less effort to write than 2B lines of code commits the fallacy that effort is proportional to number of lines of code. Come on, Wired, you can do better.
- ilurkedhere 11y agoYeah, but it's only like ~200 lines rewritten in Lisp.
- juhq 11y agoA serious question about Lisp and Google, is Lisp used within Google, and if so, in what projects and why?
- michaelwww 11y agoFor those interested, the source analyzer Steve Yegge was working on called GROK has been renamed Kythe. I don't know how useful it turned out to be for those 2B LOC. http://www.kythe.io/docs/kythe-overview.html http://www.kythe.io/docs/kythe-overview.html Steve Yegge, from Google, talks about the GROK Project - Large-Scale, Cross-Language source analysis. [2012] https://www.youtube.com/watch?v=KTJs-0EInW8 https://www.youtube.com/watch?v=KTJs-0EInW8
- deleted 11y ago[deleted]
- known 11y agoHow frequently Google does https://en.wikipedia.org/wiki/Code_refactoring https://en.wikipedia.org/wiki/Code_refactoring
- rosege 11y agoHow many lines is duckduckgo? :-)
- creshal 11y agoCan't be that many, given they outsource the actual search engine to third parties.
- wedesoft 11y agoWith 2 billion lines of code I would consider the problem of developers stepping on each other's toes essentially solved.
- izzydata 11y agoIf they were to recompile all of it on a standard desktop PC how long would it take? A week?