11 ms·
Google's monorepo, and it's not even close - primarily for the tooling: * Creating a mutable snapshot of the entire codebase takes a second or two. * Builds a
by pradn 2y ago
Google's monorepo, and it's not even close - primarily for the tooling:
* Creating a mutable snapshot of the entire codebase takes a second or two.
* Builds are perfectly reproducible, and happen on build clusters. Entire C++ servers with hundreds of thousands of lines of code can be built from scratch in a minute or two tops.
* The build config language is really simple and concise.
* Code search across the entire codebase is instant.
* File history loads in an instant.
* Line-by-line blame loads in a few seconds.
* Nearly all files in supported languages have instant symbol lookup.
* There's a consistent style enforced by a shared culture, auto-linters, and presubmits.
* Shortcuts for deep-linking to a file/version/line make sharing code easy-peasy.
* A ton of presubmit checks ensure uniform code/test quality.
* Code reviews are required, and so is pairing tests with code changes.
- pjungwir 2y ago> Entire C++ servers with hundreds of lines of code can be built from scratch in a minute or two tops. Hundreds, huh? Is this a typo? It makes me wonder if the whole comment is facetious. Or do C++ programmers just have very low expectations for build time?
- jbyers 2y agoI suspect they meant "hundreds of thousands"
- pradn 2y agoYes, oops - fixed!
- stefan_ 2y agoThat's the beauty of C++, an absurdly slow build is just an include away.
- thfuran 2y agoBoost always helps prop up falling compile times.
- robodan 2y agoThe public version of Google's build tool is Bazel (it's Blaze internally). It has some really impressive caching while maintaining correctness. The first build is slow, but subsequent builds are very fast. When you have a team working on similar code, everyone gets the benefit. As with all things Google, it's a pain to get up to speed on, but then very fast.
- citizen_friend 2y agoJust wait until you try a “modern” language
- dskloet 2y agoCode search not just across the entire code base but across all of time.
- ok_dad 2y agoI’m surprised they didn’t turn that into a product, it sounds great.
- twunde 2y agoParts have been. Sourcegraph is basically the code search post built by ex-Googlers originally. Bazel is the open source build tool. Sadly, most of these things require major work to set up yourself and manage, but there's an alternate present where Google built a true competitor to GitHub and integrated their tooling directly into it.
- snotrockets 2y agoBuilding tools for others is a competency that is under rewarded at Google. They would never.
- kridsdale3 2y agoI've published my Google proprietary stuff (when we decided to open source it) on GitLab, but they wouldn't let me do it on GitHub.
- tallowen 2y agoI always find these comments about interesting, having worked at Facebook and Google, I never quite felt this way about Google's Monorepo. Facebook had many of the features you listed and quite performantly if not more so. Compared with working at Facebook where there are no owners owners files and no readability requirements, I found abstraction boundries to be much cleaner at FB. At google, I found there was a ton of cruft in Google's monorepos that were too challenging / too much work for any one person to address.
- pradn 2y agoOWNERS files rarely get in the way - you can always send a code change to an OWNER. They are also good for finding points of contact quickly, for files where the history is in the far past and changes haven't been made recently. Readability really does help new engineers get up to speed on the style guide, and learn of common libraries they might not have known before. It can be annoying - hell, I'll have to get on the Go queue soon - but that's ok.
- deleted 2y ago[deleted]
- sawyna 2y agoThis isn't true at all for OWNERS files. If you try developing a small feature on google search, it will require plumbing data through at least four to five layers and there is a different set of OWNERS for each layer. You'll spend at least 3 days waiting for code reviews to go through for something as simple as adding a new field.
- okdood64 2y ago3 days for a new change on the biggest service on the planet? Not bad.
- tallowen 2y agoI agree that it could be worse! Facebook has significant (if not more) time spent and I found adding features to news feed a heck of a lot easier than adding features that interacted with google search. Generally a lot of this had to do with the number of people needed to be involved to ensure that the change was safe which always felt higher at Google.
- robertsdionne 2y ago* https://abseil.io/resources/swe-book/html/ch16.html https://abseil.io/resources/swe-book/html/ch16.html * https://abseil.io/resources/swe-book/html/ch17.html https://abseil.io/resources/swe-book/html/ch17.html
- WWWMMMWWW 2y agoGoogle's code, tooling and accompanying practices are developing a reputation for being largely useless outside Google ... and many are starting to suspect it's alleged value even inside Google is mostly cult dogma.
- wiseowise 2y ago> Google's code, tooling and accompanying practices are developing a reputation for being largely useless outside Google ... Not that I don’t believe you, but where do you see this?
- matthewfcarlson 2y agoI haven't worked at google, but this is something I have heard from a few people. Reputation is largely word of mouth, so it checks out for me. I suspect the skills/tools at most large companies are increasingly less transferrable as they continue to grow in scale and scope.
- deleted 2y ago[deleted]
- robertakarobin 2y agoI can vouch for it. It's the main reason I quit: none of the "hard" skills necessary to code at Google were transferrable anywhere outside of Google. It would have been easy enough to skate and use "soft" skills to move up the management ladder and cash big checks, but I wasn't interested in that. The reason it's not transferrable is that Google has its own version of EVERYTHING: version control, an IDE, build tools, JavaScript libraries, templating libraries, etc, etc. The only thing I can think of that we used that wasn't invented at Google was SCSS, and that was a very recent addition. Google didn't even use its own open-source libraries like Angular. None of the technologies were remotely usable outside Google. It might sound cool to use only in-house stuff, and I understand the arguments about licensing. But it meant that everything was poorly-documented, had bugs and missing features that lingered for years, and it was impossible to find a SME because whoever initially built a technology had moved on to other things and left a mess behind them. Some people may be able to deal with the excruciating slowness and scattered-ness, and may be OK with working on a teeny slice of the pie in the expectation that years later they'll get to own a bigger slice. But that ain't me so I noped out as soon as my shares vested.
- bbor 2y agoI’m just one incompetent dev, but I’ll throw this in the convo just to have my perspective represented: every individual part of the google code experience was awesome because everyone cared a ton about quality and efficiency, but the overall ecosystem created as a result of all these little pet projects was to a large extent unmanaged, making it difficult to operate effectively (or, in my case, basically at all). When you join, one of the go-to jokes in their little intro class these days is “TFW you’re told that the old tool is deprecated, but the new tool is still in beta”; everyone laughs along, but hopefully a few are thinking “uhhh wtf”. To end on as nice of a note as possible for the poor Googs: of all the things you bring up, the one I’d highlight the biggest difference on is Code Search. It’s just incredible having that level of deep semantic access to the huge repo, and people were way more comfortable saying “oh let’s take a look at that code” ad-hoc there than I think is typical. That was pretty awesome.
- aleksiy123 2y agoImho the reason for the deprecated and beta thing is because there is a constant forward momentum. Best practices, recommendations and tooling is constantly evolving and requires investment in uptake. I sometimes feel like everything is legacy the moment it's submitted and in a constant state of migration. This requires time and resources that can slow the pace of development for new features. The flips side is this actually makes the overall codebase less fractured. This consistency or common set of assumptions is what allows people to build tools and features that work horizontally across many teams/projects. This constant forward momentum to fight inconsistency is what allows google3 to scale and keep macro level development velocity to scale relative to complexity.
- bbor 2y agoThat’s all well said, thanks for sharing your perspective! Gives me some things to reflect on. I of course agree re:forward momentum, but I hope they’re able to regain some grace in that momentum with better organization going forward. I guess I was gesturing to people “passing the buck” on hard questions of team alignment and mutually exclusive decisions. Obviously I can’t cite specifics bc of secrecy and bad memory in equal amounts, so it’s very possible that I had a distorted view. I will say, one of the things that hit me the hardest when the layoffs finally hit was all the people who have given their professional lives to making some seriously incredible dev tools, only to be made to feel disposable and overpaid so the suits could look good to the shareholders for a quarter or two. Perhaps they have a master vision, but I’m afraid one of our best hopes for an ethical-ish megacorp—or at least vaguely pro social—is being run for short term gain :( However that turns out for society, hopefully it ends up releasing all those tools for us to enjoy! Mark my words, colab.google.com will be shockingly popular 5y from now, if they survive till then
- teaearlgraycold 2y agoMy experience with google3 was a bit different. I was shocked at how big things had gotten without collapsing, which is down to thousands of Googlers working to build world-class internal tooling. But you could see where the priorities were. Code Search was excellent - I'd rate it 10/10 if they asked. The build system always felt more like a necessary evil than anything else. In some parts of google3 you needed three separate declarations of all module dependencies. You could have Angular's runtime dependency injection graph, the Javascript ESM graph, and the Blaze graph which all need to be in sync. Now, the beautiful part was that this still worked. And The final Blaze level means you can have a Typescript codebase that depends on a Java module written in a completely unrelated part of google3, which itself depends on vendored C++ code somewhere else. Updating the vendored C++ code would cause all downstream code to rebuild and retest. But this is a multi billion dollar solution to problems that 99.99% of companies do not have. They are throwing thousands of smart people at a problem that almost everyone else has "solved" by default simply by being a smaller company. The one tooling I think every company could make use of but doesn't seem to have were all of the little hacks in the build system (maybe not technically part of Blaze?). You could require a developer who updates the file at /path/to/department/a/src/foo.java to simultaneously include a patch to /path/to/department/b/src/bar.java. Many files would have implicit dependency on each other outside of the build graph and a human is needed to review if extra changes are needed. And that's just one of a hundred little tricks project maintainers can employ. The quality of the code was uniformly at least "workable" (co-workers updating parts of the Android system would probably not agree with that - many critical system components were written by one person poorly who soon after quit).
- SR2Z 2y ago> But this is a multi billion dollar solution to problems that 99.99% of companies do not have. I know it's trendy for people to advocate for simple architectures, but the honest-to-god truth is that it's insane that builds work ANY OTHER WAY. One of the highest priorities companies should have is to reduce siloing, and I can barely think of a better way to guarantee silos than by having 300 slightly different build systems. There is a reason why Google can take a new grad SWE who barely knows how to code and turn them into a revenue machine. I've worked at several other places but none of them have had internal infrastructure as nice as the monorepo; it was the least amount of stress I've ever felt deploying huge changes. Another amazing thing that I don't see mentioned enough was how robust the automatic deployments with Boq/Annealing/Stubby were. The internal observability library would automatically capture RPC traces from both the client and server, and the canary controller would do a simple p-test on whether or not the new service had a higher error rate than the old one. If it did? The rollback CL would be automatically submitted and you'd get a ping. This might sound meh until I point out that EVEN CONFIG CHANGES were versioned and canaried.
- scubbo 2y agoInteresting to note that almost all of these are to do with tooling _around_ the codebase, not the contents _of_ the codebase!
- kridsdale3 2y agoJust like the man is the product of his genetic code, the codebase is invariably the product of the constraints on its edits enforced by tooling.
- nathan_douglas 2y agoso we beat on, commits against the tooling, borne back ceaselessly into the technical debt
- remram 2y agoIs that true, though? Is the code itself good? Because it is sorely absent from GP's list... If you are trying to say that people can't make bad code with good tools, I don't agree.
- scubbo 2y agoTo extend the previous commenter's simile - a man is a _product_ of his genetic code, but is also affected by environment. Bringing it back to the point at hand - yes, you are right that people can make bad code with good tools, but they'll be _much more likely_ to make good code with them (and vice versa).
- remram 2y agoThis Ask HN is not about "the code that should logically be best" but "the best code". There is no need for likelihood, people who have worked on it can report whether it is the case. And people here seem to praise the tooling exclusively... I would also point out that good tooling makes for good code, but big scale makes for bad legacy code. It is not at all obvious to me which of those effects should prevail at Google.
- dheera 2y agoWhy do so many people like monorepos? I tend to much prefer splitting out reusable packages into their own repos with their own packaging and unit tests and tagging to whatever version of that package. It makes it MUCH easier for someone to work on something with minimal overhead and be able to understand every line in the repo they are actually editing. It also allows reusable components to have their own maintainers, and allows for better delegation of a large team of engineers.
- Tyr42 2y agoI can change a dependency and my code at the same time and not need to wait for the change to get picked up and deployed separately. (If they are in the same binary. Still need cross binary changes to be made in order and be rollback safe and all that.)
- radicality 2y agoHave you ever worked at FB / Google / whatever other company has huge mono repo with great tooling? I went from many years at FB, to a place like you describe - hundreds of small repos, all versioned. It’s a nightmare to change anything. Endless git cloning and pulling and rebase. Endless issues since every repo ends up being configured slightly differently, and very hard to keep the repo metadata (think stuff like commit rules, merge rules, etc) up to date. It’s seriously much harder to be productive than with a well-oiled monorepo. With a monorepo, you wanna update some library code to slightly change its API? Great, put up a code change for it, and you’ll quickly see whether it’s compatible or not with the rest of the whole codebase, and you can then fix whatever build issues arise, and then be confident it works everywhere wherever it’s imported. It might sound fragile, but it really isn’t if the tooling is there.
- dheera 2y agoI have worked at a company that has huge monorepos and bad tooling. Tooling isn't the problem though, the problems are: - multiple monorepos copying code from each other, despite that code should be a library or installable python package or even deb package of its own - you will never understand the entire monorepo, so you will never understand what things you might break. with polyrepos different parts can be locked down to different versions of other parts. imagine if every machine learning model had a copy of the pytorch source in it instead of just specifying torch==2.1.0 in requirements.txt? - "dockerize the pile of mess and ship" which doesn't work well if your user wants to use it inside another container - any time you want to commit code, 50000 people have committed code in-between and you're already behind on 10 refactors. by the time you refactor so that your change works, 4000 more commits have happened - the monorepo takes 1 hour to compile, with nothing to compile and unit test only a part of it - ownership of different parts of the codebase is difficult to track; code reviews are a mess
- rkagerer 2y agoQuestion I've always wondered: Does Google's monorepo provide all its engineers access to ALL its code? If yes, given the sheer number of developers, why haven't we seen a leak of Google code in the past (disgruntled employee, accidental button, stolen laptop, etc)? Also how do they handle "Skunkworks" stlye top-secret projects that need to fly under the radar until product launch?
- rkagerer 2y agoEdit - I guess there hasn't been zero leaks: https://searchengineland.com/google-search-document-leak-ranking-442617 https://searchengineland.com/google-search-document-leak-ran...
- pradn 2y agoThe very very important stuff is hidden, and the only two examples anyone ever gives are core search ranking algorithms and the self-driving car. Even the battle-tested hyper-optimized, debugged-over-15-years implementation of Paxos is accessible. Though I’m sure folks could point out other valuable files/directories.
- kccqzy 2y agoFormer employee here. I remember a third example: the anti-DoS code is hidden. I remember this because I needed to do some very complicated custom anti-DoS configuration and as was my standard practice, I looked into how the configuration was being applied. I was denied access. Fourth example: portions of the code responsible for extracting signals from employees' computers to detect suspicious activity and intrusion. I suspect it's because if an employee wants to do something nefarious they couldn't just read the code to figure out how to evade detection. I only knew about this example because that hidden code made RPC calls to a service I owned; I changed certain aspect of my service and it broke them. Of course they fixed it on their own; I only got a post-submit breakage notification.
- robodan 2y agoPartial check outs are standard because the entire code base is enormous. People only check out the parts they might be changing and the rest magically appears during the build as needed. There are sections of the code that are High Intellectual Property. Stuff that deals with spam fighting, for example. I once worked on tooling to help make that code less likely to be accidentally exposed. Disclaimer: I used to work there, but that was a while back. They probably changed everything a few times since. The need to protect certain code will never go way, however.
- Ocerge 2y agoI recently left Google and knew it was going to be a step down from Google's build ecosystem, but I wasn't prepared for how far a step down it would be. It's the only thing I miss about the place, it' so awesome.
- pmb 2y agoPeople who have never had it have no concept of how much they are missing. It's so frustrating.
- Thaxll 2y agoMost of those arguments are not about code quality though.