3 ms·
50 SRE's sounds like they could keep borg running. There really isn't that much details on what borg is made up of outside of a 2015 research paper[1]. Just to
by ZeroSolstice 3y ago
50 SRE's sounds like they could keep borg running. There really isn't that much details on what borg is made up of outside of a 2015 research paper[1].
Just to address the parts you laid out.
1. Keep Borg running
What would entail running? There are plenty of small teams that manage hundreds of servers and virtual machines through monitoring, deployment, networking, decom, etc lifecycle. This seems counter to purpose of cloud computing, commodity hardware and abstracted compute resources that can be deployed on-demand. Not to mention redundant sites, availability zones, <insert your cloud providers name for High Availability(HA) feature here>. Is this the part that SRE's maintain?
2. Machine maintenance
From documentaries[2] and other comparable data center operations this would seem to be handled by on-site datacenter staff. Disk replacement, physical replacement, network cabling, etc. Without any other operational info I would doubt resources would just be dedicated to borg as its a bit of a overall technician site work. As for OS and software updates and again without any other info or insight would seem to fairly automated after passing testing or at least passed to another team that just handles updates.
3. Machine procurement
Again commodity hardware or if its all custom still that would seem to be more of a EE task for build out and accounting for purchasing. Otherwise at Google's scale you would just get pre-populated racks delivered and replace the entire rack when 51% of the machines have failed. This doesn't really seem SRE or developer specific. It would also be happening for other services if they are already aren't abstracted from the underlying hardware/network layer.
4. Regulatory changes
I'm not sure what this would be in relation to a job processing system? Is this checking where you are saving data? Seems like a feature that is built once.
I agree that if you have follow-the-sun and need to be up with 5 (9's) 50 people in one time zone would be hard to ensure that but once the foundation is setup its just (50) people in a different location or timezone there isn't really anything too different about what they are doing.
If you have additional information or insight that would interesting to hear about.
[1] https://research.google/pubs/pub43438/ https://research.google/pubs/pub43438/
[2] https://www.youtube.com/watch?v=XZmGGAbHqa0 https://www.youtube.com/watch?v=XZmGGAbHqa0
- kikimora 3y agoI think you oversimplify a lot and miss many real world things that a Google sized org has to deal with. Machine maintenance - someone has to coordinate updates across multiple data centers run by different organizations. Imagine you want to update networking from 1Gbs to 10Gbs across 3 DS in 3 different countries and time zones. After a while you’ll wish to own data centers, but that means multiple legal entities, bank accounts, etc. in different countries. Machine procurement - one DC can buy Intels XYZ while another cannot because there are no local vendors who can deliver. So now you have a dilemma - build software working on different hardware, complicate all future updates or go and negotiate with Intel and other vendors. Regulatory changes is when Malaysian government decides that it wants to protect children from something and you have to take extra measures before you can show search results or publish ads. Or EU votes for something called GDPR. Or Russia suddenly decides that Ads must be VAT taxed. And don’t forget about a Microsoft exec who wants to discuss how they can use Google Search in their browser. Honestly I don’t see how it can be done with 50 engineers. You’ll need hundreds of people, mostly managers and clerks to keep this thing running.
- ZeroSolstice 3y agoI appreciate the examples but none of these items are difficult regardless of size of the organization. These are standard system and network administrative items or on-site technician work. Unless you have other data about the borg architecture to show, its entire premise from the paper[1] is that it just runs jobs and removes the developers needing to know about the underlying hardware or OS that its running on it. From the paper I referenced its huge cluster and if its their own software I'm sure with the development experience at google they have accounted for loosing a cluster node. Losing any number of cluster nodes is the same scenario if you were performing an software or hardware upgrade in which case you would be taking a node offline. I would be very skeptical that there are people manually kicking off upgrade jobs for nodes in the cluster. Maybe an A/B deployment or burn-in for a week before fully pushing it out with a pipeline workflow. I'd more so say that their borg implementation probably runs at different minor versions frequently. Google runs their own data centers, thats the video documentary I referenced[2]. For machine procurement I don't see that as a SRE task. If they are making custom boards and systems they are going to have an entire group supporting that. At the size of google I'm thinking they wouldn't be running into supplier issues that they could figure out easily or account for in their borg software. Again its supposed to be one big abstraction. For the regulatory changes for services like SEO/advertising, GMAIL, search, etc they all have separate teams for those services and would seem separate from SRE's that are maintaining borg the service which is provided to those groups. [1] https://research.google/pubs/pub43438/ https://research.google/pubs/pub43438/ (research paper on borg) [2] https://www.youtube.com/watch?v=XZmGGAbHqa0 https://www.youtube.com/watch?v=XZmGGAbHqa0 (documentary at their data centers)