4 ms·
So I've been oncall at two major companies (Google and Facebook) and, at least in my experience, this covered both ends of the spectrum. Basically, Google gets
by cletus 4y ago
So I've been oncall at two major companies (Google and Facebook) and, at least in my experience, this covered both ends of the spectrum. Basically, Google gets it mostly right and Facebook gets it mostly wrong.
At Google, a new service has to be supported by the team that developed it. There'a an extensive launch checklist that includes monitoring, having a runbook, etc. Here's the most important part: you're paid when you're oncall. The amount varies depending on how important the service is and the expected response time but can easily be 5 figures a year. Oncall period varies but a week at a time varies with hopefully 8-12 people in rotation.
Too few people and people get burnt out. Even if nothing happens on an oncall shift, it's an annoyance and a restriction on what you can do. Too many and people tend to forget what to do. So with a sufficiently large team you may end up with some people in the rotation and some people not. That's why the compensation is importatnt.
Particularly large, important and mature services may enjoy SRE support. You can't throw a service over the fence and have SRE deal with it. It doesn't work that way. It typically needs to have been running for at least 6 months and SRE needs to be satisified it's sufficiently reliable, stable and monitored with a good runbook. SRE support is globally distributed and typically means 8 hour shifts during normal hours.
The owning team will often still be secondary support.
Also code has to be owned by somebody. This may be a team but when I was there (some years ago now so it may have changed) this also meant 2 actual people (not just team aliases) had to be owners. This is to avoid abandonware. This very much is a support and oncall issue.
Facebook OTOH is a dumpster fire when it comes to oncall.
Not getting paid to be oncall is (IMHO) one of the biggest mistakes. The mantra is "it's part of the job" but that responsibility is not shared equally. That's the point of compensation.
My experience at Google was that issues were relatively infrequent. What I saw at FB however was that oncall could often be the only thing you did for the week. Noisy alerts, alerts caused by issues in downstream systems that you could do nothing about or would get ignored by their oncall, a bunch of issues raised that some would just ignore until they expired (or closed just prior to going out of SLA as "could not reproduce"), etc. You may also be dealing with code that nobody owns (or, rather, nobody takes responsibility for) for features that are live.
Plus the incentive structure, at least on the product side, was to ship new features. Oncall was often treated as just extra work you have to do on top of whatever else you're doing.
Obviously I didn't see how every team did it so none of this is absolute but I did see a reasonably high number of samples.
It's also worth noting that not everything at FB is like this (eg the Web Foundation people were and I believe still are outstanding). Also, in high-visibility outage situations you have highly knowledge individuals who can and do get involved and know the right people to push.
The FB equivalent of SREs is Production Engineers ("PEs"). There are less of these and more services at FB are supported by the SWEs than at Google (IME).
I got the impression that FB processes and culture were forged when the company had less than 500 employees and they never really adjusted to the greater scale. There are a lot of things that work very well. Oncall just isn't one of them. Nor is code ownership.
- igotsideas 4y agoI had no idea you essentially get “overtime” pay for on call at google. That’s how it should be imo. I’ve avoided on call jobs for the lack of extra pay for doing more work.
- newman123 4y ago> alerts caused by issues in downstream systems that you could do nothing about How does Google deal with issues caused by downstream systems causing alerts?
- okdood64 4y agoGenerally 1) Mitigate if possible from your service's end, while simultaneously 2) paging the dependent service's team to mitigate/resolve; and if this happens too frequently or if it was a particularly bad incident you can push the other team to 3) create a postmortem with follow up AIs if they already didn't do so.
- cletus 4y agoThe deeper you go into the stack, the more reliable things tend to get, the more mature those systems tend to be and the more likely they are supported by SRE who take things very seriously. So if you're having an issue with Spanner, first it's likely not a bug in spanner. If it's an outage, somebody has probably already been paged. But if not, paging someone responsible will be answered quickly and treated seriously. You could've unexpectedly gone over quota on something. More often than not you can alleviate that with temporary quota while you resolve your issue (by reducing your usage, getting more permanent quota or both). A big part of this is that it's a cultural thing.
- dekhn 4y agoNote that the on-call bonus is not entirely known. I had managers try to put my team "on-call" for a product and after I explained to them how Google actually did it, they suddenly said "oh, it's not really on-call. You just have to be ready to answer the pager at any time and respond". I was also on what was one of the most dysfunctional on-calls at the company- keeping several distributed clusters of unique business-critical mysql instances that failed frequently with an unreliable failover method. For some reason it was a joint on-call with a neighboring team so I was responsible for systems I didn't know about or understand but were busisness critical. At times it was fun, at times it was educational, but at times, it was the worst thing in the world for me.