4 ms·
> In May 2015, an on-call engineer failed ... At the time, Uber had recently reached a valuation of $50 billion. I find that astonishing. Outside Silicon Vall
by thisrod 9y ago
> In May 2015, an on-call engineer failed ... At the time, Uber had recently reached a valuation of $50 billion.
I find that astonishing. Outside Silicon Valley, I think that it would be quite unusual to leave $50 billion of plant running overnight, but pinch pennies by not paying anyone to stay up and check the oil levels during the dog watches.
- potlee 9y agoThere are many different people on-call for many different things at any given time. This is standard industry practice.
- WalterBright 9y agoFor mission critical systems, there should be a backup plan should the on-call engineer not respond. Like an airliner has a copilot.
- cbanek 9y agoThe problem is, unlike a copilot who is trained on the system, usually on-call people are winging it at best, and after the horse has left the barn. I've been on many on call rotations for many different projects. Some of them have had such great ideas as "you'll learn a lot from being on call, so we'll add you to the rotation (on your first day)." Sadly, people don't normally document code and emergency procedures nearly as much as they should. Unlike an airliner, there's no emergency manual sitting around. Even a normal manual would be hard to find, and who knows if it's current. One place I worked had solved this problem by having an actual operations team, a practice that seems to have faded with time. Our operations team was great. They were there in 3 shifts, 24/7, to cover your ass. They knew more about how the service ran than any dev. There were standards for documenting procedures, and if you got the call (there wasn't an on call rotation, you were only called if everyone was clueless and you wrote it), it's because things were really bad. This was before "devops" was minted the silver bullet. While a dedicated operations team can save you, and your reputation, most organizations would rather not pay that, and push that cost onto the devs, and make them "feel the pain" of their mistakes.
- kevan 9y agoI can understand not doing butts in chairs 24/7, but I'm really surprised they managed to make it to that scale without adding paging escalation. If someone doesn't respond after X minutes then page their manager. Repeat until you get to the CEO.
- oh_sigh 9y agoYes....that's the way it is set up at Amazon, and I recall one time a page from my little teams project with no revenue or customer implications for downtime managed to make it all the way to one level below Bezos before it was acked (this was 1am after Thanksgiving, 2012).
- jfoutz 9y agoThat was my initial thought as well. But, low disk space on the master DB? And you need a human to go futz around with it? There is either a whole lot of missing monitoring and automation, or those are completely worthless pages. I understand the freakout, i assume it could take down the site. But, like, not responding to a page is like the very last on the list of things to fix. Calling in random engineer is pretty much last ditch, hail mary, the world is on fire, oh fuck we're going to go out of business disaster. At a certain point, you kind of have to assume one possibility is your oncall person is dead. Whatcha gonna do then? I seriously doubt all of their careful high availability planning failed. I would bet their paging is (hopefully was) just stupid.
- oh_sigh 9y agoFrom the article, it's not that the engineer never responded. The engineer acknowledged the pages but never did anything to correct the problem beyond that. I agree with you though that a random engineer might not have the tools and knowledge ready to fix something like that, but that is why you can let the pages escalate until you wake someone up who does.
- oh_sigh 9y agoFrom the article, it is stated that the oncall engineer acked the pages for low disk space repeatedly but never took any corrective actions. So, escalation wasn't exercised in this case because the engineer tricked the system into thinking they handled it.
- ncallaway 9y ago> “We try to have a blameless postmortem,” said one engineer. “That email from Thuan [...] was great example of not following that.” Besides, the employee added, “if you’ve been woken up at 3 a.m. for the last five days, and you’re only sleeping three to four hours a day, and you make a mistake, how much at fault are you, really?” > In a follow-up email two days later, Pham addressed criticism of his decision to email the staff about the engineer’s mistake. "It came to my attention that some people in the org think that my note to the company is overly critical and has the feel of throwing people under the bus,” he wrote. > “I don't want to create a culture [where] people are fearful of making mistakes or causing outages because they want to move fast and take smart risks, but I also don't want a culture where we do substandard work and cause outages that are easily avoidable,” Pham wrote. > He continued: “Feeling defensive, or feeling like a victim, is NOT the way to get ourselves better.” I find Pham's response to criticism about the blameless post-mortem to be lacking in self-reflection. The entire point point of a blameless post mortem is to ensure that people don't feel defensive or feel like a victim. To get criticism about failing to follow a blameless post-mortem process and respond by saying: "don't feel like a victim?" tells me that they haven't internalized the criticism in any way.
- oh_sigh 9y agoBlameless post mortems only work when you know that engineers or whoever is on call are operating in good faith. From Pham's first email, it seems like there is some question whether that was the case, because the engineer repeatedly acked the pages without resolving any issue: >The on-call engineer received/acknowledged three alerts about master database being low on disk space, but ignored it. This is not acceptable,” Pham wrote in the email, sent to more than 3,500 employees and obtained by BuzzFeed News. “We are looking to determine whether this is negligence or whether a different on-call engineer could have reasonably missed the alerts amidst a flood of other alerts from the systems at that time.” I think the timing of the statement was bad - it should have been a result of the post mortem if it was negligence, and not so explicitly stated publicly going in, if only because it can tarnish someone's reputation even if it turns out not true.
- wdb 9y ago