4 ms·
One of the darker reasons I've been pulled into an operational call about order drops in a region's marketplace...
by matt_daemon 4y ago
One of the darker reasons I've been pulled into an operational call about order drops in a region's marketplace...
- 2143 4y agoDo you have some numbers (of course, only if you're willing and allowed to share them)? Curious to know what's the usual on a normal day, and how much it dipped.
- zzleeper 4y agoSorry for the ignorance but what's an order drop?
- phnofive 4y agoA dip in the rate of orders being placed, often signifying some part of the checkout flow having slowed or broken.
- refulgentis 4y agoSo OnCall is mandatory for Amazon engineers and you will literally get paged just because some stat drops, not that it goes to 0? And there's no human in the loop for a basic sanity check? Sounds awful
- djbusby 4y agoThe pager is bringing humans in for the sanity check.
- mjr00 4y ago"Some stat" here translates to millions, maybe tens of millions of dollars every minute at Amazon scale. If you were running a business and suddenly started losing a million dollars a minute, wouldn't you like to have someone investigate why that's happening?
- ojbyrne 4y agoI think the human in the loop is the on call engineer.
- ncallaway 4y ago> And there's no human in the loop for a basic sanity check? I mean, I'm not going to go to bat for Amazon's worker practices, but... how would you get a human in the loop for a basic sanity check without someone getting paged? Like, you have to alert someone so that they can do a sanity check.
- majormajor 4y agoYou'd have a separate 24/7 ops teams instead of every single team being first-responders for their own things. This team would be watching the dashboards and alarms with some first-level investigation/triage playbooks so that if all that's needed is "press this button to reboot the server" or "it's a holiday, or the national team just won the world cup, this is a drop in traffic cause everyone is outside" then nobody else has to get paged. I don't know which model Amazon uses myself, but it sounds more like the "everyone is their own first responder" model which surprises me a little, but each model has its pros and cons. The justification I'm most familiar with is "if there's a separate ops team you're going to be complacent and throw shit over the wall without caring about quality" which is a somewhat lazy excuse for not actually getting your engineers to care about quality with anything but a hammer to beat them on the head with, but ....
- bee_rider 4y agoMaybe the person who made the original comment is on that 24/7 ops team?
- mjr00 4y agoThere are both models at Amazon. Individual teams have their own on-call; if for example an RDS instance is continually failing to boot, someone on-call (edit: from the RDS team) will get paged to look at it. I believe order drops are looked at by dedicated reliability team, because it's such a significant problem. But they don't watch graphs because... why would they? They set up the machines to page them when metrics aren't within an expected range. Which is what happened.
- 4y ago
- lambda_lord 4y agoParent poster most likely works at Amazon. When the rate of purchases falls below a threshold, many engineers across the company get paged to find the issue. As you can imagine a news event would decrease activity on Amazon for a while.
- deleted 4y ago[deleted]
- bamboozled 4y agoHope you’re holding up?