6 ms·
it's interesting how Amazon is now positioning themselves as a tier 1 employer. They used to be a tier2. AFAIK they also have a brutal pager culture where you
by top256 9y ago
it's interesting how Amazon is now positioning themselves as a tier 1 employer.
They used to be a tier2.
AFAIK they also have a brutal pager culture where you're awoken in the middle of the night.
- babaganoosh89 9y agoIsn't that the point of a pager system?
- jorblumesea 9y agoMany companies have overseas teams that take the night shift, eg: every on call dev has 12 on during the day, 12 off at night. Amazon does not based on cultural reasons (allegedly) because it promotes ownership and accountability for those teams. Idea being if you're woken up at 4am because of a flaky test or some issue you'll be more inclined to fix it. And if it's a recurring issue, you'll really want to get it sorted or else your life will be hell. It also pushes emphasis on serious testing methodology and the idea that Amazon is a 24/7 business. It actually makes a fair amount of sense, but it's also brutal and contributes to burnout and attrition. I also think this is starting to change for some teams.
- openasocket 9y agoAmazon encourages developers to be on call, but there are also rolling on-call systems so no one gets paged in the middle of the night. A team I work closely with does this. It isn't very stressful for me, I'm just on call for small, well-defined intervals, something like 1 week out of every 6, and only high severity issues will cause a page. In the last year or so I've been on call I was only paged after normal business hours once. Of course, Amazon is an absolutely massive company, so different teams may have wildly different experiences.
- jorblumesea 9y agoI think it's different for each team, especially the AWS groups. I know some support engineers that do 12on/12off rolling calls like you said, but others where it's one person on call, 24/7 for a week and they catch everything, not just SEV 1. It really just depends on your org/boss but the AWS org is known to be particularly pager-slaved.
- wink 9y agoMeh, I really don't buy that reasoning - but maybe it's my current job where 9/10 "got woken up at night" is stuff that's not really our fault (e.g. something in the datacenter is broken, or missing network connectivity, etc.pp. - and yes, I've heard of multihoming, but it's not the team's budget and decision to not keep stuff redundant...)
- twayamznacct 9y agoDisclaimer: Ridiculous pager duty was one of the reasons I quit Amazon (combined with massive failures in management, which I'll describe shortly). The team I was on (retail-related, not AWS), shifted away from the rolling schedule with your counterparts in India taking over for the other 12 hours in order to push the "promote ownership" BS. The only problem? Management was constantly pushing for new things to get done on extremely tight deadlines (including "emergency features" that needed to be done and in prod in days when they likely required a week of design efforts to get right, never mind actual dev time) so you have two options: (1) Develop something stable, with good test coverage and the like, and work to fix it if it breaks... and work 20 hours to get it done. (2) Shit something out as fast as you can and hope it breaks when someone else is on call (who likely will be too busy triaging SEV3-5's during business hours to even think about spending time fixing the root cause of the SEV1/2, rather than mitigating it and moving on...or, even worse (!) (/s) try to get the actual feature owners to fix it) in order to maintain some semblance of work-life balance. I'll leave it as an exercise to the reader to guess at which route was generally taken.
- snorkel 9y agoI've experienced that when the coders are getting paged they'll focus on quality and writing better proactive tests. But when you throw the ops responsilities over to another team then that other team gets to decide how and when new code gets shipped which slows down the whole operation. I've experienced that Dev vs. Ops arrangement too where the ops team was incentivized to maintain stability over allowing new code to be shipped which resulted in no code being shipped! When devs own the ops responsibility too then things move faster.
- noahl 9y agoDo you think the tier 1 thing started after the NYTimes article on them, or before?
- top256 9y agoI think it started with AWS success
- falcolas 9y ago> pager culture where you're awoken in the middle of the night For multinational corporations, this has never made sense to me. Why do we insist on relying groggy, "at the bottom of their performance curve" engineers to keep your system up, instead of someone in the middle of their work day? IMO, if you're waking someone up in the middle of the night, something is horribly wrong, and it's not the issue you woke that person up for.
- dsfyu404ed 9y agoMost don't page people at 1am for the initial page. People only get paged at 1am when the problem has been narrowed down to some specific software and all the people with expertise in that thing are in the same time zone.
- mabbo 9y agoHaving done Amazon on-call for five years, I can give some background on that. First, the front-line pager duty is usually hit by an ops eng team first. Each of those teams has an India group and an North America group. If they can't resolve the issue, they page in the developers. Having good Ops Eng guys supporting you is a blessing. There's also a very strong incentive for teams to write good software and test before deploying when you know damn well that if you're deploying something that doesn't work, you're the guy who will have to deal with it at 3am. Having strong integration and system tests suddenly becomes incredibly important. Edit: and as dsfyu404ed said, the Ops Eng guys will be smart enough to narrow down where the problem is usually. If you're being paged, it's usually your team's fault.
- colmmacc 9y agoAmazon on-call engineer here! I'm on-call for three rotations. Two are "Call leader" rotations where a call leader like me is engaged if there's any kind of large event (e.g. Amazon.com orders aren't working, or an AWS service is having issues). I volunteer for those because it's super rewarding to be able to help materially make a tangible difference now and then, wouldn't trade them away for anything. For those rotations, I am paged in the middle of the night, and there's nothing I can really do about that. The other rotation is for one of my dev-team's software, and I'm hardly ever paged - like once or twice a year max. That's how it's supposed to be: by having skin in the game and being directly on the hook for the reliability of what we produce, I'm sure that we automate to a much greater degree than we might otherwise. We have a culture of strong ownership: it's not uncommon to see senior engineers elect to be engaged on all of their team's live issues, even if they aren't on-call themselves. Some other teams make different trade-offs, and have follow-the-sun rotations or engagements that are more reliant on humans following run-books than automations. There can be a place for that, but it gets old quickly and is not scalable. Good teams will prioritize the work to get out of that above almost anything.
- hsod 9y ago> The other rotation is for one of my dev-team's software, and I'm hardly ever paged - like once or twice a year max. I'm in an on-call rotation (not at AMZ) where I hardly ever get paged, and it's cold comfort. I still have to carry my laptop with me everywhere I go, make sure I have access to the Internet, make sure I don't have one too many at the bar on Saturday night, etc.
- deleted 9y ago[deleted]
- synicalx 9y ago> AFAIK they also have a brutal pager culture where you're awoken in the middle of the night. That's a fairly normal part of Operations work and has been for decades. Most places (I assume Amazon is one of them) will have a pool of lower-tier admins on deck 24/7 who handle and investigate 99% of the issues that crop up. But for that 1% they can't deal with, then they need to be able to reach someone who CAN deal with it. It's much cheaper to occasionally call an SME every once in a while when they're needed, than it is to hire another 3 or 4 of them to cover a 24 hour roster. I've been called out in the middle of the night to fix someone else's mess more times than I can count, so I'm 100% behind Amazon's cultural reasons for sticking to this sort of on-call system as well.