4 ms·
This is great. Having the engineers who built the thing be the same people who respond to pages is a great way to incentivize robust systems. If you built it, y
by pkaeding 5y ago
This is great. Having the engineers who built the thing be the same people who respond to pages is a great way to incentivize robust systems. If you built it, you likely know how to fix it better than anyone else, and if you don't want to be disturbed in the evening, you will think about how to better deal with faults during the early stages of development.
- willcipriano 5y agoIn a world where people leave every 1 - 2 years to stay current with the market, more often than not the failure will be related to what someone else did a long time ago.
- deleted 5y ago[deleted]
- zero_shift 5y agoI'm sorry, I'm going to have to chew you out here. Firstly: when was the last time your on call rotation comprised exclusively programs that you had written? How many systems does your team look after that it "inherited" from elsewhere? Secondly, how many organisations allow developers to prioritise solving pages over other forms of revenue generating work? How many teams demand negotiation with a "product owner" before a piece of work can be added to the board? Thirdly, how many pages are truly easy or tractable to fix? I have worked in teams where we were constantly paged due to external API failures, but intra team disagreements meant we could neither punt the responsibility elsewhere nor resolve the problem. The problem wasn't technical, it was social, but there's no PagerDuty for absentee Engineering Managers. Fourthly, how often is the Real Problem™ for the reliability and uptime of a given system, actually at the program level, and not at architecture or system level? How many programs are hamstrung by internals alone? Once you get past the basics, most of the _really_ significant decisions affecting reliability are at a system design level that Johnny or Jane Developer isn't going to be empowered to fix in a Scrum "sprint". Really, this "shitty on call incentivises robust systems" argument is facile. It's paper thin. Put it out in the sunlight for a second and it crumbles. It only makes any sense, tentatively, under idealised conditions where developers alone are responsible for non functional requirements and the usual relationship of employer / employee is suspended. Think about it for a second. It's just rot.
- deleted 5y ago[deleted]
- aero142 5y agoI've been part of an on call rotation at my current job and previous job. Both had the same rules because I and other engineers insisted on it. * You are only on-call for systems you can directly change. * The person who is on call has full flexibility to work on anything that improves the life of the on-call person for that time. * Anything that wakes people up in the middle of the night gets prioritized above new feature work. * There is a pager set of monitors, and a message only set of monitors. If a page goes off, and there is nothing the on-call engineer can do about it, it gets moved to the message only channel or removed, because it is a bad monitor. I discussed this list of rules when I interviewed and the job description included being on-call. It wasn't a negotiation. If those aren't followed, I remove the monitors. I'm sorry you had a terrible work environment, but I encourage everyone to have professional standards. I hope you find a place where you can.
- ianpenney 5y ago> Firstly: when was the last time your on call rotation comprised exclusively programs that you had written? How many systems does your team look after that it "inherited" from elsewhere? Agree it's not fair to be on-call for things you can't fix. That's silly. But, you're surely not saying when you push a change to a 24/7 mission critical system, and you're on the "git blame" - you're not interested in taking responsibility if the consequences occur after 5pm? And now an SRE has to learn everything you know, in order to fix it, instead of asking for your help? Also - think about it a different way, you personally, are not the one I'd ask to participate in on-call. It's your whole team, maybe even your whole department. You and your team/department should work out how to field that together because you have more context into each others work/life balance, projects, skills, risky commits, etc. Your team would manage your own rotation. > Secondly, how many organizations allow developers to prioritise solving pages over other forms of revenue generating work? How many teams demand negotiation with a "product owner" before a piece of work can be added to the board? Tell me about it. Totally agree. This is where I focus much of my energy as a part of leadership teams - making other leadership realize they're over-promising on feature delivery without concern for their poor operations. Frankly, fuck sales-driven development for this. I even work with marketing departments to value-stream map so we include good ops & security as part of our core prop to customers and investors. If they don't get it? Well, I guess the fish is already rotting from the head, as they say. > Thirdly, how many pages are truly easy or tractable to fix? If they're repeated ad nausea? Then those alerts need to go. Now. And if they can't go, because Management don't see the pain they're causing? See my response to #2. Culture change first. > Fourthly, how often is the Real Problem™ for the reliability and uptime of a given system, actually at the program level, and not at architecture or system level? Depends on your industry/platform. Good SRE sysadmin troubleshooter types will do everything they can before they push the escalation button to engineer tier. At least they'll get all the data and metrics they can and document the situation at hand before the engineer comes online. Sometimes you don't really know what caused a problem until you all get together and retro it. Better to stay blameless even beyond when you have consensus on the technical factors. One time, it was a leap-second messing up some 3rd party vendor software, with consequences for our code base that nobody ever expected. Another time, it was a tired SRE combined with a lack of proper procedure combined with a code mistake with no test coverage. Another time, it was a binder that fell on a customer's space bar effectively DoS'ing a system that had no rate limiting. Some of those got fixed pretty quickly by an SRE. BUT - whatever the cause... the times that our MTTR got truly trashed? (We're talking DAYS...) It was because the person who wrote the code and had all the context, had already quit long ago - leaving us with no proper logging and no one to help us. The organization had never prioritized replacing their code or hiring someone who could understand it. > Really, this "shitty on call incentivises robust systems" argument is facile. You're right in the sense that, considering all of the above, I would need your buy-in. To get it, I'd seek to eliminate the "rot". I'm truly sorry you got burned in the past. I have also been there. When it was too much for me, I quit, and found somewhere more ideal. Now I just keep trying to work with people I trust as they move around in the industry, to avoid all the org B.S. you've rightly brought up.
- sdevonoes 5y agoSorry, but not. I keep my 9-5 strict. I'm not into the game of getting more money in exchange for my scarce free time. If the company needs people to work on Sunday mornings, they can hire them (i.e., SREs). I can't understand why it is becoming so normal for regular (senior) developers to become slaves of our companies by working more than 40h/week; the general excuse is "you build it, you run it, you fix it". If we are into writing robust software, I'm all in, but that's a totally different thing.
- knicholes 5y agoThat sounds amazing until you get a system that is complicated and the part of the system you work on relies on an unstable dependency upon which you have no control. Oh no, my service isn't responding to 99.9% of requests within 250ms!? Boom, page at 2am. The service that is responsible for the alert either doesn't have monitoring set up correctly or at all or they aggregate their metrics differently, so on average, all of their calls are 5ms, but for all of your calls, maybe they're taking 2-3 seconds. It's a nightmare that I escaped after three years. I was driven from a happy person to someone who hated his life. It took me a while to realize it was my job. The worst part is my management told us that the on-call pay was already "baked into your salary." I switched jobs internally. No more on-call. Strangely, I was able to keep my pay without having to wake up in the middle of the night any more. Oh yeah, and you aren't the only one who wakes up. Someone sleeping in your bed may be an insomniac and may have just fallen asleep. Your alert wakes them. Now they don't get to sleep for the rest of the night. Now you and they are both sleep deprived and irritable. It obliterates personal relationships. Maybe your kids overhear you explaining to someone what is going on in the next room. Whenever anyone asks you about your job, PTSD triggers and you spend ten minutes venting. I am so, so, so glad I'm no longer on call. I don't know what they'd have to pay me to get back on it, but it's at least 2x, and even then, I couldn't last for over a year. You don't get to go to movies. You don't get to join your friends when they go biking or hiking or camping. You get interrupted in the middle of parent-teacher conference. You have a laptop next to you at the 4th of July party. You get interrupted in the middle of a shower. You don't get to host Thanksgiving or Christmas because you may have to work in the middle of making dinner. You're about the roll the dice to take the first turn of a game you've been promising your kids you'd play with them in the first attempt in a long time to be a decent parent, and your alarm goes off, you work for the next three hours, and they play without you. There is no 40-hour work week for software engineers. There is no union. You work your normal days then you also get to work normal nights. Then you work a normal day again either fixing the problem that caused the alert or spend 6 hours in post-mortems explaining what happened when instead you should be sleeping.