4 ms·
I'm sorry, I'm going to have to chew you out here. Firstly: when was the last time your on call rotation comprised exclusively programs that you had written? H
by zero_shift 5y ago
I'm sorry, I'm going to have to chew you out here.
Firstly: when was the last time your on call rotation comprised exclusively programs that you had written? How many systems does your team look after that it "inherited" from elsewhere?
Secondly, how many organisations allow developers to prioritise solving pages over other forms of revenue generating work? How many teams demand negotiation with a "product owner" before a piece of work can be added to the board?
Thirdly, how many pages are truly easy or tractable to fix? I have worked in teams where we were constantly paged due to external API failures, but intra team disagreements meant we could neither punt the responsibility elsewhere nor resolve the problem. The problem wasn't technical, it was social, but there's no PagerDuty for absentee Engineering Managers.
Fourthly, how often is the Real Problem™ for the reliability and uptime of a given system, actually at the program level, and not at architecture or system level? How many programs are hamstrung by internals alone? Once you get past the basics, most of the _really_ significant decisions affecting reliability are at a system design level that Johnny or Jane Developer isn't going to be empowered to fix in a Scrum "sprint".
Really, this "shitty on call incentivises robust systems" argument is facile. It's paper thin. Put it out in the sunlight for a second and it crumbles. It only makes any sense, tentatively, under idealised conditions where developers alone are responsible for non functional requirements and the usual relationship of employer / employee is suspended.
Think about it for a second. It's just rot.
- deleted 5y ago[deleted]
- aero142 5y agoI've been part of an on call rotation at my current job and previous job. Both had the same rules because I and other engineers insisted on it. * You are only on-call for systems you can directly change. * The person who is on call has full flexibility to work on anything that improves the life of the on-call person for that time. * Anything that wakes people up in the middle of the night gets prioritized above new feature work. * There is a pager set of monitors, and a message only set of monitors. If a page goes off, and there is nothing the on-call engineer can do about it, it gets moved to the message only channel or removed, because it is a bad monitor. I discussed this list of rules when I interviewed and the job description included being on-call. It wasn't a negotiation. If those aren't followed, I remove the monitors. I'm sorry you had a terrible work environment, but I encourage everyone to have professional standards. I hope you find a place where you can.
- ianpenney 5y ago> Firstly: when was the last time your on call rotation comprised exclusively programs that you had written? How many systems does your team look after that it "inherited" from elsewhere? Agree it's not fair to be on-call for things you can't fix. That's silly. But, you're surely not saying when you push a change to a 24/7 mission critical system, and you're on the "git blame" - you're not interested in taking responsibility if the consequences occur after 5pm? And now an SRE has to learn everything you know, in order to fix it, instead of asking for your help? Also - think about it a different way, you personally, are not the one I'd ask to participate in on-call. It's your whole team, maybe even your whole department. You and your team/department should work out how to field that together because you have more context into each others work/life balance, projects, skills, risky commits, etc. Your team would manage your own rotation. > Secondly, how many organizations allow developers to prioritise solving pages over other forms of revenue generating work? How many teams demand negotiation with a "product owner" before a piece of work can be added to the board? Tell me about it. Totally agree. This is where I focus much of my energy as a part of leadership teams - making other leadership realize they're over-promising on feature delivery without concern for their poor operations. Frankly, fuck sales-driven development for this. I even work with marketing departments to value-stream map so we include good ops & security as part of our core prop to customers and investors. If they don't get it? Well, I guess the fish is already rotting from the head, as they say. > Thirdly, how many pages are truly easy or tractable to fix? If they're repeated ad nausea? Then those alerts need to go. Now. And if they can't go, because Management don't see the pain they're causing? See my response to #2. Culture change first. > Fourthly, how often is the Real Problem™ for the reliability and uptime of a given system, actually at the program level, and not at architecture or system level? Depends on your industry/platform. Good SRE sysadmin troubleshooter types will do everything they can before they push the escalation button to engineer tier. At least they'll get all the data and metrics they can and document the situation at hand before the engineer comes online. Sometimes you don't really know what caused a problem until you all get together and retro it. Better to stay blameless even beyond when you have consensus on the technical factors. One time, it was a leap-second messing up some 3rd party vendor software, with consequences for our code base that nobody ever expected. Another time, it was a tired SRE combined with a lack of proper procedure combined with a code mistake with no test coverage. Another time, it was a binder that fell on a customer's space bar effectively DoS'ing a system that had no rate limiting. Some of those got fixed pretty quickly by an SRE. BUT - whatever the cause... the times that our MTTR got truly trashed? (We're talking DAYS...) It was because the person who wrote the code and had all the context, had already quit long ago - leaving us with no proper logging and no one to help us. The organization had never prioritized replacing their code or hiring someone who could understand it. > Really, this "shitty on call incentivises robust systems" argument is facile. You're right in the sense that, considering all of the above, I would need your buy-in. To get it, I'd seek to eliminate the "rot". I'm truly sorry you got burned in the past. I have also been there. When it was too much for me, I quit, and found somewhere more ideal. Now I just keep trying to work with people I trust as they move around in the industry, to avoid all the org B.S. you've rightly brought up.