5 ms·
Seems like "if your shit breaks, you fix it" = on call 24/7... isn't it better to have a specific shift where you know you are responsible and then you aren't t
by s0rce 5y ago
Seems like "if your shit breaks, you fix it" = on call 24/7... isn't it better to have a specific shift where you know you are responsible and then you aren't the rest of the time.
- toast0 5y agoPersonally, I hate coordination and knowledge dissemination. It's better for me if I own my stack and support it 24/7 if that means I don't have to communicate details with other people regularly. I also gravitate towards teams that work well without a lot of documentation, so that helps in case I'm really unavailable.
- PeterisP 5y agoIf any system needs 24/7 support and more than two nines of availability, that's simply not compatible with having a single person "owning the stack" and not documenting/disseminating the details. That's a bus factor of 1, and fails whenever you will get any actual time off offline (e.g. in a plane). So if the uptime of the system and 24/7 is actually important, I would consider that your manager is unacceptably negligent in allowing this to happen. On the other hand, if 24/7/365 is more like a 'nice to have' option, then sure, such an arrangement will be more cost effective than more people in support; a 99% uptime SLA often can be maintained by one core person.
- toast0 5y agoI mean, two nines is 3 and a half days of downtime. That's not hard to get as long as I don't make a habit of breaking shit in ways that take a long time to find out, the hardware is decent and I don't disappear for days very often. I'd add that the network should be decent too, but if it's not, my stuff should be fine. A competent person can dig around and fix most anything good enough in three days and that leaves you 12+ hours for other outages. Three nines is almost 9 hours a year, but planes almost universally have wifi now, even if it's not great. If you have a bad year, with two big incidents when you're on a plane, you might not make your goal. At four nines, it's about an hour. Good luck with that one. You're kind of screwed either way. Someone with intimitate knowledge of the system can probably handle an incident or three in time, but it's an enormous amount of work to get people up to speed for that and you can't have very many live training sessions because you don't have the outage budget. At the same time, if you page me and I'm at the cinema, I'm not going to fix it in time. But, the biggest thing is working to make failure as graceful as possible. As much as possible, each component should be able to fail individually, without affecting the rest of the system. And, when there is an outage of a critical dependency, you need a way to push that outage status upstream and a method to manage load when it comes back up. (Those knobs should be stable and can certainly be documented for others).
- throwaway55421 5y agoNo, because in that case you are responsible for other people's systems you have no idea about and there is no incentive to fix them. I am always on call to secure my own house, so I have a good alarm system, cameras, locks etc which means I hopefully don't have to do much. It doesn't keep me up at night when I'm on holiday because, well, it's overwhelmingly likely that nothing will happen. If I were on call for the neighbourhood or general area once a week, I'd probably have to physically be there patrolling it because otherwise I have liability for things which are outside of my control.
- maximus-decimus 5y agoAnd if somebody ever leaves the company for any reason, who is on call for their stuff?
- throwaway55421 5y agoProjects get reassigned over time regardless. It's not as if this is some strict 1:1 relationship, we obviously had more than 1 person with knowledge of / working on a codebase. It would have been an issue if 2-3 working on the same thing left simultaneously, but for more reasons than just on-call.
- maximus-decimus 5y agomaybe there is a big difference between my workplace and yours, but for me a lot of stuff never gets improved much after it's written so it's basically impossible to get up to speed until the person who owned it leaves the company and now you're she one fixing their bugs. Onboarding is just really hard on legacy projects.