5 ms·
Most jobs I've had there was an implicit expectation that if your shit breaks, you fix it. This resulted in systems which didn't break because, well, that keeps
by throwaway55421 5y ago
Most jobs I've had there was an implicit expectation that if your shit breaks, you fix it. This resulted in systems which didn't break because, well, that keeps work to working hours.
A few years back I took a job with no mention of on call and then was told there would be a fortnightly 24 hour slot, which later turned into ~weekly as team members left.
I left too.
I have no idea why they felt it reasonable to not mention this before I joined, they effectively just wasted a few months pay on me on a gamble that I had no life outside of work.
They seemed to find the idea that I don't take my laptop to the pub, or up a mountain, or on holiday, some sort of bizarre way of living. Sick system effect? Who knows.
- Wiseacre 5y agoSign of a wider lack of trust between employer and employee. Yet the idea of unionization is still very unpopular.
- throwaway55421 5y agoI don't follow. Most jobs I've had were high trust, this was the weird exception. I didn't need a union to fix that.
- Wiseacre 5y agoIt's not a rare phenomenon across the industry. Plenty of my colleagues and I have experienced something similar.
- cameronh90 5y ago> if your shit breaks, you fix it What if you're drunk?
- throwaway55421 5y agoYou think about that when you design it. I used to work on trading systems which were managed in real time by a rotating team of traders (i.e. mathematically oriented people as opposed to software engineers). You design the system so that if it breaks there are at least workarounds that anyone can employ without needing to write code. For example, a big stop button, manual adjustments, manual trading, etc. Maybe it stops printing money overnight and you take an opportunity cost loss. That's fine, post mortem at work, fix it. In the worst case if you're not available then someone with ownership steps in like a CTO/founder level (who are of course always on call almost by definition, though they generally have the executive power to say - sod this, we'll just leave it down for a while).
- 2-718-281-828 5y agothis is such a bullshit you are writing here. maybe justifiable if your payment was really decent - but only then. with your line of thinking reintroducing public flogging for bugs would probably also work very well. also you are naive with your argument that people just do their job right so they don't get called in the middle of the night. bugs and problems can arise out of nothing (dns problems in google cloud and stuff like that) also bugs aren't always immediately attributable. there you have your group pressure dynamics. no, thank you.
- odonnellryan 5y agoYou're being harsh but I do think it ignores many of the realities of most software positions where you are often pressured into deploying on Friday, without time to write many tests, etc.
- 2-718-281-828 5y agoyes - you're right I'm a bit harsh. but colleagues with this attitude help egotistic employers to turn work into hell for the rest and are even proud of it.
- odonnellryan 5y agoWell, all the yelling you and I would like to do isn't going to change the fact that the problem is with leadership.
- s0rce 5y agoSeems like "if your shit breaks, you fix it" = on call 24/7... isn't it better to have a specific shift where you know you are responsible and then you aren't the rest of the time.
- toast0 5y agoPersonally, I hate coordination and knowledge dissemination. It's better for me if I own my stack and support it 24/7 if that means I don't have to communicate details with other people regularly. I also gravitate towards teams that work well without a lot of documentation, so that helps in case I'm really unavailable.
- PeterisP 5y agoIf any system needs 24/7 support and more than two nines of availability, that's simply not compatible with having a single person "owning the stack" and not documenting/disseminating the details. That's a bus factor of 1, and fails whenever you will get any actual time off offline (e.g. in a plane). So if the uptime of the system and 24/7 is actually important, I would consider that your manager is unacceptably negligent in allowing this to happen. On the other hand, if 24/7/365 is more like a 'nice to have' option, then sure, such an arrangement will be more cost effective than more people in support; a 99% uptime SLA often can be maintained by one core person.
- toast0 5y agoI mean, two nines is 3 and a half days of downtime. That's not hard to get as long as I don't make a habit of breaking shit in ways that take a long time to find out, the hardware is decent and I don't disappear for days very often. I'd add that the network should be decent too, but if it's not, my stuff should be fine. A competent person can dig around and fix most anything good enough in three days and that leaves you 12+ hours for other outages. Three nines is almost 9 hours a year, but planes almost universally have wifi now, even if it's not great. If you have a bad year, with two big incidents when you're on a plane, you might not make your goal. At four nines, it's about an hour. Good luck with that one. You're kind of screwed either way. Someone with intimitate knowledge of the system can probably handle an incident or three in time, but it's an enormous amount of work to get people up to speed for that and you can't have very many live training sessions because you don't have the outage budget. At the same time, if you page me and I'm at the cinema, I'm not going to fix it in time. But, the biggest thing is working to make failure as graceful as possible. As much as possible, each component should be able to fail individually, without affecting the rest of the system. And, when there is an outage of a critical dependency, you need a way to push that outage status upstream and a method to manage load when it comes back up. (Those knobs should be stable and can certainly be documented for others).
- ozarkerD 5y agoGood thing I was drunk when I made it!
- NikolaeVarius 5y agoStop deploying shit then. Also, I am very good at fixing systems even when drunk. Lots of practice.
- doktorhladnjak 5y agoReminds me of my manager at a job a few years back. His standard was that any steps an on-call needs to take should be doable while they are drunk. You're waking people up in the middle of the night or they could be in who knows what sort of cognitive state. The runbook needs to be dead simple to follow. Mitigate now, fix it all up tomorrow when someone's awake and coherent.
- cameronh90 5y agoIf it is simple enough to do when drunk, why does it need to be done at all? Typically for us, when something goes wrong that needs intervention, it's the first time that issue has ever happened and it needs to be debugged and understood. Then we'll put a mitigation in so it hopefully doesn't happen again. In my last place, we had a runbook more like you describe, but I prefer prioritising bug fixes over features as it means I haven't had an out of hours call in a year!
- doktorhladnjak 5y agoThat was the other half of it though. Once it’s simple to mitigate, automate it! To be fair, this was on a team that had A TON of operational debt. Like you’d get paged 30 times _a day_ when on call. Nobody wanted to move to weekly on call from a daily rotation because it would be a total nightmare, but daily meant the incentive to fix something was low because tomorrow it would be someone else’s problem. Of course, the solution was to get the alerts and underlying software under control, which we did and things got better. Many of the non automated run books that remained were things like “look at these three graphs, if they’re doing x, escalate to team y”. We were somewhat hamstrung by political considerations that prevented us from simply directly paging those teams, or limitations of our alerting system that couldn’t link alerts (“if this other alarm is already going off, you don’t have to do anything because team y is already being paged for the root cause”).
- ipaddr 5y agoIf you can't fix your stuff while drunk you have never tried.
- 0x0000000 5y ago> Most jobs I've had there was an implicit expectation that if your shit breaks, you fix it. This resulted in systems which didn't break because, well, that keeps work to working hours. My experience was similar in my last long-term software role. We were not explicitly on call, but we developed/deployed/supported an app which was critical to a 24×7×365 business process. So we were careful in our testing, code reviews, and deployments. One time in 7 years I had to answer an out of hours call because of an issue in this app. I'll take that over semimonthly on-call rotas for suites of apps my team isn't isn't responsible for.