4 ms·
Your comment gives me the impression that you massively under-estimate the level of specialization in either software or operations, or perhaps both. Some syst
by throwawaygh 5y ago
Your comment gives me the impression that you massively under-estimate the level of specialization in either software or operations, or perhaps both.
Some systems are simple enough that a single person can understand everything about the entire stack. Hell, I still run downright ancient LAMP stacks on colocated servers where I'm the only one familiar with the code base and I interact directly with the dude on the floor when I need a hard powercycle.
But there are a lot of systems in which the ops side is incredibly complex (think compute-intensive features with constant churn in the underlying hardware, or fleets of embedded devices that are always failing in new and fun ways, or finicky lab equipment, etc. Have you ever run a fleet of machines each hooked up to and controlling piece of $x,xxx,xxx equipment, where each of those pieces of equipment is a snowflake? And where all the data from that fleet of equipment gets pulled off the machines attached to the equipment and sent over to a computer cluster the size of a small DC?).
There are also lots of systems in which the software requires a lot of specialization, either technical or subject matter (in the above case: the software running on that cluster was written by and requires the expertise of phds in a few different natural sciences + CS, and the hardware at the time was so bleeding edge that we were essentially the chip maker's and system integrator's beta testers.)
The real issue is when these two types of systems very often overlap.
And, like, this sounds special, but it's really not. There are projects like this is every sector -- life sciences, finance, tech, basically anything in a hospital is a 50/50 chance of being like this (mostly due to externally imposed complexity), agri of course, even property insurance. I'd say that "ops and software need separate on-calls" probably describes most successful software products/orgs.
Even if one person did have enough time in their career to learn how to be the oncall for both the software and the ops, there's no way they could keep up with all the moving parts of both jobs to be able to do both jobs at the same time. At least, not without burning out pretty quickly. At some point these have to be separate jobs because just keeping up with all the moving parts well enough to do an oncall shift is by itself a >60 hour/week job. And that's just to stand still in the feature set. At some point we wanna get new work done.
FWIW, I haven't been paid to work on the former type of system -- where both aspects are simple enough for one human to first grok and then more importantly keep up with -- in a long time. I sort of