4 ms·
Spent about a decade on-call...It is an interesting and perhaps even worthwhile experience to do it for a little bit. IMHO there are two angles to achieve high
by srer 5y ago
Spent about a decade on-call...It is an interesting and perhaps even worthwhile experience to do it for a little bit.
IMHO there are two angles to achieve higher reliability:
1. Build systems which go down less on their own (which is quite difficult)
2. Get good at fixing them fast when they go down (which is much easier)
A trivial example of 1 might be moving from no RAID to RAID1, an example of 2 might be getting on-call proficient at quickly responding to a dead disk and restoring from backups.
(But I did say 1. was hard, even in this simple example maybe you used hardware raid, and are finding out just how crappy HW raid can be ;)
Companies attempt both angle 1 & 2 to differing degrees. I worked a lot in 2, and it sucked. We were a proficient fire department in a city of straw houses and gas stoves.
Aside from the general suckage of on-call (24 hour days, many, many nights lost sleep), the work is by it's nature high risk, high urgency, high impact, high risk - but is rarely rewarded sufficiently in pay or promotion opportunities.
Having spent so much time on angle 2, I've decided it is largely shallow work, and to grow intellectually I needed harder more interesting problems, which meant trying angle 1.
Consequently now my title is SWE and I don't do out of hours on-call, and it is glorious. I do notice I see the world differently to my less scarred fellow SWE. It definitely changed how I write software, and view the product priorities.
At heart I still consider myself an SRE. It's just I had to place myself in the most impactful place to achieve it, as a regular product/feature dev.