4 ms·
And I'm curious about the work culture park. They're a small organisation in terms of workforce, but somehow managed to respond within minutes on Saturday. How
by d33 7y ago
And I'm curious about the work culture park. They're a small organisation in terms of workforce, but somehow managed to respond within minutes on Saturday. How does that work? Do they have shifts or somehow the workers are so devoid of private life that they have to respond during early morning on weekend?
- icebraining 7y agoIn California, it was Friday evening still.
- m_eiman 7y agoI assume they probably have 24h support, but also consider time zones: 3 AM Saturday in UTC is Friday evening in San Fransisco: https://duckduckgo.com/?q=03%3A00+utc+to+pst&ia=answer https://duckduckgo.com/?q=03%3A00+utc+to+pst&ia=answer
- michaelt 7y agoThe blog post only says they halted issuance within minutes of confirming the bug - not that they confirmed the bug within minutes of receiving a bug report.
- pgporada 7y agoIt takes time to correctly determine if a bug report is real and to determine the possible scope of the bug.
- lidHanteyk 7y agoThey don't have many incidents or outages. As a result, it's much easier to respond to the incidents that do occur.
- tialaramex 7y agoFor me at least 24/7 incident response is completely acceptable in a properly compensated role so long as it's accompanied by the culture that says preventing such incidents in the first place is Job #1 That is, I'm OK with being woken at 0200 to try to understand and if appropriate fix or recover from a disaster only so long as if I'd suspected this might happen the people expecting me to be awake at 0200 would have given me the resource (money, people, whatever) to fix it. If I feel like I don't have that support, I'll only start looking at your disaster during my working day. My impression is that ISRG pays a lot of attention to preventing disasters, so if I worked for ISRG (not very practical since they're based on the US West Coast and I live in England) I'd be comfortable taking a call in the middle of the night to fix things.
- ithkuil 7y agooperations based on US west coast could definitely benefit from a few people on the other side of the world to achieve a 24/7 coverage while keeping a good work-life balance.
- myself248 7y agoLocal nerds can be noctournal too. Letting people pick their preferred shift is just as important as accommodating other kinds of physiological diversity.
- pgporada 7y agoCan confirm.
- folmar 7y agoYou'd normally want the 24/7 people to be part of the day to day operations, otherwise they will quickly stop being up to date, so their don't go out of the knowledge loop, so selecting the reasonable timezone set is not trivial.
- ithkuil 7y agoyeah, my assumption was that the team in another TZ would also work together on the same thing. Yes it has some challenges, but there are a few upsides beside the oncall coverage (e.g. increase in talent pool)
- bcrosby95 7y agoYeah, as long as it doesn't happen often. I'm technically always on call but we haven't had an on call incident in close to a year. Basically I keep a phone and laptop on me at all times. This is in comparison to a friend that works somewhere that always has daily on call incidents that are not actually problems 95% of the time. That would piss me off even if I weren't always on call.
- 7y ago
- dboreham 7y agoOut of hours response to a critical problem is standard. You can achieve it in various ways but they all boil down to people who know what they're doing having a professional ethic. Typically it isn't possible to have shifts of engineers with deep understanding of the code on call so ultimately you need to wake someone up. So remember to keep a note (up to date) of key staff home phone numbers, their home addresses.
- pgporada 7y agoSpeaking for myself here. The workforce is spread across the states. When you're drawn to the mission like I am, late nights here and there don't matter at all. I communicated with my wife and I'm sure others informed their significant others what was going on etc and why Friday, Saturday, Sunday, Monday, and Tuesday would be thrown out of wack. Members of the team put in much more hours than I did and that is truly impressive. It takes all of us with our different specialties to make an accurate and effective response. Some of the things we do are internal post mortems and find ways to prevent the issue from happening again by either improving alerting/monitoring, writing a runbook, fixing code, and fixing misconceptions about a part of the entire system. We do weekly readings of various RFCs, the Baseline Requirements, and other CAs CP and CPS documents to again better understand our system and Web PKI as a whole. This is an understatement, but we heavily rely on automation. From the moment the call was made to stop issuance, an SRE was ready to run the code that disables the issuance pipeline. The biggest takeaway is that communication and leadership makes all the difference. I have to go, there's work to be done.
- jaas 7y agoHead of Let's Encrypt here. We have an on-call rotation and a system for getting others notified and online quickly when necessary. We make sure not to bring too many people online so that some people are fresh and can rotate in later if the incident lasts longer. It's not often that staff have to put in time at night or on weekends, and when it happens we work hard to make sure the problem doesn't happen again.