4 ms·
While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it
by gregdoesit 4y ago
While the post-mortem is thorough, it misses key details on what companies experienced who were unlucky enough to be caught out by this outage. For example, it fails to mention how impacted customers lost access to certain Atlassian services for up to ~2 weeks: JIRA, Confluence, OpsGenie. But not others like Trello or BitBucket.
Of these, losing access to OpsGenie for this long was a massive problem, dwarfing most other systems. OpsGenie is like PagerDuty in the Atlassian world.
I spoke to several engineers at impacted companies who could not believe their incident management system was “deleted” and had no ETA on when it would be back, or Atlassian could not prioritise restoring this critical system ASAP. JIRA and Confluence being down was trouble enough, but those systems being down for some time was things most teams worked around. However, suddenly flying blind, with no pager alerting for their own systems? That is not acceptable for any decent company.
Most I talked with moved rapidly to an alternative service, building up oncall rosters from memory and emails - as Confluence which stored these details was also down. Imagine being a billion dollar company suddenly without pager system: and no ETA on when that system would be back, your vendor not responding to your queries.
I talked to engineers at such a company and it was a long night to move rapidly over to PagerDuty. It would be another 7 days they could get through to a human at Atlassian. By that time, they were a lost customer for this product. Ironically, this company moved to OpsGenie a few years before from PagerDuty because OpsGenie was cheaper and they were on so many Atlassian services already.
The post-mortem has no actions on prioritising services like OpsGenie in reliability or restoration, which is a miss. I can’t tell if Atlassian staff are unaware of the critical nature of this system or if they treat all their products - including paging systems - as equals in terms of SLAs on principle.
Worth keeping in mind when choosing paging vendors - some might recognise these systems are more critical ones than others.
I wrote about this outage from the viewpoint of the customers as it entered its 10th day and it was discussed in HN, with comments from people impacted by the outage. [1]
[1] https://news.ycombinator.com/item?id=31015813 https://news.ycombinator.com/item?id=31015813
- Aeolun 4y agoBut you don’t need a paging system if your services don’t go down. Isn’t the fact you can’t deal without one for two weeks an indictment of your own practices?
- na85 4y agoSure and you don't need emergency locator transmitters if your aircraft doesn't crash. When you're ready to prove that your services "don't go down" send me an email and I'll come work for you.
- ramraj07 4y agoEven if they do, I wouldn’t want to work for them
- pixeltopic 4y agoWell said - you don't need to write tests for your code if you don't write bugs!
- Aeolun 4y agoThat’s not equivalent. Code changes every day (potentialy quite extensive), but the same thing is not true for infra, which hopefully stays mostly the same from day to day. If your incident reporting is down, hopefully you completely stop changing anything about your infra.
- pixeltopic 4y agoIt's naive to expect that things will just stay fine and dandy if you "stop changing" your infra. Consider scenarios such as traffic spikes, network outages, and under-provisioned resources.
- throwaway787544 4y agoIt should be shouted from the rooftops: don't switch services just to save money if the result is potentially worse business outcomes. Why save a tiny bit of cash if it puts your business at risk?
- spondyl 4y agoThe saying "Cheap is expensive" comes to mind
- Griffinsauce 4y agoAlso "don't put all your eggs in one basket" This is not very far away from AWS hosting their own status pages.
- giraffe_lady 4y agoPresumably no one who made that decision is fucking stupid and thought they were putting their business at risk. Good lord. I wasn't affected by this in the slightest and just found out opsgenie exists from the parent comment but even I can understand that this decision would almost certainly be driven by things like "we're already using atlassian for everything else and will benefit from the interop" and "we already trust them with everything else and they haven't let us down or we wouldn't be using them for any of that stuff either."
- johnzabroski 4y agoIn my experience, the decision is driven by a VP who wants a bullet point on their year-end review, namely "notional cost savings achieved".
- gtirloni 4y agoIn my experience there isn't much interop that's worth anything with OpsGenie. It's a different story with JIRA/Confluence.
- throwaway787544 4y agoThere's no need to get that upset over it... When we switched to OpsGenie it was also to consolidate billing with other Atlassian products, but we all knew PagerDuty worked better and that OpsGenie was still pretty new and rough around the edges. We certainly didn't gain any business advantage by switching, but we did have to do a lot of work to switch which took away from other things we needed to get done. But ultimately we had no say because somebody just wanted to save a little money. I also doubt that there was a thorough vendor assessment before we picked it up, since it was a vendor we already used.
- BonoboIO 4y agoAs a customer I would not buy or use an Atlassian product in a 1000 years. 14 days without pagers ... From friends and people I know I heard nothing good about Atlassian products. And that was before that 2 weeks downtime. It looks like the products are duct together with duct tape, spit and a little bit of dirt.
- bpicolo 4y ago> 14 days without pagers There’s an interpretation of that where life is great