6 ms·
As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking. 75% of the post-mortem talks about the power outage
by yowai 3y ago
As someone who was slightly affected by this outage, I personally also find this post-mortem to be lacking.
75% of the post-mortem talks about the power outage at PDX-04 and blames Flexential. Okay, fair - it was a bit of a disaster what was happening there judging from the text.
But by end of November 2 (UTC), power was fully restored. It still took ~30 hours according to the post-mortem for Cloudflare to fully recover service. This was longer than the outage, and the text just states that too many services were dependent from each other. But I'd wish they go into more detail here why the operation as a whole took that long. Are there any take-aways from the recovery process, too? Or was it really just syncing data from the edges back to the "brain" that took this long?
Also one aspect I am missing here is the lack of communication - especially to Enterprise customers.
Cloudflare support was basically radio silent during this outage except for the status page. Realistically, they couldn't do much anyway. But at least any attempt at communication would be appreciated - especially for Enterprise customers, and even more especially after the post-mortem blames Flexential for a lack of communication.
While I like Cloudflare since it's a great product, I think there are still a few more things that should be taken as a conclusion for CF to take away from this incident.
That being said, glad you managed to recover, and thanks for the post-mortem.
- vb-8448 3y ago> Also one aspect I am missing here is the lack of communication - especially to Enterprise customers. They blame Flexential for lack of communication, but were the first one not saying anything.
- iAMkenough 3y agoEven "we don't know why our data center is failing, but we're sending a team over to physically investigate now" would have been A+ communication in the moment.
- NicoJuicy 3y agoEverything was on the status page since the start? DC related updates: > Update - Power to Cloudflare’s core North America data center has been partially restored. Cloudflare has failed over some core services to a backup data center, which has partially remediated impact. Cloudflare is currently working to restore the remaining affected services and bring the core North America data center back online. Nov 02, 2023 - 17:08 UTC > Identified - Cloudflare is assessing a loss of power impacting data centres while simultaneously failing over services. We will keep providing regular updates until the issue is resolved, thank you for your patience as we work on mitigating the problem. Nov 02, 2023 - 13:40 UTC
- yowai 3y agoAs an enterprise customer, I would expect a CSM reaching out to us informing us about the impact, getting into more details about any restoration plans and potentially even ETAs or rough prioritization to resolution on them. In reality, Cloudflare's support team was essentially completely unavailable on Nov 2, leaving only the status page. And for most of the day, the updates on the status page were very sparse except "we are working on it", and "We are still seeing gradual improvements and working to restore full functionality.". Yet clearer status updates were only giving starting on Nov 3. However, I still don't think I heard anything from support or a CSM during that time.
- NicoJuicy 3y ago? 1) Were you affected on the data plane? Which product? As far as I can tell, while the outage was in the core dc's. The impact was minor. 2) Both examples were exactly from 2 November. Not 3 November. 3) What method of support did you try? I thought that their support was impacted ( email?). The status page explicitly mentioned to get in contact with your account manager for some config changes on some products, if you wanted changes. 4) I have never heard of Enterprise customers being contacted by a cloud company during an outage. Which company does that? Do you have an example? 5) I would think it's absolutely a nogo to contact every preemptively Enterprise customer with: "hey, the product works, but if you change xyz, atm that doesn't.". Since most customers weren't affected and some others were minorly impacted. There is not a single cloud company that does that. Feel free to correct me if I'm wrong...
- MatthiasPortzel 3y ago> In particular, two critical services that process logs and power our analytics — Kafka and ClickHouse — were only available in PDX-04 but had services that depended on them that were running in the high availability cluster. Those dependencies shouldn’t have been so tight, should have failed more gracefully, and we should have caught them. This paragraph similarly leaves out juicy details. Exactly what services fail if logging is down? Were they built that way inadvertently? Why did no one notice?
- DylanSp 3y agoI'm not that surprised at the relative lack of detail, given how quickly they released this; I'm surprised they published this much info so quickly. Calling it a postmortem is a bit of a misnomer, though. I'd expect a full postmortem to have the kind of detail you mention.
- ecs78 3y agoI think they just wanted a quick post-mortem. I'm sure they will add more to the blog later in the year when they implement mitigations.