5 ms·
> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had no
by Dunedan 3y ago
> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster.
> The other two data centers running in the area would take over responsibility for the high availability cluster and keep critical services online. Generally that worked as planned. Unfortunately, we discovered that a subset of services that were supposed to be on the high availability cluster had dependencies on services exclusively running in PDX-04.
> A handful of products did not properly get stood up on our disaster recovery sites. These tended to be newer products where we had not fully implemented and tested a disaster recovery procedure.
So the root cause for the outage was that they relied on a single data center. I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet.
- SushiHippie 3y agoAnd the top comment on the other HN post called it: https://news.ycombinator.com/item?id=38113503 https://news.ycombinator.com/item?id=38113503
- thelastparadise 3y ago> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Absolute lack of faith in cloudflare rn. This is amateur hour stuff. It's especially egregious that these are new services that were rolled out without HA.
- NicoJuicy 3y ago? Tbh. As far as I can see, their data plane worked at the edge. Cloudflare released a lot of new products and the ones that affected were: streams, new image upload and logpush. Their control plane was bad though. But since most products worked, that's more redundancy than most products. The proposed solution is simple: - GA requires to be in the high availability cluster - test entire DC outages
- throwaway6920 3y ago> Tbh. As far as I can see, their data plane worked at the edge. Arguable, it's best to think of the edge as a buffering point in addition to processing. Aggregation has to happen somewhere, and that's where shit hit the fan.
- NicoJuicy 3y ago? That would mean their data is at the core cluster. That's not true or I haven't seen any evidence to support that statement. Cloudflare's data lives in the edge and is constantly moving. The only thing not living in the edge ( as was noticed), is stream, logpush and new image resize requests ( existing ones worked fine) from the data plane
- throwaway6920 3y ago>That would mean their data is at the core cluster. That's not true or I haven't seen any evidence to support that statement. You're being loose in your usage of 'data'. No one is talking about cached copies of an upstream, but you probably are. Read the post mortem a bit more closely. They explicitly state that the control plane(s) source of truth lives in core, and that logs aggregate back to core for analytics and service ingestion. Think through the implications on that one.
- e1g 3y agoThat’s my interpretation as well. There is one central brain, and “the edge” is like the nervous system that collects signals, sends it to the brain, and is _eventually consistent_ with instructions/config generated by the brain.
- yowai 3y agoFrom what I was reading on the status page & customers here on HN, WARP + Zero Trust were also majorly affected, which would be quite impactful for a company using these products for their internal authentication. It's not just streams, image upload & Logpush.
- simion314 3y ago[flagged]
- sophacles 3y agoSounds like chatgpt doesn't want your business and tuned thier cloudflare settings accordingly. Conveniently cloudflare is getting the blame, which is presumably part of what they're paying for.
- rezonant 3y agoYep, it's easy to spot folks who have never configured Cloudflare's WAF when they suggest Cloudflare is blocking their browser of choice instead of the website itself.
- simion314 3y ago>Sounds like chatgpt doesn't want your business and tuned thier cloudflare settings accordingly. Conveniently cloudflare is getting the blame, which is presumably part of what they're paying for. The issue is fixed now. But as I mentioned CloudFlare still has a shit captcha, and the one for disabilities was broken as I mentioned.
- creshal 3y ago> I find that pretty embarrassing for a company like Cloudflare, which powers such relevant parts of the internet. Bah, who cares about such unimportant details, what's important is that ~dev velocity~ was reaaally high right until that moment! > We were also far too lax about requiring new products and their associated databases to integrate with the high availability cluster. Cloudflare allows multiple teams to innovate quickly. As such, products often take different paths toward their initial alpha. While, over time, our practice is to migrate the backend for these services to our best practices, we did not formally require that before products were declared generally available (GA). That was a mistake as it meant that the redundancy protections we had in place worked inconsistently depending on the product. Complete and utter management failure. And customers apparently are sold what Cloudflare internally considers to be alpha quality software?
- marcinzm 3y ago> Complete and utter management failure. And customers apparently are sold what Cloudflare internally considers to be alpha quality software? This has been my experience with AWS and GCP as well. Assume anything that's under 3 years old is not really GA quality no matter what they say publicly.
- arrakeenrevived 3y agoI've been involved with some new service launches at AWS, and it's a strict requirement that everything goes through some rigorous operational and security reviews that cover exactly these issues before the service can be launched as GA. Feature-wise people might consider them "alpha", but when it comes to the resilience and security of the launched features, they are held to much higher standards than what is being described in this post-mortem.
- phan 3y agoYour operational reviews must be lacking at AWS then (surprise surprise) then because there are so many instances where something will be released in alpha yet the documentation will still be outdated, stale and incorrect LOL.
- davedx 3y agoAnd that this was unironically written in the same post mortem: “We are good at distributed systems.” There’s a lack of awareness there.
- steve1977 3y agoWell, they did distribute their systems. Some were in the running DC, some were not ;)
- ZiiS 3y agoThey are good at systems that are distributed; they are very bad at ensuring systems they sell thier custoners are distributed.
- emadda 3y agoTheir uptime was eventually consistent
- ecs78 3y agohaha. The control plane was eventually consistent after 3 days
- belter 3y agoThey distributed the faults across all their customers....
- brookst 3y agoGood != infallible
- troyvit 3y ago> I am sorry and embarrassed for this incident and the pain that it caused our customers and our team. So do they.
- cyberax 3y ago> While most of our critical control plane systems had been migrated to the high availability cluster, some services, especially for some newer products, had not yet been added to the high availability cluster. It's amazing that they don't have standards that mandate all new systems to use HA from the beginning.