7 ms·
Major data center power failure (again): Cloudflare Code Orange tested
- Olesya000 3y ago[dead]
- qmarchi 3y agoTwo major outages less than half a year a part, but with wildly different outcomes. It's definitely showing their engineering capabilities were targeted at the correct outcomes. Would definitely be interested to see the detailed RCA on the power side of things. Not many people really think about Layer 0 on the stack.
- mike_d 2y agoCloudflare is in EdgeConneX Portland, you can try poking around but I haven't seen a RCA of what happened. Details are usually only shared with direct customers because it is bad for the brand. https://www.edgeconnex.com/wp-content/uploads/2018/10/ECX-22-47-PORTLAND-2023-DATA-SHEET-V5.pdf https://www.edgeconnex.com/wp-content/uploads/2018/10/ECX-22...
- virtuallynathan 2y agoThey called out Flexential in the post?
- Jamie9912 2y agoYes
- mannyv 3y agoSomeone set up the breakers incorrectly way back when, and they were never adjusted. I'll bet it's not possible to adjust those without powering off the downstream equipment. It reminds me of the amazon guy discovering that there was no way to fail back power without an outage, then them going off and building their own equipment.
- SteveNuts 3y ago> there was no way to fail back power without an outage, then them going off and building their own equipment. Anywhere I can read about that?
- michaelt 3y agoMaybe [1] which is about per-rack UPSes, to reduce the blast radius of UPS failure? Pretty sensible IMHO - I live in a country with a reliable electricity grid, and outages due to UPS malfunction are about as common as power outages. [1] https://www.datacenterdynamics.com/en/news/aws-develops-its-own-rack-ups/ https://www.datacenterdynamics.com/en/news/aws-develops-its-...
- nyrikki 3y agoIn large data centers, rack level UPSs are impractical for many reasons like cost and efficiency, but the big problem is that modern power densities are so high that you want rows to fail if cooling isn't available. It doesn't take long without cooling to cook equipment to the point of failure or reduced reliability. 7 to 16kW per rack is common even in these older colo facilities. And there never would have been enough UPS to make up for not enough replacement breakers on site.
- michaelt 2y agoBut isn't the UPS only expected to last for 15 minutes or so, to give the backup generators time to start up? Or to perform a fast-but-graceful migration when the generator doesn't start up? I thought most DCs just pause the cooling until the generator comes up, rather than running the cooling on battery power?
- throw0101d 2y ago> But isn't the UPS only expected to last for 15 minutes or so, to give the backup generators time to start up? They are expected to last that long, but if the batteries are on year 4 of their 5-year life, that may not happen. What also may not happen is the generator starting up. Or the automatic transfer switch (ATS) not working properly: it should be on either input Feed A or Feed B, but when it tries to throw itself over (making a loud kah-chunk sound), it gets stuck in between—this happened to us once. "The perversity of the universe always tends toward a maximum." — Finagle's Law of Dynamic Negatives
- emmanueloga_ 3y ago[flagged]
- decasia 3y agoAs always, it's really impressive to see how much technical detail they release publicly in their RCAs. It sets a good example for the industry. Also — quite impressive to make major infrastructure and architecture changes in a few months. Not every organization can pull that off.
- Waterluvian 3y agoI feel there’s a sweet spot where if you do it too quickly, it’s a bad sign. And if it takes years, risk just keeps going up and up until it becomes basically impossible to do smoothly.
- zamalek 2y ago> quite impressive to make major infrastructure and architecture changes in a few months And have it work first time round.
- belter 2y agoThere is not much impressive here. Their architecture seems to be relying on a couple of datacenters, and if somebody have turned on or not the right switches in the right places. This means you will continue to hear from regular outages and maybe nice postmortem blogs. I don't see a fundamental approach to reliability based on concepts around blast radius, expecting anything to fail anytime, and fundamental principles of bulkhead patterns. Here is a free hint: By talking so much about where their data centers are located, on my view, they already failed item one on my check list. Principle number one of Physical Security is, you don't say where your Data Centers are, except of course to very restricted number of "need to know group". As predictions have no value, unless they are made prior to events, I predict the next outage will be some common core component, with some on/off type of config, with some common core configuration, that "could not be foreseen". Then to the next blog...On to the next outage....
- mleonhard 2y agoThe blog post contain no details at all about how they achieved high availability. I'm disappointed.
- andrewaylett 3y agoI can very definitely empathise with the experience of having worked hard at fixing the issues underpinning high priority incidents, then noticing that what previously would have taken hours to fix is now only visible as a blip on a graph.
- llbeansandrice 3y agoA single k8s cluster spanning multiple datacenters feels mind boggling to me. I know it's not exactly uncommon for HA even if you just have a little one in your cloud provider of choice but I'm sure it's a totally different beast than the toy ones I've created.
- Havoc 2y agoI guess if you have really fast fibre connections between them then the separation isn’t all that separate even with distance
- ekimekim 2y agoAt its core Kubernetes is mostly a HTTP API being used to sync state between nodes. I see no reason that part shouldn't work even across the world, albeit at the cost of slower syncing of state (eg. pods taking longer to be created after a scale-up). That HTTP API is backed by etcd which uses Raft and that is where running over a large area is more likely to cause problems. One approach would be to keep the etcd instances in one region (and probably also co-locate the scheduler, controller manager there), while having far-flung worker nodes. This creates a risk of losing the control plane but in most cases services would keep running (but would be unable to react to any further issues until the control plane recovers). You would also want to carefully design your workloads with topology constraints and region-specific services to avoid high application layer latencies though. Overall it's a fun thought experiment. In practical terms I think cross-datacenter in a small geographic area would work fine but I probably wouldn't want to run a single worldwide cluster, both for the reasons above and for other scaling reasons.
- oceanplexian 2y agoYou could implement geographic taints and tolerations and constrain certain workloads to certain regions. Lots of places span clusters across several AZs and have the same problem in theory. However I don’t personally like it, because you’re engineering a cross-datacenter failure domains which is usually a bad design decision unless done for specific reasons. And as for the the control plane, for it to reliably span multiple locations while also tolerating failures, you now have to deal with all kinds of weird split brain scenarios, so most just run multiple clusters instead of rewriting k8s for this kind of design.
- alberth 2y agoSingle Point of Failure Is PDX still a single-point-of-failure for Cloudflare services? It was 5-months ago [0], and if I understand the post - it sounds like it still is. If anyone knows, I'd be curious to hear. [0] https://news.ycombinator.com/item?id=38113503 https://news.ycombinator.com/item?id=38113503
- internetter 2y agoFrom what I understand of what I read, it was a single point of failure for one product instead of 15
- CodeWriter23 2y ago...and they were already on fixing the 1.
- vlovich123 2y agoIt’s complicated and it’s not supposed to be. PDX is where the config plane for the edge network and services is stored. Some of this information is transported to the edge via QuickSilver [1] (e.g. auth tokens). This information is replicated to Europe and fail over is possible. The challenge with the previous outage was a combination of things (as it always is) whereby certain services had a dependency that relied on PDX as a single point of failure (if I recall correctly). That underpinned enough services that a good chunk of Cloudflare’s config plane went down. Additionally, the cutover to Europe didn’t go smoothly because it was instantaneous instead of gradual which resulted in traffic amplification as retries for previous requests & current requests were shuttled into the online data center resulting in a thundering herd. What this blog post talks about it is how this time nothing went down (or at least cut over within minutes due to presumably automated systems noticing & doing it) with the exception of analytics data which does have a single dependency and that’s determined to be “ok” (i.e. it’s an acceptable failure mode for that product & that product only). There are additional failure complications required because PDX (& Europe) is composed of several independent data centers which aren’t supposed to all fail simultaneously. What’s pretty clear from the implied “we’re not happy either” is that Cloudflare isn’t pleased with their vendor’s separately located data centers still having correlated failures. [1] https://blog.cloudflare.com/introducing-quicksilver-configuration-distribution-at-internet-scale https://blog.cloudflare.com/introducing-quicksilver-configur...
- Jamie9912 2y agoInteresting, I didn't even hear about that second outage
- slyall 2y agoSame. My employer was highly impacted by the November 2nd outage but this latest one didn't appear to have affected us at all.
- Terretta 2y ago> "When one or more of these breakers tripped, a cascading failure of the remaining active CSB boards resulted, thus causing a total loss of power serving Cloudflare’s cage and others on the shared infrastructure." Background note to HN readers: Almost zero SaaS providers (or even CDNs) using the term "our datacenter" or showing their datacenters on maps etc. have their own datacenters. It's universal and normal. In general they have a server, a rack, a cage, in shared space, subject to others' policies and practices, and their neighbors. This can adjust your mental model for accountability and your designs for resilience. You can even exploit this by colo-ing at the same addresses to get LAN latencies to your SaaS provider, CDN, or (sometimes) even cloud provider.