3 ms·
GCP has some nice things but AWS is simply more reliable. The GCP outage [1] in April is a great example. A fire in europe-west9-a "zone" took down the entire
by Shakahs 3y ago
GCP has some nice things but AWS is simply more reliable.
The GCP outage [1] in April is a great example. A fire in europe-west9-a "zone" took down the entire europe-west9 "region" (because this entire "region" is actually housed in a single datacenter), which then caused a global GCP console/API outage because GCP's single global control plane couldn't reach europe-west9.
The fact that a zonal issue escalated to a regional and then global outage shows that GCP is not serious about limiting blast radius and all their talk of independent regional/zonal infrastructure is marketing fluff.
In comparison, AWS AZs can be up to 60 miles apart, share no physical infrastructure, and every region hosts its own API/console/control plane.
Also, Nitro. AWS has essentially eliminated the performance overhead of virtualization with Nitro, and years later GCP still has no equivalent so GCP customers still have to pay for that overhead.
1: https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPfkY https://status.cloud.google.com/incidents/dS9ps52MUnxQfyDGPf...
- esprehn 3y agoThat Incident Report is phrased awkwardly. GCP does not have a single global control plane. Each product has a separate control plane, and many/most of those are regional. GCE (Compute) does have a global control plane today. The GCP Console, unlike other cloud providers, displays an aggregated view over all the regions. That means calling the control plane for a particular service which then in turn will do fan-out to all regions and merge the results. For example a page might call the GCE global control plane: https://cloud.google.com/compute/docs/reference/rest/v1/instances/aggregatedList https://cloud.google.com/compute/docs/reference/rest/v1/inst... Or a different page might call this with - as the location to globally fan-out instead of talking to a particular region's control plane: https://cloud.google.com/functions/docs/reference/rest/v2/projects.locations.functions/list https://cloud.google.com/functions/docs/reference/rest/v2/pr... (Specifying the location would call that particular region though) During the incident the small (but critical) subset of Console pages which needed aggregated data from GCE's global control plane failed to load. I agree it's unacceptable that the Console doesn't handle regional failures more gracefully, and it's very much been a priority to improve region down handling. Source: I work on the GCP Console.
- rwiggins 3y ago> because this entire "region" is actually housed in a single datacenter From the incident report you linked: "a cooling system water pipe leak occurred in one of the data centers in the europe-west9 region [...] Europe-west9 contains three buildings with independent cooling, power, and networking" Also, > a global GCP console/API outage because GCP's single global control plane couldn't reach europe-west9. again, from that page, "A small number of methods within the GCE control plane API must collect information from multiple regions or zones by making requests to each regional control plane (called fanout requests). Google Cloud services including Cloud Console depend on these methods. When the GCE control plane for the europe-west9 region and zones went offline, some of these fanout methods did not operate correctly. During the outage, this led to global unavailability for some pages and control plane operations within Cloud Console" It's certainly not great that a regional problem had global impact, but there is some nuance. For example, that paragraph talks about "each" regional control plane, in addition to a global control plane. I don't mean to belittle the importance of the incident. Suffice it to say lessons were learned and follow-up changes were made. As for Nitro - take a look at C3. https://cloud.google.com/blog/products/compute/introducing-c3-machines-with-googles-custom-intel-ipu https://cloud.google.com/blog/products/compute/introducing-c... Disclosure: I work on GCE.
- retinaros 3y agobigquery or table was down for a week
- deadmutex 3y ago> Nitro. AWS has essentially eliminated the performance overhead of virtualization with Nitro, and years later GCP still has no equivalent FYI, per https://cloud.google.com/blog/products/compute/introducing-c3-machines-with-googles-custom-intel-ipu https://cloud.google.com/blog/products/compute/introducing-c... The Compute Engine C3 machine series, now available in Private Preview, is the first VM in the public cloud with the 4th Gen Intel Xeon Scalable processor and with Google’s custom Intel IPU. C3 machine instances use offload hardware for more predictable and efficient compute, high-performance storage, and a programmable packet processing capability for low latency and accelerated, secure networking Disclosure: I work for GCP