3 ms·
Word of caution - we've observed 2 global (all regions down concurrently) networking outages on GCE this year. 18 minutes on April 11 and 1.5-1.7 hours last nig
by jread 10y ago
Word of caution - we've observed 2 global (all regions down concurrently) networking outages on GCE this year. 18 minutes on April 11 and 1.5-1.7 hours last night (except for us-west1 - down only 10 minutes):
https://status.cloud.google.com/incident/compute/16015 https://status.cloud.google.com/incident/compute/16015
https://status.cloud.google.com/incident/compute/16007 https://status.cloud.google.com/incident/compute/16007
https://cloudharmony.com/status-for-google https://cloudharmony.com/status-for-google
- ben_jones 10y agoWouldn't the phrase be "down simultaneously"?
- jread 10y agoI think concurrent is better in this context because the outage periods overlapped but were not synchronized.
- snewman 10y agoWait: seriously, a 1.5 hour multi-region GCE outage just last night? That would be huge, how is it that we haven't heard about this? The linked /16015 status report doesn't have much information.
- jread 10y agoI think partially due to the timing - 12-2AM PT vs 6-7PM on April 11. Google hasn't provided specifics - but our VMs in every region were 100% inaccessible during that period.
- Twirrim 10y ago"how is it that we haven't heard about this?" Not enough customers to complain / notice. Gartner report in the last week or two pointed out that Amazon is way in the lead with more cloud capacity than all the other providers combined. The top two clouds are AWS and Azure. When either of those have major incidents you hear about it all over twitter etc. because that's a large user base that gets impacted. Google is in third place (according to them), but dramatically far behind. They do praise Google's big data tools, but point out even then people use AWS for most stuff and Google just for their big data bits.
- boulos 10y agoThe incident report isn't done yet, so please be patient as the team finalizes the root cause and next steps. However, it didn't affect all routes, VMs, etc. and you'll see that when they publish the IR next week.
- jread 10y agoFYI - we verified outages from hundreds of last mile routes using Ripe Atlas probes with failure rates in the range of 86-96%.