3 ms·
According to network checks we have running against VMs in each GCE region, this event resulted in about 1.6 hours of concurrent ICMP timeouts for every region
by jread 10y ago
According to network checks we have running against VMs in each GCE region, this event resulted in about 1.6 hours of concurrent ICMP timeouts for every region (except us-west1 for only 10 minutes). We use Panopta for monitoring which verifies outages from multiple external network paths. When outages trigger we also use Ripe Atlas to confirm them using 100s of last mile network paths of which 85-95% resulted in timeouts. This is the second global GCE networking outage this year. These outages are particularly problematic because even multi-region load balancing will not avert downtime.
https://cloudharmony.com/status-for-google https://cloudharmony.com/status-for-google
Prior global outage - April 11:
https://status.cloud.google.com/incident/compute/16007 https://status.cloud.google.com/incident/compute/16007
Disclaimer: I am the founder of CloudHarmony
Edit: Outages triggered due to ICMP timeouts
- sshykes 10y agoNot accusing you of being intentionally misleading, but it would be nice if you put a clear disclaimer pointing out that you work for CloudHarmony.
- jread 10y agoI founded CloudHarmony and now work for Garter which acquired CloudHarmony in 2015.
- simonebrunozzi 10y agoI think it would still be nice to mention that you work there.
- deleted 10y ago[deleted]
- jread 10y agoIs there a reason for the downvotes?
- slau 10y agoMost probably because you didn't edit to post a disclaimer. It just looks like someone trying to profiteer off of GCE's incident.
- jread 10y agoI've added a disclaimer. We do not profit in any way from these incidents.
- GilbertErik 10y agoWhile it's true that an auto insurance company doesn't profit off car crashes... if people _never_ had car crashes, we might not need auto insurance companies.
- ajkjk 10y agoPresumably because your post feels like an ad.
- lima 10y agoTheir post-mortem states that only traffic for protocols other than TCP and UDP was dropped. Does your monitoring take this into account? It states "1.67 hours" of downtime, which is not the same as increased latency due to sub-optimal routing.