12 ms·
Ongoing Incident in Google Cloud
- lokl 4y agomail.google.com showed error messages for me intermittently during the past hour.
- asicsp 4y agodiscussed here: https://news.ycombinator.com/item?id=34955906 https://news.ycombinator.com/item?id=34955906
- jakedata 4y agohttps://www.google.com/appsstatus/dashboard/incidents/5ML14kN7xvCATksYHh5w https://www.google.com/appsstatus/dashboard/incidents/5ML14k... They claim the Gmail specific issues are resolved. We shall see... Feb 27, 2023 2:03 PM UTC We experienced a brief network outage with package loss, impacting a number of workspace services. The impact is over. We are investigating and monitoring.
- fastest963 4y agoThis affected us starting at 4:57am US/Pacific with a significant drop in traffic through the HTTPS Global Load Balancer across all regions and Pub/Sub 502 errors but there was nothing on the status page for another 45 minutes. Things returned to normal by 5:05am from what I can tell.
- dixie_land 4y agoYup we saw the exact same symptoms with some GCLBs getting 100% 502 ( our upstream QPS graph looks scary with 5 mins of 0 QPS )
- 0x0000000 4y agoOutages at the hyperscalers can have a huge blast radius, is anyone encountering other services with outages because they're built on GCP?
- hellcow 4y agoWe are in us-central1 and didn't have an outage, so it appears not to have affected everyone.
- pictur 4y ago[flagged]
- andsoitis 4y agoSolving this sort of thing is not about throwing more people at it. That would be brute force and not strategic. Instead, you want to architect systems like these in a way that strikes a good balance between resilience and things like cost/efficiency/etc.
- bushbaba 4y agoThis demonstrates yet again why global configurations, global services, and global anycast VIP routing should be considered an anti pattern. gcp should be designed in a way where the term “global outage” isn’t a word in their vocabulary.
- consumer451 4y agoMy knowledge level: can use AWS console to do < 5% of what is possible. How much more work would Google create for themselves if they had not globalized their stack? Are we talking something like 5 subsets to manage instead of 1?
- NineStarPoint 4y agoAssuming good automation, most of the work comes in being able to do a second of something instead of just having one. The difference in work between “single point” and “multiple point” is a lot, but increasing the multiple points beyond that isn’t too bad. Of course, if you deploy a change to all of your separated stacks at once through some sort of automated pipeline it doesn’t matter too much. Easy to break everything simultaneously that way if there’s some difference between test and prod you didn’t realize was there.
- jjoonathan 4y agoMy biggest AWS surprise bill (so far!) was due to a bug in AWS console region switching.
- singron 4y agoMost of it is cellular or regional, but there are a few critical global services. The global network load balancing, network qos, and ddos prevention are more functional because they are global (i.e. you couldn't replace them with equivalent regional versions), but are often causes of issues like this. There was a push a few years ago to ensure global services had at least 99.999% uptime or make them regional. This was a 48 minute outage, so it blows that five 9 budget for 9 years. Ex-googler, no particular knowledge of this event, information might be out of date.
- typaty 4y agohttps://packages.cloud.google.com/apt/doc/apt-key.gpg https://packages.cloud.google.com/apt/doc/apt-key.gpg Even the public apt key for signing Google's cloud packages is unavailable (returns 500 for me). This is insane
- nicholasklem 4y agoThis key was 500 some hours before the incident started, I hope it's unrelated.
- typaty 4y agoinb4 it turns out an intern was tasked with updating the apt key, which brought a cascading outage of all their services
- roseway4 4y agoDownloading the key has been erroring since at least ~5pm PT yesterday, 2/27. It’s likely unrelated. Though I’d be unsurprised if the recent layoffs contributed to the situation.
- chedabob 4y agoCurrently being tracked here: https://github.com/GoogleCloudPlatform/gcsfuse/issues/961 https://github.com/GoogleCloudPlatform/gcsfuse/issues/961
- deleted 4y ago[deleted]
- eik3_de 4y agoGoogle cloud bugtracker bug: https://issuetracker.google.com/issues/270782614?pli=1 https://issuetracker.google.com/issues/270782614?pli=1
- m00dy 4y agoOur workloads are fully functional, DK/EU
- Dave3of5 4y agoOuch some pain at google today then. I hate to wake up on a Monday morning to this. <3 To the engineers trying to fix it at the moment.
- zamnos 4y agoGoogle has follows-the-sun on-call rotations for large rotations, so this hit the UK team just after lunch.
- bongobingo1 4y agoAh so the rotation rotates to match the current rotation. Very smart.
- LewisVerstappen 4y agoThe sun never sets on the Google empire
- Leszek 4y agoI like the mental image of this being a very precise matching -- as the sun traces across the sky, the responsibility of on-call passes from desk to desk, town to town, country to country; two engineers on a boat in the Atlantic race to keep up with their rotation...
- opportune 4y agoThe logical conclusion is SREpiercer
- deleted 4y ago[deleted]
- abc20230215 4y ago[flagged]
- uniformlyrandom 4y ago05:41 - 06:26 PT, 45 min total. Not great, not terrible.
- throwaway892238 4y agoYep. Of course there's no detail yet so we don't know what exactly was affected. All we can see is "Multiple services are being impacted globally" and a list of services (Build, Firestore, Container Registry, BigQuery, Bigtable, Networking, Pub/Sub, Storage, Compute Engine, Identity and Access Management) but there's no indication of what specifically was impacted. Could you still see status for your VMs, but not launch new ones? Was it mostly affecting only a couple regions? No idea. All we know is they're now below four nines in February for a handful of critical services. Let's take a gander at incident history: https://status.cloud.google.com/summary https://status.cloud.google.com/summary Cloud Build looks bad... three multi-hour incidents this year, four in fall/winter last year. Cloud Developer Tools have had four multi-hour incidents this year, many last fall/winter. Cloud Firestore looks abysmal... Six multi-hour incidents this year, one of them 23 hours. Cloud App Engine had three multi-hour incidents this year, many in fall/winter last year. BigQuery had three multi-hour incidents this year, many in fall/winter last year. Cloud Console had five multi-hour incidents this year, many in fall/winter last year. (And from my personal experience, their console blows pretty much all the time) Cloud Networking has had nine incidents this year, one of them was eight days long. What the fuck. Compute Engine has had five multi-hour incidents this year, many last fall/winter. GKE had 3 incidents this year, multiple the past winter. Can somebody do a comparison to AWS? This seems shitty but maybe it's par for the course?
- Rebelgecko 4y agoIt's weird, I did a cursory search and can't find people complaining about that 8 day long networking issue. I wonder if the latency was just barely out of SLO so people didn't notice? Or since it was a telecom problem, maybe it was part of one of the recent undersea cable outages so people weren't surprised enough to remark on it? Or maybe I'm just not searching well. (full disclosure, work at Google but not on cloud stuff)
- 4y ago
- monero-xmr 4y agoThis is why any criticism of AWS reliability is meaningless to me. All the cloud providers go down - all of them. Either you are multi-cloud, or you run your own hardware, but these events are inevitable.
- dymk 4y agoInevitable != immune to criticism
- MuffinFlavored 4y ago> you run your own hardware in multiple datacenters?
- yjftsjthsd-h 4y agoThis is why any criticism of AWS > reliability is meaningless to me. Er, we absolutely can and should compare rates of problems and overall reliability.
- ctvo 4y ago> This is why any criticism of AWS reliability is meaningless to me. Is anyone tracking reliability for these public providers? Would be curious how AWS compares to Azure and GCP. My experience is it's better, but we may have avoided Kinesis or whatever that keeps going down.
- WaxProlix 4y agoThere's Cloudharmony, https://cloudharmony.com/status https://cloudharmony.com/status
- vhiremath4 4y agoThe amount of time you are down vs. up dictates your SLOs and SLAs. Criticism of how reliable one vs. another is is not only valid, it's backed by hundreds of millions of contractual dollars and credits every year. We spend tens of millions on AWS per year. We have several SLAs with them. Our Elasticache SLA was breached once (localized to us - not whole customer base) and we got credits which were commensurate with the amount of business we lost during that downtime period. If one provider is down more than the others, the criticism is not only valid, it results in real business loss for the provider and its customers. On multi-cloud: it's one way to reduce the amount of downtime you have, but it comes with a significant operational cost depending on how your application is architected and how your teams internal to your company are formed. It is totally practical for someone to bank on AWS' reliability until they're at a significant amount of traction or revenue where the added uptime of going multicloud is worth the investment. I know you're not saying this isn't the case (I think you're saying "do that if you're going to complain about 1 providers' uptime"), but thought it was worth putting the context into the HN ether.
- Aldipower 4y agoCertainly a problem with a BGP misconfiguration. :)
- zamnos 4y agoNope. A BGP misconfiguration would manifest in more broad/different ways.
- 2OEH8eoCRo0 4y agoI fantasize that it's a three-letter agency with a warrant making them either start shitting their chat logs or pulling drives and recovering them the hard way. https://arstechnica.com/tech-policy/2023/02/us-says-google-routinely-destroyed-evidence-and-lied-about-use-of-auto-delete/ https://arstechnica.com/tech-policy/2023/02/us-says-google-r...
- JosephRedfern 4y agoSounds painful.
- kkielhofner 4y agoAs has happened many times throughout history (back to mainframes and thin clients of the 90s) there are swings/trends in how infrastructure is hosted. Listening to the “All In Podcast” yesterday even those guys were talking about revenue drops in the big cloud services and noting we’re currently in the midst of a swing back to self-hosting/co-location/whatever thinking and migrations out. IMHO those building greenfield solution today should take a hard look at whether the default approach from the last ~10 years “of course you build in $BIGCLOUD” makes sense for the application - in many cases it does not. It also has the added benefit of de-centralizing the internet a bit (even if only a little).
- knorker 4y agoRevenue drop? Google Cloud is still growing 30-40% year on year.
- rcme 4y agoAWS also had 20% revenue growth last quarter.
- knorker 4y agoYeah. Google Cloud was the only one I knew off the top of my head. I'm sure other clouds (not just MSFT) are growing too. Maybe the second derivative is going down, but it's absolutely not even close to a drop. Public cloud will grow a lot more. I'd expect a slowdown when they're all ~10x what they are now.
- deleted 4y ago[deleted]
- camhart 4y agoAll in podcast mentioned growth slowing, but not revenue dropping.
- lolinder 4y agoAs others have mentioned, there was no revenue drop, there's been a reduction in growth. AWS's 20% growth rate is still very respectable, more than double the 9% growth rate the company had overall. I would be hesitant to attribute slowed growth to a return to self hosting, it's much more likely that it's caused by companies dialing back their cloud growth after spending a few years going ham digitizing everything during the pandemic.
- lee101 4y ago[dead]
- oars 4y agoIs it likely this outage still would've have occurred even without their 12,000 layoffs in January?
- deleted 4y ago[deleted]