5 ms·
I'm curious to hear if anyone's multi-az setup (RDS, ECS, etc) handled this event without much of an issue? I assume so but would be nice to know its working a
by plasma 5y ago
I'm curious to hear if anyone's multi-az setup (RDS, ECS, etc) handled this event without much of an issue?
I assume so but would be nice to know its working as expected!
- adrianpike 5y agoA handful of AZ-less managed AWS services were having a bad time(tm) during the peak of the incident, so if anyone was reliant on those they were having a bad time as well.
- clutchdude 5y agoIt affected services we have that are not located in the impacted AZ.
- gkop 5y agoLikewise, our RDS replica experienced degenerate lag, when neither the primary nor replica were in the scorched AZ.
- plasma 5y agoThat's really interesting. I wonder whether its because a large load of capacity needs suddenly in unaffected AZs put abnormal stress on things like networking to the whole unaffected AZs. My only experience there is in Azure, when they deployed patches for the Heartbleed etc issues, for a few weeks things were much slower in CPU power (our response times just shot up 20% for no reason then recovered a few weeks later) and there were network related timeouts that were abnormal and it all settled down eventually.
- xmodem 5y agoHeartbleed? I assume you mean Spectre and Meltdown
- celticninja 5y agoYour comment is not clear, were you unaware of heartbleed? Or were you questioning if the GP was remembering correctly? For reference https://heartbleed.com/ https://heartbleed.com/
- chrisandchris 5y agoIt is clear. Heartbleed does not affect CPU like Spectre/Meltdown did (both migitations increased CPU usage quite a lot).
- TheP1000 5y agoKinesis data streams and firehouse were down. All but one out of 300 RDSs failed over. One was down for 5 hours.
- plasma 5y agoAre you able to find out why it didn’t work? Very surprising.
- mercora 5y agonote that every AZ assignment is mapped randomly for each account [0]; "To ensure that resources are distributed across the Availability Zones for a Region, we independently map Availability Zones to names for each account.". my eu-central-1a is not necessarily yours. [0] https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids.html https://docs.aws.amazon.com/ram/latest/userguide/working-wit...
- deleted 5y ago[deleted]
- jen20 5y agoThis is true, but there is a consistent identifier cross account (likely to make it possible to build multi-account architectures without having to measure network latency between zones manually) - this is the AZ ID instead of the AZ Name [1]. [1]: https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids.html https://docs.aws.amazon.com/ram/latest/userguide/working-wit...
- ununoctium87 5y agoSome instances of our services went down but our deployments are multi-AZ by default so minimal perceptible outage to our customers
- tomw1808 5y agoWe had a few interesting "bugs" appear: Mostly logging went down, but EC2 machines kept running. E.g. We are running a lot in ECS with EC2 machines, like a RabbitMQ cluster with 20 instances all across all AZs. None of the machines died, none of the containers had degraded performance as far as I can tell, but the performance logs just disappear from 21:50 to 22:30 GMT+2. It's just blank int he Metrics and Cloudwatch dashboard. Same with other cloudwatch logs. We also do have some EC2s running with nodejs applications and there the aws-sdk just errored out with "UnknownError: 503" and simply stopped logging until we restarted the machines. The machines itself were not stopped at all. Other than that, I can't see any effects across our accounts. Also not RDS or anything else. Fascinating. Glad its under control and seemingly nobody died or so.
- tetha 5y agoAs said in another comment, we had a dozen instances or two affected. Most of the hashi stack just lost a node and chugged along at reduced redundancy. A patroni/postgres cluster lost a replica, but automatically re-integrated it into the cluster. Very nice and smooth. We have mostly found one or two classes of jobs in the orchestration for which nomad stopped retrying deployments before the ec2 instances running the allocations were fully failed and removed from the cluster - and our on-call was unsure how to handle that situation in nomad right. Network and routing were really weird at some point. Additionally, we ended up with a couple of container instances orphaned from the container management, which was strange for a moment. This was made a bit more hectic over here because a second hoster apparently fried their own network at the same time so we needed some time to realize we have two issues. Overall, 5/5 Outage, would fail again once we've updated our jobs. We're happily close to not caring about such an incident.
- kawsper 5y ago> Most of the hashi stack just lost a node and chugged along at reduced redundancy. Amazon had released a version of their AWS Linux edition that rebooted randomly due to a kernel bug, and I was working on our staging cluster, but I didn't even notice that I had EC2 instances that randomly rebooted and dropped because Nomad just kept the workload up.
- uniformlyrandom 5y agoLost one node in a small ElasicSearch cluster. The incident did not affect the cluster availability. The node was reintegrated into the cluster once it came back online.