3 ms·
This year has been really bad for the cloud vendors in terms of stability.
by mangatmodi 7y ago
This year has been really bad for the cloud vendors in terms of stability.
- karma_fountain 7y agoProbably unrelated but there is also an issue with JIRA cloud based solution. https://status.atlassian.com/ https://status.atlassian.com/
- mbo 7y agoIt's related. Atlassian runs on AWS.
- StreamBright 7y agoThis is a single AZ. Your application should tolerate single AZ outages. That is the first rule of building reliable, highly available services on AWS. 12:08 AM PST We are investigating increased network connectivity errors to instances in a single Availability Zone in the EU-CENTRAL-1 Region.
- mvanbaak 7y agoTrue, but autoscaling at the moment is unable to scale in/out because of this issue. If your autoscaling group includes the affected AZ, you are out of luck because those instances are being terminated since their health checks fail. But because of the failures, autoscaling is unable to complete this, and unable to launch new instances in the other AZ's as it is stuck on the terminating part.
- StreamBright 7y agoThat is an interesting detail. Would you consider having 3 separate autoscaling groups (one per AZ) or this is not feasible for some reason? One of the interesting aspects of running services on AWS was to remove the AZ from "rotation" while the outage lasts. Meaning, having a DNS change and exclude the public endpoint from taking any traffic. If you have 3 separate groups doing 1/3 of the load and having independent DNS entries, autoscaling groups, etc. then moving traffic from one AZ to another is probably easier. Not sure about the details of your setup though, you might have reasons not to do this.
- mvanbaak 7y agoA setup with an ASG per AZ makes it a lot harder to do real autoscaling based on load/mem/connections. If it was purely for running a fixed amount of instances equally spread acros AZ's this would probably work, but not in our setup where we have unpredictable traffic and load patterns. [edit] I know it can be done with combined metrics etc, but it would make it a lot more complicated ;-)
- StreamBright 7y agoIt is most certainly more complicated. We were ok to be put up with that additional complication because service reliability (especially tolerating single AZ outages with ease) was higher on the requirements than avoiding complication. :)