3 ms·
This is very different from my experience. In my years with AWS I’ve only had an instance get stopped once for a reason that was weird AWS background stuff that
by NineStarPoint 3y ago
This is very different from my experience. In my years with AWS I’ve only had an instance get stopped once for a reason that was weird AWS background stuff that had nothing to do with my application. I don’t think I’ve ever had or even heard of an instance just disappearing.
- vel0city 3y agoBy "disappear" I mean the instance failed hard and couldn't be restarted. It's just gone. Usually related to the EBS volume dying. But yeah, usually when they die they can just be relaunched. Still they die way more often on AWS than in GCP, and will just end up staying stopped. Until very recently they couldn't even migrate the instances when the underlying hardware had some maintenance, you had to stop and relaunch it on your own. FFS most decent hypervisors have had live migrations for decades and yet I still get notifications of "this instance will stop on x day..." emails. I should never see that. The cloud provider should keep the instance running forever. There's no excuse.
- yolovoe 3y agoI don’t know why you’re getting downvotes. What you’re saying sounds true to me, and I work in the core of EC2. I am guessing you’re using newer instance types if their reliability is still questionable. Or you have a huge fleet of instances so you see a steady rate of failures every year. Our failure rate on the commonly used instance types if fairly low. We have several types of failures and in some bad failure cases, live migration isn’t possible and your instance won’t even be restarted. AWS already asks people to expect failures and plan around this with multi AZ deployments. If you want stability, sign an NDA with AWS and ask for fleet wide reliability metrics for various instance types. There’s a surprisingly huge variance.
- berniedurfee 3y agoSame. 12+ years of using AWS and there’s been 1 instance of a server (RDS) going down due to something outside of our control. Restoring a snapshot got us back running quickly. If we were multi-az, we probably wouldn’t have noticed.