4 ms·
agree. like you imply, too many people rely heavily on their providers for business critical things like backups and redundancy. while generally big providers t
by vectorEQ 7y ago
agree.
like you imply, too many people rely heavily on their providers for business critical things like backups and redundancy. while generally big providers to a good job at this (and aws certainly does a good job at this), it does not mean there is a guarantee of any kind failures won't ever occur. thus the need to heavily invest in failure resistant technologies upon this borrowed infrastructure is arguably more important than picking 'the best' provider for the job.
I would like to note that the response time suggested by the tweet is a bit bad, 4 days to realise something and send out response / alerts to customers is a bit slow even for amazon. But then again, everything goes slower for bigger things, and amazon is quite big i'd say. not sure what the SLA response time to such an incident is, so it might be within agreed times...
- romaaeterna 7y agoBut the reason I pay AWS is so that I don't have to hire a team to take care of backups and redundancy on my side. If they can't be relied on, a lot of the justification for their cost markup goes out the window.
- cloakandswagger 7y agoIt can be relied on, but it's up to you to configure it properly. AWS has no way of knowing how critical your application is and what level of redundancy it needs, and this has cost implications so they can't do it automatically.
- romaaeterna 7y agoI'm pretty sure that RDS is EBS-backed.
- acdha 7y ago… and RDS has a multi-AZ checkbox which does exactly what it claims. Anyone who used it did not have a problem with this outage.
- kevan 7y agoIf you don't want to think about things like redundancy then use higher-abstraction services. Lambda for example takes care of multi-AZ redundancy so you don't have to think about it. The lower level building blocks like EC2 don't. They expose the fault boundaries so that you can build HA applications on top of them, but it's still your responsibility to do so.
- thruhiker 7y agoRespectfully, that's not a good reason to use public cloud providers like AWS. They provide features and tooling that make building redundancy into your services easier but for many of these redundancy features you must integrate them into your infrastructure design to take advantage of them.
- bdcravens 7y agoAWS gives you access to redundant resources inexpensively. If you have your application in a single AZ, you’ve elected to bypass the redundancy.
- skywhopper 7y agoIt sounds like you might misunderstand the product you are buying from them. They are very clear about the reliability of EBS (1 in 1000 volumes will fail during a year of uptime), and they provide a really easy way to back things up, and there are tools available to schedule automated backup rotation. So I'm not sure what more you expect. AWS can't possibly know what your needs are for backup and restore for a particular EBS volume. If you want data durability, use S3.
- tnolet 7y agoIf you replace AWS with Heroku in your statement, I agree. Heroku abstracts the redundant AWS resources for you so you can just “run your app”. However, Heroku also had a huge outage. That is way more problematic as far as I am concerned.
- blihp 7y agoAWS is only selling you infrastructure as a service, not a turnkey solution. It's up to you to combine and coordinate these services into a solution that delivers the capabilities (including backup, recovery and fault tolerance) appropriate to your needs. So while you don't need to hire a team to take care of backups and redundancy, you do need to provision and configure what is required so that their team can.
- skywhopper 7y agoI suspect the four days was how long it took to confirm that their particular EBS volume was definitely not recoverable. Soon after the outage ended on Saturday morning it was clear that some EBS volumes were not recovering quickly, and support's advice at that time was to rebuild/recover if you needed to still be online. When a datacenter loses power like this, a few of the storage arrays will just not come back online. But another few will take time to run through their corruption recovery process, and it may take a long time and some service by a human (eg parts replacement, etc) before they can be certain a particular volume is not recoverable. Given their scale and the timing, at the beginning of a holiday weekend, four days is annoying, but not bad.