5 ms·
As far as I remember, previous failure which happen with Amazon earlier this year have also affected Multi-AZ deployments too. Anyway, I don't think that we ar
by akhkharu 14y ago
As far as I remember, previous failure which happen with Amazon earlier this year have also affected Multi-AZ deployments too.
Anyway, I don't think that we are ready to invest large amount of money on Multi-AZ deployments to the doubtful reliability. Cloud solutions even with single AZ should not loss data.
- nevinera 14y ago>Cloud solutions even with single AZ should not loss data. You mean you think all cloud db solutions should implement replication for you? There aren't very many backup solutions that never lose any data.
- akhkharu 14y agoNo, I understand that replication is the double cost. I meant cloud solution should not have storage failures which causes data loss.
- kiallmacinnes 14y agoI disagree. Amazon simply gave you the choice over what level of reliability you require. Would you have preferred they only offered Multi-AZ databases? (and in the process doubled your development, staging and QA environment costs..)
- nevinera 14y agoAnd google should give us all ponies. Would you mind defending your expectation that other people and companies will give you extra services for free? Amazon has been very clear on the expected failure rate of ebs volumes and the attendant rds failures. If you want data safety, multi-az deployment offers it.
- dspillett 14y agoA lot of people do (incorrectly) assume this sort of thing, even in cases when it would only take cursory understanding/research to spot the limitations a given service has and where it might fail in a disaster recovery situation. No matter how good you think a given "cloud" solution is, no matter how much too big to fail you thing the company responsible is, you should make use of the redundancy options they provide and make sure you have reasonable backups elsewhere too (local to you or on another completely separate remote service). This is what made me dismiss Google's App Engine when it first turned up (I've not looked for some time, they may have addressed this concern long ago): there was no easy way to backup all your data to another location/service and the not-so-easy ways would probably all end up costing a fair bit in bandwidth charges.
- seanp2k2 14y agoThis kinda highlights the problem with "cloud"...many people, even engineers, don't really understand what they're getting into. At least with a single server, you know what you're getting, and you have only yourself to blame if you didn't plan for a typical, known, documented failure mode.
- acdha 14y ago… and magically do it at no extra charge, too. Some learning experiences are in order.
- acdha 14y ago> As far as I remember, previous failure which happen with Amazon earlier this year have also affected Multi-AZ deployments too. Which failure? The networking issue which had nothing to do with RDS and left your data unaffected? > Anyway, I don't think that we are ready to invest large amount of money on Multi-AZ deployments to the doubtful reliability. Cloud solutions even with single AZ should not loss data. Any server can go down. The very modest increase for a multi-AZ setup buys you real, meaningful improvements as you just learned. I'm sorry that you had to learn a lesson the hard way but there's a reason why AWS recommends a multi-AZ deployment for failover and it's not revenue. The next step up would require you having multiple widely separated servers, which is where you really start talking about large amounts of money because you're talking about non-trivial engineering and taking on the operational overhead of 24x7 support.
- akhkharu 14y agoYes, networking issue which brought down even multi-az deployments. I don't want to setup highly available fault tolerant systems, I just want a good level of reliability of a service provider. Probably, we will migrate to the dedicated servers out of Amazon soon. It will be harder to maintain, but cheaper and, as practice shows, more reliable.
- ceejayoz 14y ago"Brought down" and "Brought down and lost data" are very different severities. Many businesses can handle occasional downtime as long as data's not disappearing into the ether.
- acdha 14y ago> Yes, networking issue which brought down even multi-az deployments. You might want to learn more about this before making business decisions on it. RDS was completely unaffected, as were all of my EC2 servers. They didn't receive any traffic from the internet but the systems were running fine throughout the brief outage interval. A quick Google search will reveal that this is not uncommon for any hosting setup - data centers have lost network connections, routers can fail or be misconfigured, etc. - which is why anyone with serious uptime requirements has multiple widely separated data centers. Using AWS doesn't magically remove the need to avoid single points of failure in your system design. > Probably, we will migrate to the dedicated servers out of Amazon soon. It will be harder to maintain, but cheaper and, as practice shows, more reliable. I hope you have a good ops team and extra engineering resources; otherwise you'll learn very quickly that dedicated servers have the same failure modes. So far we're at ~18 minutes of AWS downtime this year - that's not going to be easy to beat.