12 ms·
For most businesses a little down time here and there is a calculated risk versus more complex infrastructure. You can’t assume all the cloud architects are idi
by dangero 4y ago
For most businesses a little down time here and there is a calculated risk versus more complex infrastructure. You can’t assume all the cloud architects are idiots — they have to report their task list and cost of infrastructure to someone who can give feedback on various options based on comparative resource requirements and risks.
Zone downtime still falls under an AWS SLA so you know about how much downtime to accept and for a lot of businesses that downtime is acceptable.
- ilikecakeandpie 4y agoThis is true, but I think it would be more acceptable if the region were down vs the single AZ
- water-your-self 4y agoIt gives me a bad gut feeling when you imply that multiple instances of a service is more complex than a single instance which cannot be duplicated easily. I also disagree that it is inherently more costly to run a service in multiple locations.
- mrits 4y agoYou should get into the database business. A lot of money to be made there if things are so trivial for you.
- willmadden 4y agoThe sounds of crickets is deafening!
- wkoszek 4y agoHow do you NOT pay more for running double of everything + load balancers?
- poxrud 4y agoYou do not need to pay double for everything, that might have been true with traditional VPS providers but it is not the way it works with cloud services. You decide on what kind of failure you're willing to tolerate and then architect based on those requirements (loss of multiple AZ's, loss of a region, etc..). Let's say your website requires 4 application servers, you can then tolerate a single AZ failure by using 5 application servers and spreading them among 5 AZs.
- deleted 4y ago[deleted]
- nemothekid 4y agoIf you already have 4 application servers you are probably already AZ tolerant; most people concerned about "doubling everything" are only running 1 instance. Going by your example, If your website requires 1 application server, to tolerate a single AZ failure, it requires you to double the number of application servers. Example - we have a service that used Kafka in the affected region that went down. Our primary kafka instance (R=3) survived but this auxiliary one failed and caused downtime. There's no way around this other than doubling the cost.
- poxrud 4y agoThat's true, when only dealing with 1 server, you technically double the cost by adding a second server. My original comment was about "popular sites/services", that should be able to tolerate the costs and are most likely dealing with multiple servers. For a single server deployment you can still reduce your downtime (with minimal costs) by having the ASG redeploy into another AZ on a failed health check.
- Nextgrid 4y agoIn most cases the elephant* in the room is your DB - it doesn't matter where your stateless application servers are, if your stateful DB goes down you're in trouble. It's also often 1) the hardest to replicate, as replication involves tradeoffs - see CAP theorem & co and 2) the most expensive, since it needs to be pretty beefy in terms of CPU, RAM and IO - all very expensive on AWS. *: https://commons.wikimedia.org/wiki/File:Postgresql_elephant.svg https://commons.wikimedia.org/wiki/File:Postgresql_elephant....
- _joel 4y agoOf course it's more costly, you need to ensure state between locations so by virtue there's more infra to pay for. It's not just a single instance too, there's generally a lot more infrastructure (db servers, app servers, logging and monitoring backends, message queues, auth servers... etc)
- ethbr0 4y agoAlso, people who can configure and maintain that infrastructure. It is more complicated, and it does require a different sort of person. (And checkbox-easy is sweeping edge cases and failure modes under the rug)
- aledalgrande 4y agoalso inter region replication costs bandwidth money
- throwbigdata 4y agoLots and lots of money.
- throwbigdata 4y agoI’m sorry about your feelings but you are wrong. its more expensive to have more things and it’s more expensive to have more complicated things that are also complex. And things that can fall over are inherently more complicated.
- boomer918 4y agoA multi-az deployment is a checkbox in most AWS services, e.g. ASGs, RDS, load balancers, etc. Someone didn't check that box because they didn't know about it, there isn't much complexity in it.
- schroeding 4y agoAren't multi-az deployments more expensive? That would be a valid reason not to check this checkbox, if your business can survive a bit of downtime here and there.
- phamilton 4y agoMost of that expense is just the cost of a hot failover, but there is some additional cost around inter-AZ data transfer. If someone is not checking the boxes for cost reasons, I would be surprised if they had failovers in the same AZ. It seems more likely they just don't have failovers.
- jenny91 4y agoA checkbox that might 3-4x the cost.
- retinaros 4y agomulti az brings multi complexity in terms of data duplication, consistency, if your app wasnt designed to handle those kind of scenarios and experience high users loads then you are in for a lot of problems. designing for those scenarios increase complexity; cost; architecture style and most of the time it will bring you in microservices territory where most of the companies lack experience and just are following best practices in a field where engineers are expensive and few
- boomer918 4y agoRDS just has a button for multi-AZ primaries. No complexity or microservices.
- 4y ago
- happymellon 4y agoConsidering almost all of the services are multi-zone, it's not hard to add in a couple of lines to make them resilient against this. People are just unaware, and probably making bad calls in the name of being "portable".
- blamarvt 4y agoIf your application and infra can magically utilize multiple zones with “a couple lines”… then I would say you are miles ahead of just about every other web company.
- happymellon 4y ago> you are miles ahead of just about every other web company. I'm curious who these web companies are. Use something like Lambda and you get multi-az for free. https://docs.aws.amazon.com/lambda/latest/dg/security-resilience.html https://docs.aws.amazon.com/lambda/latest/dg/security-resili... Dynamo is another service that wouldn't be impacted as it is multi-az. Getting postgres RDS multi-region would require the extra couple of lines in your CDK, but is fairly straightforward.
- leesalminen 4y agoToday, a SaaS I’m familiar with that runs ~10 Aurora clusters in us-east-2 with 2-3 nodes each (1 writer, 1-2 readers) in different AZs had prolonged issues. At least 1 cluster had a node on “affected” hardware (per AWS). Aurora failed to failover properly and the cluster ended up in a weird error state, requiring intervention from AWS. Could not write to the db at all. This took several hours to resolve. All that to say that it’s never straightforward. In today’s event, it was pure luck of the draw as to whether a multi-AZ Aurora cluster was going to have >60 seconds of pain. That SaaS has been running Aurora for years and has never experienced anything similar. I was very surprised when I heard the cluster was in a non-customer-fixable state and required manual intervention. I’ve shilled Aurora hard. Now I’m unsure. Thank goodness they had an enterprise support deal or who knows if they’d still have issues now.
- 4y ago
- justapassenger 4y agoThis. People working in IT naturally think keeping IT systems up 100% time is most important. And depending on the business it often is, but it all costs money. Running a business is about managing costs and risks. - Is it worth to spend 20% more on IT to keep our site up 99.99% vs 99%? - Is it worth to have 3 suppliers for every part that our business depends, with each of them being contracted to be able to supply 2x more, in case other supplier has issues? And pay a big premium for that? - Is it worth to have offices across the globe, fully staffed and trained to be able to take on any problem, in case there's big electrical outage/pandemic/etc in other part of the world? I'm not saying that some of those outages aren't results of clowny/incompetent design. But "site sometimes goes down" can be often a very valid option.
- aledalgrande 4y agoI think a good trade off, if your infra is in TF, is to be able to run your scripts with a parameterized AZ/region. That way you can reduce the downtime even more at a fraction of the cost. (assuming the services that are down are not the base layers of AWS, like the 2020 outage)
- Sebb767 4y agoIf you can get the data out of the downed AZ, don't have state you need to transfer and are not shot in the foot once the primary replica comes online again. I've rarely deployed an app where it was as easy as just to change a region variable.
- aledalgrande 4y agoYeah the data stores are the ones that I would always keep multi AZ no matter what. Everything else is stateless and can be moved quickly.
- xchaotic 4y agoWrite an article on that because you make it sound simple. Or better yet, start a company that configures this for companies.
- cromulent 4y agoYeah, makes sense if explicitly stated. Not everything is worth the money. However, in my experience, the people doing the calculations on that risk have no incentive to cover it. Their bonus has no link to the uptime and they can blame $INFRA for the lost millions and still meet their targets and get promoted / crosshired. The people who warned them and asked for funding are the ones working late and having conf calls with the true stakeholders.