38 ms·
Just wanted to add a quick note before we get the usual deluge of "you should be running in multiple AZs and regions" posts: These outages are relatively rare a
by matt2000 7y ago
Just wanted to add a quick note before we get the usual deluge of "you should be running in multiple AZs and regions" posts: These outages are relatively rare and your best decision might just be to accept the tiny amount of downtime and keep your app simple and inexpensive to run.
I of course don't know the tradeoffs involved in running your system, but I know for a lot of my situations the simplicity of single AZ with a straightforward failover option is usually the right tradeoff.
- jen_h 7y agoAn occasional outage is sometimes good for an app (depending, of course, on how mission-critical it is): 1. People don't realize how much they love and depend on you until you're gone. 2. Keeps you on your toes, it's easy to get complacent when everything just runs along happily for months and years on end. I do wish there was a way to train users that millions of them reloading constantly as service ramps back up doesn't accelerate the ramp-up time, though. ;)
- cle 7y agoYes! If you haven’t had an operational issue with a service in a while, you should force one on a testing to stack to make sure it fails like you expect, you can recover gracefully, your monitors work, etc. (lots of folks call these “game days”). A service that hasn’t had an ops issue in a long time is a ticking time bomb: when it does fail, nobody will be very familiar with it, and environmental assumptions might have changed, which could result in taking a much longer time to recover. When I design services these days, I try to design them so these failure scenarios are constantly exercised. Eg if I care about multi-AZ resiliency, I try to design it so that it’s forced to fail over to other AZs all the time. Or at the least, write tests for the scenario. Exceptional behavior or code paths are dangerous.
- bob1029 7y agoFor us this is exactly the correct approach. We could have spent millions of dollars and thousands of man hours hardening things to be resilient to single region outages. But for what? We aren't GE or Google. If our conference line goes down for 2-3 hours per year because we don't have apocalypse-proof infrastructure, literally nothing bad happens to our business. In this exact outage we are discussing, all of my coworkers are at home having breakfast with their families and doing various weekend activities. No one but me will know there was even a problem until I log into the AWS console and find the alerts. Worst case, I have to reboot or restore a few affected instances on Tuesday morning. It seems like a lot of businesses are chasing this ideal of perfect and end up much worse off than if they had just stuck whatever application on a single server in a semi-reliable part of the world.
- Johnny555 7y agoIf your business is so complicated and mature that it takes millions of dollars to build multi-region tolerance, you probably need that tolerance. For more simpler sites, having multi-region failover (even if it's a manual failover and you lose a few in-flight transactions) is much easier to build.
- hhw 7y ago2-3 hours per year is a lot of downtime. Most competent bare metal providers see maybe one major outage of less than an hour every 3-5 years. Nothing other than a facility wide power outage, if the load somehow gets dropped because the generators don't start right away as they should, or a misbehaving (only partially failing) core network infrastructure device should result in major outages when all the proper redundancies are in place. Specific providers aside, there's more complexity involved in a large cloud provider's infrastructure and much more that can go wrong as a result. Having a code update, or some orchestration issue from your infrastructure provider be potential points of major outages are huge and unnecessary risks. You don't need that much scale, just utilizing enough resources to fill up a few whole physical machines for a few hundred dollars a month. Add some globally distributed BGP Anycast DNS and database replication and you have enough redundancy to withstand most of the worst major infrastructure failures. I would understand if AWS was super simple and convenient, but these days the learning curve seems far greater than setting up the above described bare metal solution. While being almost an order of magnitude more expensive for the equivalent amount of resources. How did we end up here? Does brand recognition just trump all technical and economic factors, or what am I missing? Disclaimer: I run a bare metal hosting provider
- SPascareli13 7y agoI don't want to advertise your particular company, but if we are talking about numbers, how does your bare metal offer compares to a Amazon ec2 offer for example? And how would a customer that need to scale their load do it?
- hhw 7y ago
- manojlds 7y agoBeing on AWS is also easy to explain to customers about downtimes - AWS was down and customers are pretty understanding in that case and don't demand why you aren't multi AZ etc ( of course YMMV based on sensitivity of your business)
- jen_h 7y agoYep. I call it the "AWS Chicken Pox Party." We got lucky this time, RDS, ELB, EC2 and Lightsail instances all in US-East-1 across multiple accounts and no issues (knock-on-wood and understand that we'll get it the next time). Especially happy as I had a two-day running neural net training job running and it's still going, that would have been depressing. Phew.
- kahnjw 7y agoWhy don't you just checkpoint the model every n steps? NNs fail for a myriad of reasons, you can easily reduce risk by routinely saving state.
- jen_h 7y agoAfter my last oom-party, I now have it checkpointing every 1000 steps (way too often, I think, but there's plenty of disk), but I just really really want it to complete a full run. ;)
- freehunter 7y agoKind of like the "no one ever got fired for recommending IBM", if you have significant downtime on Linode or Hetzner, people are going to ask "why weren't you on AWS?!". If you're on AWS and AWS goes down, you get to skip that question entirely. You were already using the logical choice, you don't need to defend anything. If Netflix can go down when AWS goes down, so can your app. AWS outages impact so much of the Internet that people will just accept it. Of course here I am with a site on Heroku which uses AWS... impacted by the AWS outage... fielding questions about why I didn't pick AWS if Heroku suffers outages like this. Can't please them all.
- bdcravens 7y agoIndeed. Our company's core business occurs in batch, background processing, that needn't be real time. If it doesn't run right now, there's literally no damage to the business or our customers if it runs in an hour. We have a customer-facing website, but there's very little there that can't be served by a cache. tbh, I can't recall a support request from a customer that was caused by our infrastructure vendor and not our product.
- acdha 7y agoThis should always be a business calculation but I do think you should note that there’s at least an order of magnitude difficulty increase between multiple AZs and regions, especially if you’re using services like RDS where it’s designed in, so I’d consider that a solid bridge step. I’m trying to put some numbers into that, I’ve been running a relatively well trafficked website in multiple AZs since 2011. We had ~20 minutes of downtime when they had a network routing issue for us-east-1 and a few hours of degraded service when S3 had a region-wide outage. I haven’t added up the number of single AZ outages during that period but based on the RSS feeds I think it’s a good bit more relative to the very modest additional cost.
- matt2000 7y agoGood point, the multi-AZ RDS feature is a nice way to get most of the resilience upsides without any additional app complexity. You do double your database cost, but that might be worth it.
- runamok 7y agoNot necessarily. You could keep the reader a smaller size and scale it up only if needed.
- acdha 7y agoThat’s assuming that you a) religiously test with the smaller size and b) are comfortable that scaling up will work when lots of other people are shifting workloads, too. I usually work on projects where we haven’t wanted to deal with that but that’s a judgement call.
- liveoneggs 7y ago"dual az" is a checkbox that doubles your cost for transparent failover; it's different from the read only replica
- js2 7y agoMaybe just because it's been around the longest, but my impression is that us-east-1 seems to have more than its fair share of outages. Personally, for my single-region applications focusing on US customers, I go with us-east-2. Knock on wood. I'm interested in any evidence to back up my impression if anyone has bothered to do the proper data gathering. (Aside, stink eye on whoever made a breaking change over a holiday weekend, if this turns out not to be random.)
- crgwbr 7y agoIt’s because us-east-1 is the oldest and by far the largest of any AWS region. Issues get caught and fixed there before they show up at other regions.
- deleted 7y ago[deleted]
- tus88 7y agoOr how about no AZ where possible using a serverless architecture/lambda?
- scarface74 7y ago“Serverless” is not magic. The minute you need to attach to your VPC (I refuse to say “run inside your VPC, that’s not correct), you still have to worry about multi AZ. If you are using RDS, as opposed to DynamoDB, you have to configure it for multi AZ. I use lambda all of the time and I’m definitely not afraid of the “lock in” boogeyman, but I always architect my lambda’s to make moving away from lambda to either Fargate or just an EC2 instance as easy as possible. On another note, it’s just as easy to architect your regular old EC2 instances running stateless servers to be AZ failure resilient. Just set up an autoscaling group with a min/max of 1 and configure it to work across multiple AZ’s. Also lambda comes with its own set of limitation - maximum runtimes of 15 minutes, cold start times, temporary storage space of only a half a gig, limited CPU/memory options, etc.
- austinshea 7y agoI can understand if the networking implications and data replication issues are too complicated for you, but, if you have failover in the same region, you’re already paying the extra cost. Working with multiple regions is cost friendly on AWS. You should put in the time and learn how that stuff is done, it’s not as complicated as you think.