9 ms·
The reliability pillar of the AWS Well-Architected Framework [pdf]
- crankylinuxuser 7y ago(erased)
- klodolph 7y agoYep. Amazon (GCP, Azure) are making bank off the idea that you can just pay opex to run in multiple regions and lay off your SREs, but if you really wanted high availability (and not just some service credits when things go wrong) you would spend the engineering effort to go multi-provider. At that point on-prem looks better, but at least with multi-provider you’re no longer in the business of ordering hardware and power. The idea that failures between different systems are uncorrelated doesn’t even work at a hardware level (e.g. two different sticks of RAM), it’s pure fantasy when you’re talking about software stacks, especially when you have things like the blow-up-the-world button called “BGP.” I’m deeply uncomfortable reading formulas for availability that add or subtract nines, if there’s real money on the line, you can afford to do some better math. Disclosure: Work at a cloud provider, opinions are my own.
- cheeze 7y agoSomething that often goes ignored is that in high availability situations like that, humans are the biggest risk by a huge margin. Computers (almost always) don't accidentally wipe out the prod database thinking it was beta, that's usually a human that made a simple mistake. Dumping money into IaaS while neglecting to recognize that high availability requires extreme quality of both engineering, as well as testing/qa is a super common mistake that's being made these days. This means slowing down, which most upper management doesn't want to do. Things like executing DR plans regularly to test that they wofk are extremely costly to the business,and are one of those "IT" things that you hopefully never have to actually do in production. But without things like that, what's the point of 3xing the cost of your fleet and adding complexity (another great vector to cause failure) for "availability"? I think the answer often comes down tok the fact that people are hard to hire. Computers are easy to provision.
- FigmentEngine 7y agoYes, humans are really important in these situations, thats why "Operational Excellence" is the first pillar of AWS Well-Architected - you can only get so far in terms of reliability and security if you don't consider people and process. https://d1.awsstatic.com/whitepapers/architecture/AWS-Operational-Excellence-Pillar.pdf https://d1.awsstatic.com/whitepapers/architecture/AWS-Operat... (I work on the AWS Well-Architected team)
- crankylinuxuser 7y ago(erased)
- FigmentEngine 7y agoI can't say for your employeer, but at AWS this is not our intent. Our job is to help customers build better architectures, and if you read the framework you will see its agnostic of vendor. You can't approach reliability with the "the sky is falling down, so theres no point" approach. You are actually going to have think about component failure, blast radius and what happens afterwards a failure. rather than spreading FUD, why not help customers by showing them the better math?
- klodolph 7y ago> I can't say for your employeer, but at AWS this is not our intent. Lost the antecedent, here. What is not your intent? Granted my understanding of cloud economics is somewhat murky. My general impression is that the big selling point of cloud is the move from capex to opex, and the second selling point is that you don’t need the same level of operations expertise to run cloud compared to on-prem. My hot take is that for high reliability you still need tons of operations expertise, and in these scenarios, solutions like (partial) on-prem and multi-provider become much more favorable. > rather than spreading FUD, why not help customers by showing them the better math? I feel like I’m being accused of spreading FUD, and I want to know why? All I am really trying to say here is that you can’t just do simple arithmetic on published #nines and end up with something that approximates the truth for your service in a useful way. Depending on the operation of your business this approximation may be acceptable or it may not be. Just to recap, the bad math is to come up with some threshold for acceptable performance in all components, model each as an independent Bernoulli variable, and then plug them into some boolean formula. This is the math published in the guide here, and you can do it with arithmetic on #nines. The reason why this is bad is because this leads you towards a shallow understanding of your system and creates pressure to inflate estimates of availability beyond actual availability, sometimes much so. Unfortunately, if you are really interested in calculating availability you need to come up with a model for your particular system. This is a complex subject that involves coming up with a model which strikes some balance between accuracy, simplicity (so it can be understood and used to inform strategy), usefulness (to customers / downstream), and supportability (to engineers working on the service). I can't tell you how to do that, my best guess is to hire someone who knows enough statistics to be dangerous and lock them in a room for three months with a terminal and access to your metrics. (As a side note, the tools for this kind of analysis are much better for systems like electronic circuits. For example, a part might be labeled as "1% tolerance" but when you are running simulations you use a probability distribution.) When you actually do this, you often find out that your system is much less available (reliable, durable) than it was designed to be. It’s common that your cloud provider could stay within SLO and your service would still be “down” as far as your definition went. So then you have the engineering problem of figuring out how to improve things, which may involve multi-region, multi-cloud, on-prem, redesigning parts of your system, changing utilization targets, etc. The model helps because it can reveal key insights like “if you improve latency in this part of the system, it improves availability in this other part of the system”, so you can decide where to spend engineering resources. I can guarantee you that the folks who are in charge of, say, EC2 are not just doing math with #nines of the systems underlying EC2. They are measuring and modeling. If I really wanted to figure out how to calculate availability of my systems, I would want to read case studies of real-world systems.
- lrem 7y agoUh, don't you need to HIRE some SREs to get that multi-everything setup engineered correctly?
- klodolph 7y agoDepends on the complexity of your service. The question is also not about what is “true” (you need to hire SREs if you want high reliability) but what is believed. Beliefs are what drive sales. Cloud is so pervasive that customer beliefs about what cloud can/cannot do are different from both reality and from marketing materials.
- tbanks22701 7y agoAre you in search of a reliable Hacking Services? Then he offer the best of hacking service with his dedicated hackers with track records. He offer various Services 1.School Grades Change 2.Drivers License 3.Provide solutions on professional exams 4.Hack email, Database hack & Facebook, Whatsapp 5.Retrieve, deleted data and recovery of messages on cell phone 6.Crediting , Money Transfer. 7. Clearing of criminal records and many others . He Provide high grades techs and hacking chips and gadgets if you are interested in Spying on anyone. Contact : guruhacks007 at gmail dot com OR info at guruhacks007 dot com WhatsApp: +1 (847) 497 0407 ICQ: 704422091
- cheeze 7y agoOf course AWS wants you to use them. They would never market for a competitor, that just makes 0 sense. In terms of reliability, AWS and GCP have solid track records in terms of multi-region availability (I can't say the same about Azure... They seem to be years behind still.) Cost absolutely goes up when you're building out regional fault tolerance, but that's true whether you use a mixture of cloud providers or any specific provider. If you need high availability, you need to waste some level of resources (and in turn, money), it's part of the design. I'd argue almost anyone would be fine running out of a single aws region. There are specific cases where extreme availability matters, but in that case you're making the tradeoff of cost and simplicity for reliability. It would be interesting to see something like a study on whether using multiple regions in one cloud provider makes more sense over having your load spread across providers in terms of complexity versus fault tolerance.
- FigmentEngine 7y agoThis is not the intent at all, I work in the AWS Well-Architected team, and we are engineers trying to help other engineers. If you follow the best practices you will spend less (there are performance and cost optimization pillars).
- crankylinuxuser 7y ago(erased)
- privateSFacct 7y agoI think the way the forward looking CFO positions approach this (that keeps development velocity high without the 20 layers of paperwork to get CFO approval in old system) is as follows. New feature coming online, team sits with cost/acctg side and says, we expect our budget needs to go up by $X / month. Either third party or now AWS budgets for a given tag/project/account are prepared daily. If the daily / hourly rate > expected, inquiries as to why are made, discipline if needed (rare). Many CFO's are so happy to be out of the CapEX game, out of the Oracle audit game they will probably put up with a fair bit of a tradeoff there. Those same CFO's have had a lot of trouble when a project doesn't work out under old approach. Now they just spin down everything in AWS over a few hours. Govt side in particular, the datacenter buildout costs are CRAZY and the utilization often terrible.
- klodolph 7y agoI think this really hits it on the nose. CapEx is “cheap” long-run but often insanely difficult. It only makes sense if you have the scale (1) to control variance (2) to amortize planning costs (3) and it provides a competitive advantage. Meanwhile sysadmins with no knowledge of accounting will complain that on-prem is cheaper. This is why cloud wins.
- FigmentEngine 7y agoThanks for the clarification, I agree we can do more around cost, we do offer advice here that should help https://d1.awsstatic.com/whitepapers/architecture/AWS-Cost-Optimization-Pillar.pdf https://d1.awsstatic.com/whitepapers/architecture/AWS-Cost-O... As an engineer I would prefer to see a workload work well on a single vendor, before adding the complexity of running it across multiple vendors. Complexity is rarely your friend in Reliabilty or Security
- cavisne 7y agoAWS has not had a global outage in a very long time (I think there was one S3 outage that was global because at the time there was only 2 regions). The venn diagram between services that need very high availability, and yet have such low traffic that distributing resources across 2 or 3 AZ's is a waste of money, should be pretty small. Multi cloud for a individual service will never make sense as all the providers know how to price network transit to make this infeasible.
- iblaine 7y agoInteresting that security is 1 of 5 pillars and hardly gets a mention in this paper. Who wrote this and why?
- FigmentEngine 7y agoIt was written by the AWS Well-Architected team, and is based on curating best practices from Solutions Architects working with customers and engineering teams. Its one of Six main papers, with the main framework, and then one per pillar https://aws.amazon.com/architecture/well-architected/ https://aws.amazon.com/architecture/well-architected/ There is also a free tool available to review architectures. https://aws.amazon.com/well-architected-tool/ https://aws.amazon.com/well-architected-tool/ The aim is to help customers learn best practices for building and operating in the cloud. (I work in the AWS Well-Architected team)
- maxmcd 7y agohttps://d1.awsstatic.com/whitepapers/architecture/AWS-Security-Pillar.pdf https://d1.awsstatic.com/whitepapers/architecture/AWS-Securi...
- iblaine 7y agoThank you
- inlined 7y agoThe section about calculating the availability of hard and redundant dependencies ignores the fact that systems often fail in tandem. For e.g. you might have a primary and secondary database in different AZs, each with 99.9% availability. This gives read operations a hypothetical uptime of 6-9s. But hidden SPOFs like operator error, VM infra failing, load balancer or other networking outages can make the 0.01% failure time overlap, blowing away your 6-9s dependency guarantee.
- Tinned_Tuna 7y agop.14: "Because 99.999% availability provides for less than 5 minutes of downtime per year, every operation performed to the system in production will need to be automated and tested with the same level of care. With a 5 minute per year budget, human judgement and action is completely off the table for failure recovery"
- candiodari 7y agoThat just gets you back to automated systems, for example, causing their own failure, then responding to it by causing more failure. For example, a relatively common occurence is a BGP link getting saturated. Get that situation bad enough and the BGP session will go down, which will redirect all that traffic to another link with another BGP session, which then proceeds to go down. Meanwhile the original session comes back up and ... And then the failures synchronize and cause a third link to go down (each time taking more traffic with it and therefore causing failures faster). The second issue discussed is that the math in the statistics only works if the failures never synchronize. That's true for a lot of statistical analyses and mostly people ... just don't care. Yes that makes those analyses wrong. But we don't have a better way of doing those analyses.
- crankylinuxuser 7y ago(erased)
- FigmentEngine 7y agoThis is not true. AWS provides multiple constucts to help isolate failures, such as Availability Zones that provide seperate locations, and Regions that are designed to be completely isolated from each other. You can also use different AWS Accounts to provide logical seperation of your different workloads and limit management.
- random42 7y ago> "... For example, if a system makes use of two independent components, each with an availability of 99.9%, the resulting system availability is >99.999% ... " This does not seem correct.
- necro 7y agoI suppose it depends if the use/need is in serial or in parallel.
- alextheparrot 7y agoIndependent and redundant components is the missing context the quote was taken from, wherein the quote is correct.
- maltalex 7y agoWhat they probably mean is that if each component can fail independently with a probability of 0.001 (0.1%) then the probability of both of them failing is 0.001 * 0.001 = 0.000001 (0.0001%) If the system depends on just one of the components working then 1 - 0.000001 = 0.999999 (99.9999%)
- pkolaczk 7y agoThis is correct, however in reality completely independent components are very rare. Even things that seem independent and truly redundant e.g. jet engines of an airliner, are much more likely to fail after one of them fails. Therefore this line of reasoning must be applied with extreme care.
- Jedi72 7y agoThis is total submarine marketing by AWS. Im sure some very smart people with good intentions at AWS produced this, but the fact is they're not an independent academic providing objectively "the best" practices, they're a company trying to sell you something. They probably in full belief in their own abilities believe the AWS way of doing things is best, but frankly, I know as much as you do, or at least if you want to win me over you need to comprehensively prove it, not just preach from the hill. TL:DR stop telling me you're doing better architectures in the same of selling me something
- sciurus 7y agoI don't look at the well-architected framework as "the AWS way is the best way of doing things" but rather "this is the best way of doing things on AWS".