3 ms·
RTO is so hard to state properly even with regular testing. If someone blows away a critical database, sure you can meet your published RTO. What if we lose 300
by mdeeks 4y ago
RTO is so hard to state properly even with regular testing. If someone blows away a critical database, sure you can meet your published RTO. What if we lose 300 of our databases and need to copy snapshots from another region. AWS limits you to 20 concurrent snapshot copies cross region. Which of those databases should you do first? Do you know your entire dependency graph for all 1000 of your services to make the right call? Meanwhile tick-tock your "6 hours" is slipping away. And what if someone nukes our entire AWS account with all of our prod resources? Databases, load balancers, S3 (no such thing as snapshotting there), EC2 instances, etc, etc
Those last two are examples are very unlikely but no company is going to say RTO = "probably 6 hours but it could be three weeks if we get ransomwared"
- manquer 4y agoFor what you suggest some combination of these things should have happened. - Some employee has root access to AWS account and uses it operationally - Given wildcard S3 permissions to an IAM user and allowing delete bucket - Not enabled object versioning - Cross Region replication not enabled - no large bucket protection - don't have basic security monitoring and setup of Cloudtail alerts - have not invested in full fledged tools for IDS and so on. If some vendor have any of these issues I don't think any customer would approve these software to be used, these are not normal or best practices . Large apps have detailed playbooks on how and what gets turned on in what order, and most do DR drills and time those runs periodically. These are well established workflows in any large org. Yes in a real world downtime you can't have planned for every scenario, maybe you miss the target by 25 % like 2 hours more, or maybe in a very situation you double or even triple it say 12-18 hours. You don't go from 6 to 600+ . The way RTO is calculated starts by looking at limits on cloud/ hardware / bandwidth/ machine sizes, if basic limits are not factored in like cross region concurrency there is no point in RTO being computed. Even if something like that was missed and you spend tens of millions of dollars on AWS then AWS will work with you and relax those limits . 100x missing the plan either means extremely poor planning or they screwed up something very very badly.
- mdeeks 4y agoI haven't ever seen it be as perfectly done as you've described. It is always shades of gray across teams and companies. They have most of what you described, but not uniformly across the company. e.g. versioning may be enabled, but not cross region replication because it is cost prohibitive. Someone runs a job to clean up a bucket that includes deleting old versions. They point it at the wrong bucket or wrong path in the bucket. Or a malicious user does it on purpose. Monitors and alerts really tell you after the fact that you now have a major problem. Also limits (like cross region concurrency) may not be known about until it is time to actually do a mass scale restore. DR tests might have been done but only in isolation of one app at a time. By the time you realize your mistake you're dealing with physics. Maybe AWS can bump it a bit to help you in that particular circumstance though. No idea what happened at Atlassian. My only point is it is very hard to get it right without a huge amount of effort.
- manquer 4y agoIt is hard to get it perfectly right yes , missing by small margins or even doubled/trippled the declared time would be reasonable if it is just prediction problem. However going like 100x is not probably cause this is hard to get 100 % accurate it look more likely deleted data as being rumoured and more importantly not actually having functioning backups that were ever tested and manually reconstructing from logs and other sources. More than just RTO, they are not going to be able to meet RPO objectives for affected customers , depending on how much loss that is going to pretty bad.
- bigiain 4y ago> 100x missing the plan either means extremely poor planning or they screwed up something very very badly. Like most airliner accidents, this is probably an unfortunate combination of both of those things happening at the same time. My guess would be they have fairly decent planning overall but there's one (or more) small-ish areas where their planning is extremely poor - which crossed over with a screwup in a very specific fashion that laser focussed on that particular piece of poor planning. The "this can never happen" immovable object and the "You can't do that" irresistible force.
- syshum 4y agoNormally for short RTO you do not try to recover by to the orginal location you fail over to warm replica;s that are already staged with data with in your RPO.
- bigiain 4y agoRTO = "probably 6 hours unless we've fucked up so badly we just cut our losses and close the entire company down"