5 ms·
> If you really care about high availability, you want to bring all of the points of failure under your direct control. This is what exactly I'm talking about.
by newobj 10y ago
> If you really care about high availability, you want to bring all of the points of failure under your direct control.
This is what exactly I'm talking about. People can't accept that trying to control it is not in any way guaranteed to make it more highly available. Bringing points of failure under your control makes no innate guarantee of improving anything. It can easily make it worse.
The only thing it can do is let you blame yourself of blaming S3.
You can only be hardened for what you have anticipated or experienced before.
There are tens thousands of wall clock hours of operational experience w/ S3. Availability is one of the top concerns of all AWS.
Thinking you can be more available is just fooling yourself. Believing you are more available will entail a willful ignorance or distortion of metrics.
Do your engineers carry pagers and have a <15 minute engagement time? Do your engineers sit at home when they are on-call because they know they can't simply let a page slide because they were in the middle of dinner? Or is your company more lenient than Amazon when it comes to operations?
Do your engineers spend a quarter fixing some failure mode of your infrastructure, or are they too busy working on features?
Is your team's performance measured by availability of your service? Or your actual core business?
- mattrobenolt 10y agoIronically, in our case, our availability is very core to our business. In this exact scenario, if S3 would have blocked our processing pipeline because it was down, that means we couldn't have alerted users as reliably that they were in fact, having issues because of S3 being down as well. So in our case, this is massively important to us and worth any additional risk that may be introduced.
- IAmGraydon 10y agoAre you realizing that risk spread out over time is what creates your downtime average? So what you are saying is that your uptime is so important that it is worth additional downtime. It's an illogical statement.
- mattrobenolt 10y agoYou should actually read the blog post. We are strictly increasing availability in addition to S3's already amazing availability. Not adding another point of failure and assuming we're better. In fact, I literally assume I'm shitty, so I build defenses so my fuck ups don't cause any issues.
- mnutt 10y agoI still think a lot of this depends on your availability requirements. I don't think I could run an S3 service at the scale AWS does with higher reliability, but I do believe I can run a pool of redundant HAProxies with higher reliability, and in fact we already run pools of HAProxies for other reasons so we have quite a bit of operational experience with it. If my company had three engineers I certainly would not go this route, but if you are big enough to have a dedicated ops team that already has experience with this sort of thing, you can architect something that is more reliable than just relying on S3 alone.
- IAmGraydon 10y agoExactly. I'm sorry OP, but with all due respect, you're delusional if you think your strategy would improve upon AWS's availability. The only way you would do that is to deploy a solution like you mentioned on infrastructure that has higher availability than AWS. Hint: that doesn't exist. AWS may have low statistical availability for the month, but on any longer timescale they're still the very best in the game. You need to remember that whatever caused this was a black swan event.
- mnutt 10y agoI disagree. In the last few years S3 has had multiple of these black swan events, while the reverse proxies I am talking about have pushed through hundreds of billions of responses and have had significantly fewer incidents. (In this case, none) I think the fallacy here is that you're not comparing apples to apples: I would be the last to argue that I could run a globally distributed S3 competitor better than Amazon. But I can (and have) run a massively simpler service with better overall uptime because it increases our options during upstream black swan events.
- newobj 10y ago"But I can (and have) run a massively simpler service with better overall uptime because it increases our options during upstream black swan events." I'm not saying it's impossible. But I am saying it's dangerous to omit from this conversation the idea that introducing the very --point of option-- can cause worse reliability than just using the downstream thing in the first place.
- mnutt 10y agoSure, any new piece of infrastructure we add has the possibility for reducing reliability. We only introduce things we think we can support, and that add value for the company. YMMV. We also have less than 15 minute incident engagement times, and don't let important pages slide through dinner. It's totally standard ops stuff: if one of the servers is down, we'll replace it when we get around to it. If they're all down, pages are going off.