4 ms·
I disagree. In the last few years S3 has had multiple of these black swan events, while the reverse proxies I am talking about have pushed through hundreds of b
by mnutt 10y ago
I disagree. In the last few years S3 has had multiple of these black swan events, while the reverse proxies I am talking about have pushed through hundreds of billions of responses and have had significantly fewer incidents. (In this case, none)
I think the fallacy here is that you're not comparing apples to apples: I would be the last to argue that I could run a globally distributed S3 competitor better than Amazon. But I can (and have) run a massively simpler service with better overall uptime because it increases our options during upstream black swan events.
- newobj 10y ago"But I can (and have) run a massively simpler service with better overall uptime because it increases our options during upstream black swan events." I'm not saying it's impossible. But I am saying it's dangerous to omit from this conversation the idea that introducing the very --point of option-- can cause worse reliability than just using the downstream thing in the first place.
- mnutt 10y agoSure, any new piece of infrastructure we add has the possibility for reducing reliability. We only introduce things we think we can support, and that add value for the company. YMMV. We also have less than 15 minute incident engagement times, and don't let important pages slide through dinner. It's totally standard ops stuff: if one of the servers is down, we'll replace it when we get around to it. If they're all down, pages are going off.