4 ms·
This incident lasted for three hours. The last incident was in 2015 and was also a significant amount of time. Sentry can likely provision a new HAProxy node in
by mnutt 10y ago
This incident lasted for three hours. The last incident was in 2015 and was also a significant amount of time. Sentry can likely provision a new HAProxy node in minutes.
Most importantly, while S3 is relatively stable it's a black box. If you really care about high availability, you want to bring all of the points of failure under your direct control. On the other hand if availability is just a nice-to-have, relying on S3 is probably a better use of time.
- mattrobenolt 10y agoThis is mostly unrelated to our original goals, which was reliability and performance. This wasn't to protect ourselves in the event of S3 going down, the timing just worked out that it saved us during that as well while also accomplishing the original goals.
- mattrobenolt 10y agoSorry, and by reliable, I explicitly mean, the network connection from our datacenter to halfway across the country hiccuped more than I liked and was slower than I liked. Not reliability of the service itself.
- mnutt 10y agoAh, fortuitous timing! :) If you were looking to protect yourself against S3 going down and had bigger drives, I assume you could use cron to sync the entire bucket and use `try_files` to prefer local and fall back to S3 if the file was missing?
- mattrobenolt 10y agoYeah, we could do something like this as well if we cared. We are easily caching 95+% of our active data in a small amount of disk space. It's not that valuable for us to have a complete replica of the full data set.
- newobj 10y ago> If you really care about high availability, you want to bring all of the points of failure under your direct control. This is what exactly I'm talking about. People can't accept that trying to control it is not in any way guaranteed to make it more highly available. Bringing points of failure under your control makes no innate guarantee of improving anything. It can easily make it worse. The only thing it can do is let you blame yourself of blaming S3. You can only be hardened for what you have anticipated or experienced before. There are tens thousands of wall clock hours of operational experience w/ S3. Availability is one of the top concerns of all AWS. Thinking you can be more available is just fooling yourself. Believing you are more available will entail a willful ignorance or distortion of metrics. Do your engineers carry pagers and have a <15 minute engagement time? Do your engineers sit at home when they are on-call because they know they can't simply let a page slide because they were in the middle of dinner? Or is your company more lenient than Amazon when it comes to operations? Do your engineers spend a quarter fixing some failure mode of your infrastructure, or are they too busy working on features? Is your team's performance measured by availability of your service? Or your actual core business?
- mattrobenolt 10y agoIronically, in our case, our availability is very core to our business. In this exact scenario, if S3 would have blocked our processing pipeline because it was down, that means we couldn't have alerted users as reliably that they were in fact, having issues because of S3 being down as well. So in our case, this is massively important to us and worth any additional risk that may be introduced.
- IAmGraydon 10y agoAre you realizing that risk spread out over time is what creates your downtime average? So what you are saying is that your uptime is so important that it is worth additional downtime. It's an illogical statement.
- mattrobenolt 10y agoYou should actually read the blog post. We are strictly increasing availability in addition to S3's already amazing availability. Not adding another point of failure and assuming we're better. In fact, I literally assume I'm shitty, so I build defenses so my fuck ups don't cause any issues.