5 ms·
Alternate title: "Replacing S3 downtime for vastly greater amounts of your own downtime" What is the name for this phenomenon where folks think they can out-av
by newobj 10y ago
Alternate title: "Replacing S3 downtime for vastly greater amounts of your own downtime"
What is the name for this phenomenon where folks think they can out-available a thing that has multiple engineers singularly dedicated to nothing more than its availability /and/ operation? It it just hubris? Surely there must be a more clinical name.
- openasocket 10y agoIf the caching server fails, the infrastructure automatically switches to use S3 directly.
- mattrobenolt 10y agoYeah, one of our goals was to not add a new/weaker point of failure. We gracefully fall back to S3 directly if our cache server is down without a hiccup. So there is no operational overhead of this additional cog. If the server has a failure, we'd go back to slightly degraded performance by talking across the country until we brought it back online.
- wimagguc 10y agoMaybe I missed this in the docs, but why isn't HAProxy considered a new point of failure?
- mattrobenolt 10y agoIt's running on localhost to each server. So the failure event here is that somehow haproxy process would explode with the rest of the server being fine. It's much much more likely that a whole machine will die instead or a network issue between machines, etc.
- wimagguc 10y agoSure, that's well understood. Being a low-risk point of failure, isn't it still a new one? It does come with setup and maintenance costs, test scenarios etc., so it's only fair to recognise this as a risk.
- matt_wulfeck 10y agoDon't worry, they run HAproxy in front of haproxy incase the haproxy to S3 service goes down.
- mattrobenolt 10y agoTechnically yes. But we're pretty accustomed to this level of risk. For something like this, the pros far outweigh the cons involved. Yeah, it could fail. The maintenance overhead of this is absolutely minimal and took a handful of hours to have tested and in production. Also worth noting, that this isn't really a single point of failure as a system wide thing. It'd only be a single point of failure on that single node. So if haproxy decided to explode, only that one machine would have a problem momentarily, while the process got started back up with our process manager. The worst case scenario is a human error where we ship a bad config and break everything.
- greenleafjacob 10y agoNot really true. If you for example mistune maxconn haproxy will stop accepting new connections and that's likely to happen cluster wide.
- mattrobenolt 10y agoThis is equivalent to shipping bad application code that takes everything down. Except the config is only a handful of lines of code and will very likely never change again. Also, we don't blindly roll out changes cluster wide for things like this without testing explicitly on staging or test nodes.
- datums 10y agoShipping it with the app, you lose the cluster wide cached objects. A SPOF is the resolver. It's google but it's a SPOF. Is the failover to s3 automatic ? Or do you make a code change ? What kind of latency does that add ?
- justinsaccount 10y ago> Each application server that’s running our Sentry code has an HAProxy process running on localhost.
- newobj 10y agoThe idea that it being a local proxy meaning there is no operational overhead is a dangerous fallacy. If that's an earnestly literal statement from you, then it means you simply haven't encountered the failure modes that these kinds of set up are inclined towards. I've worked at several BigCo's, seen them all implement this pattern, and seen every single one of them have fleet-wide outages due to these innocent "local proxies". Remember FB's 2-3 hour outage 2 years ago? https://www.facebook.com/notes/facebook-engineering/more-details-on-todays-outage/431441338919/ https://www.facebook.com/notes/facebook-engineering/more-det... It was /exactly/ this kind of "local proxy for higher availability/caching over the downstream thing" that caused the outage.
- newobj 10y ago(To be clear, I'm not saying it's an anti-pattern, I'm just saying that calling it "no operational overhead" is naive)
- wmccullough 10y agoMince words all you want, they stayed operational when many others did not. That's success in my book.
- soft_dev_person 10y agoFor an increased risk of going down when everybody else is up. Is that still success? And at what other costs? It all comes down to risk vs. cost vs. gain.
- mattrobenolt 10y agoSure, there's definitely risk. I'm not asserting that it's literally 0 chance. But this risk of this is also tied up with other things that leverage this proxy. So it's not adding a new dependency or a new point of failure. If this has a problem, we'll also have problems talking to other services in our network. And for what it's worth, I've definitely fucked this up in the past and caused downtime as a result of a setup like this. The pros still outweigh the cons in practice.
- arrty88 10y agoYou are assuming your haproxy server is always operational, no?
- mattrobenolt 10y agoIt's running on localhost, so it's not it's own machine. It's local to the servers running the application code.
- mnutt 10y agoThis incident lasted for three hours. The last incident was in 2015 and was also a significant amount of time. Sentry can likely provision a new HAProxy node in minutes. Most importantly, while S3 is relatively stable it's a black box. If you really care about high availability, you want to bring all of the points of failure under your direct control. On the other hand if availability is just a nice-to-have, relying on S3 is probably a better use of time.
- mattrobenolt 10y agoThis is mostly unrelated to our original goals, which was reliability and performance. This wasn't to protect ourselves in the event of S3 going down, the timing just worked out that it saved us during that as well while also accomplishing the original goals.
- mattrobenolt 10y agoSorry, and by reliable, I explicitly mean, the network connection from our datacenter to halfway across the country hiccuped more than I liked and was slower than I liked. Not reliability of the service itself.
- mnutt 10y agoAh, fortuitous timing! :) If you were looking to protect yourself against S3 going down and had bigger drives, I assume you could use cron to sync the entire bucket and use `try_files` to prefer local and fall back to S3 if the file was missing?
- mattrobenolt 10y agoYeah, we could do something like this as well if we cared. We are easily caching 95+% of our active data in a small amount of disk space. It's not that valuable for us to have a complete replica of the full data set.
- newobj 10y ago> If you really care about high availability, you want to bring all of the points of failure under your direct control. This is what exactly I'm talking about. People can't accept that trying to control it is not in any way guaranteed to make it more highly available. Bringing points of failure under your control makes no innate guarantee of improving anything. It can easily make it worse. The only thing it can do is let you blame yourself of blaming S3. You can only be hardened for what you have anticipated or experienced before. There are tens thousands of wall clock hours of operational experience w/ S3. Availability is one of the top concerns of all AWS. Thinking you can be more available is just fooling yourself. Believing you are more available will entail a willful ignorance or distortion of metrics. Do your engineers carry pagers and have a <15 minute engagement time? Do your engineers sit at home when they are on-call because they know they can't simply let a page slide because they were in the middle of dinner? Or is your company more lenient than Amazon when it comes to operations? Do your engineers spend a quarter fixing some failure mode of your infrastructure, or are they too busy working on features? Is your team's performance measured by availability of your service? Or your actual core business?
- Klathmon 10y agoI don't necessarily agree with it, but the common response to this is all about timing. With AWS you don't have any control over "higher risk" times. If you have a massive launch coming up, or you are nearing peak usage for the year, or your clients need you to be stable for the next few months, you can't put updates on hold, you don't know to get a few more people on standby, you can't choose to not make changes to your system, because it's not your system. With an in house solution you can choose to lock it down for a month, or do the risky upgrades/changes at your lowest traffic time, or even give your customers a heads up if needed. Hell even just being able to mak e sure that your best sysadmin isn't out getting hammered when you go to make changes could go a long way. There is some merit to that idea, but I personally feel the track record of many of these services is so near perfect that the chances of unexpected downtime is still smaller than most could realistically manage.
- Veratyr 10y ago> What is the name for this phenomenon where folks think they can out-available a thing that has multiple engineers singularly dedicated to nothing more than its availability /and/ operation? S3 is an object storage system, they're adding a proxy. It's pretty easy to make a proxy that has better uptime than S3 because it's far, far less complex.
- Scaevolus 10y agoS3 is an object storage system. EBS is block storage.
- eropple 10y agoMaybe I'm reading this wrong, but to me it looks like their solution has a strict positive impact on reliability unless you are concerned about the local HAProxy node causing a problem. (Local to the running service, that is--it looks like it's on the same box?) It caches or falls through as appropriate, does it not?
- mattrobenolt 10y agoThis is correct assessment. HAProxy is running on localhost, and it strictly falls back to hitting public S3 directly if our cache is down.
- matt_wulfeck 10y agoYes! All of the time! When people talk about moving away to their own datacenter, I like to ask how many of the top engineers in the world will be on-call at any time to monitor it?
- mattrobenolt 10y agoFortunately, I am top engineer and am on-call for what I implement and rely on. :) This is also why I build things in a way that don't add more risk to production infrastructure. If you'd read the post, you'd see that this is a strict improvement without risk on our side from introducing a new dependency.
- Sanddancer 10y agoCaching proxies are old hat. Things like squid have existed pretty much the entire lifespan of the web. Keeping things close to your servers means you don't have the inherent delay of pulling things from a remote server, regardless of how fast amazon makes things. Additionally, not keeping all your eggs in Amazon's basket means you're not SOL when they have a datacenter hosting all your content go down. It also means that if and when a service that better fits your needs comes along, you are more readily able to migrate without problems. Finally, site reliability is not something that takes a team the size of Amazon's. A lot of the things that are required for availability on AWS -- redundant systems providing services, standbys, etc -- are things that sysadmins were doing before AWS was extant. AWS' biggest gift to reliability is that its instances are less stable than most dedicated servers; you're taught from day one not to rely on a single server, so you build it right the first time. So no, it's not hubris. It's calculating price/performance, it's applying things you're probably already doing to a new problem, and figuring out what the best solution really is, which rarely involves just throwing money at Amazon.
- exclusiv 10y agoI was thinking the same thing... but if your proxy goes down, couldn't you update code quickly at the app level to go direct to S3 then? Whereas - if you just rely on S3 directly, if they have problems, there's not much you can do unless you also have all assets locally on your servers.
- mattrobenolt 10y agoWe do this automatically already with HAProxy. So we don't even have to change our application.
- patrickg_zill 10y agoI think you have never run servers in a good datacenter before. I know of a datacenter in Denver CO with 14 YEARS of uptime, for instance.
- qaq 10y agoIt's called real world our clients running on dedicated hardware at 2 dcs have consistently less issues than those on AWS. The aws control layer and infrastructure is too complex and results in fairly significant outages.
- CodingGuy 10y agoThat's it. I'm running my own dedicated servers for my business and had zero (ZERO!) downtime in the last three years. How much downtime had the big cloud players?
- meirelles 10y agoMe too. I've been renting servers for over 10 years. By the way is fairly common DCs with very high uptime (between 99.999% - 100% over 5 years or more), especially in cities with high connectivity like Ashburn/VA or Dallas/TX. Failures on servers with less than 4 years of use is _really_ unusual.
- paulddraper 10y agoWell, Dyn kinda sucked for some big players. Do you host your own DNS?
- grey-area 10y agoEvery time cloud services have an outage, this line of reasoning becomes less and less appealing. Your analysis assumes all other factors are constant. Change causes downtime, cloud services have a high rate of change (constantly pushing new configs, new code), many other servers don't.