6 ms·
Maybe I missed this in the docs, but why isn't HAProxy considered a new point of failure?
by wimagguc 10y ago
Maybe I missed this in the docs, but why isn't HAProxy considered a new point of failure?
- mattrobenolt 10y agoIt's running on localhost to each server. So the failure event here is that somehow haproxy process would explode with the rest of the server being fine. It's much much more likely that a whole machine will die instead or a network issue between machines, etc.
- wimagguc 10y agoSure, that's well understood. Being a low-risk point of failure, isn't it still a new one? It does come with setup and maintenance costs, test scenarios etc., so it's only fair to recognise this as a risk.
- matt_wulfeck 10y agoDon't worry, they run HAproxy in front of haproxy incase the haproxy to S3 service goes down.
- mattrobenolt 10y agoTechnically yes. But we're pretty accustomed to this level of risk. For something like this, the pros far outweigh the cons involved. Yeah, it could fail. The maintenance overhead of this is absolutely minimal and took a handful of hours to have tested and in production. Also worth noting, that this isn't really a single point of failure as a system wide thing. It'd only be a single point of failure on that single node. So if haproxy decided to explode, only that one machine would have a problem momentarily, while the process got started back up with our process manager. The worst case scenario is a human error where we ship a bad config and break everything.
- greenleafjacob 10y agoNot really true. If you for example mistune maxconn haproxy will stop accepting new connections and that's likely to happen cluster wide.
- mattrobenolt 10y agoThis is equivalent to shipping bad application code that takes everything down. Except the config is only a handful of lines of code and will very likely never change again. Also, we don't blindly roll out changes cluster wide for things like this without testing explicitly on staging or test nodes.
- datums 10y agoShipping it with the app, you lose the cluster wide cached objects. A SPOF is the resolver. It's google but it's a SPOF. Is the failover to s3 automatic ? Or do you make a code change ? What kind of latency does that add ?
- nalllar 10y agoHaproxy is localhost. Caching nginx is nearby but not local, so the cache is shared. Haproxy sends to caching nginx if available, else directly to s3.
- datums 10y agoThx. Got it. I don't know enough about the app, but I would have it serve directly from CF to users. Instead of hitting this environment for static assets. Good job and good conversation.
- justinsaccount 10y ago> Each application server that’s running our Sentry code has an HAProxy process running on localhost.