5 ms·
Can someone ELI5 the difference between using AWS availability zone affinity and then simply dropping the downed AZ at the top most routing point? Wouldn't tha
by purpleturtle22 3y ago
Can someone ELI5 the difference between using AWS availability zone affinity and then simply dropping the downed AZ at the top most routing point?
Wouldn't that be the same thing, with the obvious caveat you are t using the routing technology Slack is using (We don't - We use vanilla AWS offerings)
- ec109685 3y agoIsn’t that exactly what they are doing? Keeping requests within an AZ and instead of using DNS at the first hop into AZ, they use envoy to control traffic shaping and making that initial decision if traffic needs to be routed away.
- t0mas88 3y agoThey decided to use every routing tool available at least once in their setup, so they can't do this. But there is no explanation in the blog about why they use so many platforms and so many routing tools. Sounds to me like they got themselves into a mess and decided to continue on that path.
- jonathankoren 3y agoSomewhere, an engineering “leader” is going to point to this blog post and then say, “Well, that’s how Slack did it!” and promptly copy this overwrought system
- Terretta 3y agoYou're doing it right.
- ec109685 3y agoIsn’t that exactly what they are doing? Keeping requests within an AZ and using global DNS at the first hop into AZ.
- esprehn 3y agoCells are not about guarding against AZ failure, but about partitioning the production infra to protect against bad deploys and configuration changes. Every AZ is split into many different cells.
- hliyan 3y agoSo, guarding against human errors / process failures, and not hardware failures?