3 ms·
Admittedly this is armchair architecture talk, but it seems like either consul or Roblox's use of Consul is falling into a CAP-trap: they are using a CP system
by NightMKoder 5y ago
Admittedly this is armchair architecture talk, but it seems like either consul or Roblox's use of Consul is falling into a CAP-trap: they are using a CP system when what they need is an eventually-consistent AP system. Granted, the use of consul seems heterogenous, but it seems like the main root cause was service discovery. And service discovery loves stale data.
Service discovery largely doesn't change that often. Especially in an outage where a lot of things that churn service discovery are disabled (e.g. deploys), returning stale responses should work fine. There's a reason DNS works this way - it prioritizes having any response, even if stale, since most DNS entries don't change that frequently. That said, DNS is not a great service discovery mechanism for other reasons. Not sure if there's an off-the-shelf solution that relies more on fast invalidation rather than distributed consistent stores.
- tptacek 5y agoCan you say more about service discovery "loving stale data"? Loves in the sense of "generates a lot of it; is constantly plagued by it"?
- boulos 5y agoTheir comment implies "are totally fine with stale data". Their argument is that the membership set for a service (especially on-prem) doesn't change all that frequently, and even if it's out of date, it's likely that most of the endpoints are still actually servicing the thing you were looking for. That plus client retries and you're often pretty good.
- tptacek 5y agoMaybe I'm just working on an idiosyncratic version of the service discovery problem, but "stale data" is basically my bête noire. Part of it is that I don't control all my clients, and I can't guarantee they have sane retry logic; what service discovery tells them is the best place to go had better be responsive, or we're effectively having an outage. For us, service discovery is exquisitely sensitive to stale data.
- boulos 5y agoYep! I'm not saying I totally agree with the original comment there, just confirming that they meant it. If you own your clients, sometimes you can say "it's on you to retry" (deadlines and retries are effectively mandatory, and often automatic, at Google). Having user facing services hand out bad addresses / endpoints would be really bad. However, even for things like databases, you really want to know who the leader / primary is (and it's not really okay to get a stale answer). So I dunno, some things are just fine with it, and some aren't. It's better if it just works :). Besides, if the data isn't changing, the write rate isn't high!
- NightMKoder 5y agoYes - exactly what boulos said - I’m coming from the Google “you control your clients” perspective. That said, in some sense you always control your clients - you can always set up localhost proxies that speak your protocol or just tcp proxies. The thing is service discovery is _always_ inconsistent. By the time you get your set of endpoints from your discovery, it can be out of date by the time you open your socket. Certainly for something like databases you need followers or standby leaders to reject writes from clients - service discovery can’t 100% save you here.
- throwdbaaway 5y agoGood catch. If Roblox only uses consul for service discovery, things should continue to work, just slowly degrade over the hours/days. There should at least be one consul agent running on each physical hosts, and this consul agent has cache and can continue to provide service discovery functionality with stale data. Dissecting this paragraph from the post-mortem... > When a Roblox service wants to talk to another service, it relies on Consul to have up-to-date knowledge of the location of the service it wants to talk to. OK. > However, if Consul is unhealthy, servers struggle to connect. Why? The local "client-side" consul agents running on each hosts should be the authoritative source for service discovery, not the "server-side" consul agents running on the 5 voter nodes. > Furthermore, Nomad and Vault rely on Consul, so when Consul is unhealthy, the system cannot schedule new containers or retrieve production secrets used for authentication. Now that's one very bad setup, similar to deploying all services in a single k8s cluster.
- NightMKoder 5y agoDidn’t realize consul had that. Seems like the right approach - though I wonder why Roblox wasn’t using it. Fwiw I believe kubernetes did this right - if you shoot the entire set of leaders, nothing really happens. Yes if containers die they aren’t restarted and things that create new pods (eg cron jobs) won’t run, but you don’t immediately lose cluster connectivity or the (built-in) service discovery. Not to say you can survive az failures or the like - or that kubernetes upgrades are easy/fun. And don’t run dev stuff in your prod kube cluster. Just…don’t.