3 ms·
I'm not sure Raft is the best distributed consensus algorithm for the situation of a global, unstable, frequently-partitioning network. I think it is in its nic
by kortex 4y ago
I'm not sure Raft is the best distributed consensus algorithm for the situation of a global, unstable, frequently-partitioning network. I think it is in its niche when leaders are running on fairly stable networks (>1-2 nines), and the main source of node failures are due to task cycles / rolling deploys.
I've played around with Hashicorp Consul on "edge boxes" - long-haul, wirelessly-connected embedded computers, with unreliable power supplies. Allowing edge boxes be Consul leaders results in all kinds of mayhem: split brain situations, corrupted state, stale DNS resolutions (Consul handles DNS as well), cats and dogs living together, mass hysteria. A much better topology is to have 3 server nodes on a LAN as the "head cluster" and letting all the edge boxes be clients of the head.
I haven't used it but Consul has a multi-datacenter mode, which I believe is designed to better handle such a situation, which I believe has a dedicated raft cluster per datacenter.
https://learn.hashicorp.com/tutorials/consul/federation-gossip-wan https://learn.hashicorp.com/tutorials/consul/federation-goss...