3 ms·
(author of the original post here) Raft can be pretty relaxed about heartbeats. It doesn't, strictly speaking, require every node to talk to everyone else. All
by bdarnell 11y ago
(author of the original post here) Raft can be pretty relaxed about heartbeats. It doesn't, strictly speaking, require every node to talk to everyone else. All you really need is some health signal about the other nodes so you A) don't call an election while the leader is still alive and B) call an election promptly when a leader disappears. As a first step, we've reduced our heartbeat traffic from one per group to one per node, and we can reduce it further (e.g. by having each node send heartbeats only to a subset of its peers and sharing the results to other nodes).
Also, we are using both Raft and two-phase commit: Raft manages the consistency of individual ranges, but transactions that span multiple ranges require 2PC.
- ccleve 11y agoThere's an interesting comment on the blog post: > Does handling all range's consensus traffic in bulk would be the most effective way to handle this? Instead could all the ranges from a node be represented as a single state machine to be handled by plain etcd raft implementation, like the approach taken in spanner paper? Do you have any comment on this? I've been trying to wrap my head around the implications. What's the downside?
- bdarnell 11y agoI replied on the blog post. The downside is basically that if you have one giant consensus group, then to add a new replica you have to replicate that entire consensus group. This hurts load balancing and recovery times, since you can't scatter a dead node's ranges across all the other nodes in the cluster.