11 ms·
> It’s 1/2 + 1 isn’t it? The parent post is talking about the number that can go down while maintaining quorum, and you're talking about the number that need t
by simtel20 6y ago
> It’s 1/2 + 1 isn’t it?
The parent post is talking about the number that can go down while maintaining quorum, and you're talking about the number that need to remain up to maintain quorum. So you're both correct.
However:
> That would mean in 3 servers you need 2.5 aka 3 machines to commit a change.
That seems wrong. You need N//2 +1 where "//" is floor division, so in a 3 node cluster, you need 3//2 +1, or 1+1 or 2 nodes to commit a change.
- hinkley 6y agoI think I see the problem. 'Simple majority' is based on the number of the machines that the leader knows about. You can only change the membership by issuing a write. Write quorum and leadership quorum are two different things, and if I've got it right, they can diverge after a partition. I'm also thinking of double faults, because the point of Raft is to get past single fault tolerance. [edit: shortened] After a permanent fault (broken hardware) in a cluster of 5, the replacement quorum member can't vote for writes until it has caught up. It can vote for leaders, but it can't nominate itself. Catching up leaves a window for additional faults. It's always 3/5 for writes and elections, the difference is that the ratio of original machines that have to confirm a write can go to 100% of survivors, instead of the 3/4 of reachable machines. Meaning network jitter and packet loss, slows down writes until it recovers, and an additional partition can block writes altogether, even with 3/5 surviving the partition.