5 ms·
There is a popular alternative as well, Raft. https://ramcloud.stanford.edu/wiki/download/attachments/11370504/raft.pdf https://ramcloud.stanford.edu/wiki/downl
by zphds 11y ago
There is a popular alternative as well, Raft. https://ramcloud.stanford.edu/wiki/download/attachments/11370504/raft.pdf https://ramcloud.stanford.edu/wiki/download/attachments/1137...
Has anyone used both? What are the pros and cons?
- sargun 11y agoPaxos is poorly explained in the multi-decree method, because the explanation isn't particularly opinionated about how SMR and logging is done. Raft only has a multi-decree variant, and it has opinions on how to do leader election, etc... Fundamentally, both protocols (multi-decree, multi-paxos, and Raft are equivalent (in performance, and capability)). There are other variants of Paxos (ePaxos) that have advantages to Paxos in some cases, especially in the WAN.
- krenoten 11y agoRaft has a dirty little secret: your cluster breaks when two nodes become partitioned from each other (but not the rest of the cluster) and one of them is the leader. Script: nodes {1, 2, 3*} where 2 and 3 are partitioned, and 3 is the current leader. Node 2 fires its leader election timer, broadcasts RequestVotes to 1 and 3, only 1 gets it, but now 3 is ignored for the term of its leadership and 2 is the only node capable of quorum. In a decent implementation 3 will be nack'd the next time it sends an AppendEntries to 1, and 1 will pass along the current term. 3 isn't hearing from the leader so broadcasts RequestVotes (either failing once due to failure to hear the current term from 1, or jumping straight into the next term, which actually increases the ratio of cluster livelock). Leadership bounces back and forth rapidly, making your cluster worthless. The current hotness is spec paxos, which gets the positive trade-offs of both cheap paxos and fast paxos, if you're able to actually implement it in your DC (has networking assumptions).
- r0naa 11y agoHow's that a "dirty little secret"? Raft not supporting asymmetric partitions is somewhere on the third line of the original paper. Besides that, I genuinely wonder how often does this happen in the real world? I see a typical case where parts of your cluster is behind a NAT. Otherwise, what could cause an asymmetric partition?
- kelnos 11y agoThat sounds easy, actually. Take a cloud provider like AWS. They divide their regions into "availability zones", which are supposedly in different buildings, don't share the same power, internet connectivity, etc. Say you have 3 AZs, and you put a node in each. A simple network connectivity issue between AZ 1 and 2 (but not between 1 and 3 or 2 and 3) would cause this scenario to happen, assuming I'm understanding this correctly.
- vidarh 11y agoConsider a 3 location network with routes between each pair of locations. Then the link between two of them fails without mechanisms in place to automatically re-route. That's not such an uncommon situation. E.g. quite a few large to mid-sized companies with servers in three locations does not have the in-house skillset to either get BGP set up and/or set up VPNs between the locations and a means to update routes automatically on outages. Some larger shops will have "metro LAN" type setups with separate ports for dedicated connections between racks in different data centres, and won't even have capacity enough on their public ports to handle a failover of traffic normally going on the private connection between two of the data centres over the public port - the expectation is often that the data centre operator will handle redundancy... (until they don't...) It's not a terribly hard thing to fix, but it's also something it seems few people think about until they reach much larger size.
- toolslive 11y agoI always considered Mencius to be the "next hotness in consensus algorithms".
- zerebubuth 11y ago> The current hotness is spec paxos, which gets the positive trade-offs of both cheap paxos and fast paxos, if you're able to actually implement it in your DC (has networking assumptions). Sounds interesting, is "spec paxos" like Spanner in requiring special hardware? If you could share some links or references to more information about it, that would be great, thanks.
- krenoten 11y agohttps://syslab.cs.washington.edu/research/specpaxos/ https://syslab.cs.washington.edu/research/specpaxos/
- nicolast 11y agoRaft is 'simpler' than Paxos because it's not only a consensus protocol, it include log management etc. On the other hand, Raft could be (with a stretch) considered one of the protocols in the Paxos family, whilst other family members have their specific strengths (and complexities), e.g. being more efficient in WAN networks. If you're interested in Raft, take a look at Kontiki (https://github.com/NicolasT/kontiki https://github.com/NicolasT/kontiki). Could use some maintenance though...