4 ms·
I'd like to know what you base your statement on that the Raft implementations in etcd or CockroachDB are incorrect. Your original paper does not mention those
by Thaxll 2mo ago
I'd like to know what you base your statement on that the Raft implementations in etcd or CockroachDB are incorrect. Your original paper does not mention those implementations, so where does that claim come from?
- gyesxnuibh 2mo agoHaving run a fleet of 100s of etcd clusters for 10000s of rps, and the fact that upstream runs tests similar to antithesis and recently partnered with antithesis [0], and jepsen has tested it long ago as well [1]. Etcd's raft algorithm is fine. Someone even did a TLA+ proof on it in the last couple years[2]. Yes there was a correctness issue a few years ago but otherwise the person you're replying to doesn't know what they're talking about. Also those bugs have nothing to do with the raft implementation, but instead the state machine implemented on top. 0: https://etcd.io/blog/2025/autonomus_testing_with_antithesis/ https://etcd.io/blog/2025/autonomus_testing_with_antithesis/ 1: https://jepsen.io/analyses/etcd-3.4.3 https://jepsen.io/analyses/etcd-3.4.3 2: https://github.com/etcd-io/raft/pull/113 https://github.com/etcd-io/raft/pull/113
- cobbzilla 2mo agoIs the correctness of its implementation of the algorithm unaffected by bugs in this state machine? Maybe I missed something.
- gyesxnuibh 2mo agoThe raft algorithm works and if you implemented a less complex state machine (like using a simpler kv store that doesn't need global event ordering via revisions and watches) it would work. That's what antithesis said they did to test the raft algorithms in the other article linked
- iscoelho 2mo ago"there was a correction issue" is downplaying it. Etcd is truly the worst example of Raft. Etcd corruption and loss of quorum is extremely common in practice and the GitHub issues sit for years. The design is simple, the performance is modest, yet it still has still never been reliable, despite being marketed as so. I can't speak to whether this is specifically due to their Raft implementation, but I'd argue the entire codebase is over-engineered and questionable.
- gyesxnuibh 2mo agoTheir lock, leader election, sessions, and leases are all awful and I'd never recommend anyone to use those. But as a strongly consistent kv store and if you need the watch mechanics, its useful. It has its place and that's mostly being used by kubernetes.
- AlphaSite 2mo agoMy beef with etcd is that its neither performant nor reliable. Its very much {reliable, performant, flexible} pick none.
- otterley 2mo ago> Etcd corruption and loss of quorum is extremely common in practice and the GitHub issues sit for years. Do these have reproducible test cases?
- dnautics 2mo agodoesnt raft have a problem that it assumes no hysteresis? and that in general you can construct a latency graph that deterministically causes a permanent lock in the leadership election phase?
- gyesxnuibh 2mo agoI'm not sure I fully understand your question, but the heartbeat time outs and leader election time outs are static. And yeah if you make votes and heartbeats time out in a way that nobody can be elected, then raft can't make progress. I've never had this be a problem in reality but AWS has a pretty good backbone. Maybe if you were running it over a pretty unreliable network this would be an issue?