5 ms·
We published etcd 2.x earlier this year, and it is super stable in current stage. We run harsh failure injection tests on it all the time, and the cluster could
by unihorn 11y ago
We published etcd 2.x earlier this year, and it is super stable in current stage. We run harsh failure injection tests on it all the time, and the cluster could survive well. You could check https://coreos.com/blog/new-functional-testing-in-etcd/ https://coreos.com/blog/new-functional-testing-in-etcd/ for more details.
- mattkrea 11y agoWe're about to start running with it but I am definitely a little concerned at the failover table. I'm a lot more comfortable when I can achieve 2/3 failure and still be okay although the cost of adding a few more instances isn't too bad.
- philips 11y agoTo avoid split-brain a consensus system like etcd cannot make progress with less than 50% + 1 members operating. This is just a hard constraint of this type of distributed system. You can find resources that explain all of the constraints in depth at the Raft homepage: https://raft.github.io/ https://raft.github.io/
- mattkrea 11y agoThanks I'll check it out. Again, the cost of spinning up even a 9-node cluster these days is negligible but my initial plan was to separate etcd clusters by application rather than sharing between but we'll obviously be doing some serious experimentation as we move forward.
- ecnahc515 11y agoYou'll be glad to know that etcd has added basic ACL support to keys. If you were worried about isolation/access to the etcd cluster across apps, this might help you avoid running a cluster per-app.
- mattkrea 11y agoYep, that's exactly what I have some staff internally looking into. Very happy to see it.
- joshuak 11y agoIt's important to note, because it's often missed, that etcd can operate either as a participant in the high reliability / consensus portion of the cluster, but also as a proxy that does not participate. So you don't have to think in terms of 50%+1 of an entire cluster needing to stay up for consistency, only 50%+1 of the etcd cluster. This is nicely illustrated in the the cluster architecture docs at CoreOS.com. https://coreos.com/os/docs/latest/cluster-architectures.html https://coreos.com/os/docs/latest/cluster-architectures.html