5 ms·
We've been using CoreOS in production with etcd and fleet for over 1 year now on 500+ machines. They're have been some growth pains, specifically with etcd - bu
by bkruse 11y ago
We've been using CoreOS in production with etcd and fleet for over 1 year now on 500+ machines. They're have been some growth pains, specifically with etcd - but now they are mostly gone (with the new raft implementation in 2.x). I really appreciate the CoreOS team and all related contributors.
First persistent storage tackled, then networking and now resource-aware orchestration.
- tracker1 11y agoHow has etcd been running for you more recently? One of the things that kept me away was my test run of etcd failed gloriously (twice), and just wasn't stable/resiliant enough for me.
- unihorn 11y agoWe published etcd 2.x earlier this year, and it is super stable in current stage. We run harsh failure injection tests on it all the time, and the cluster could survive well. You could check https://coreos.com/blog/new-functional-testing-in-etcd/ https://coreos.com/blog/new-functional-testing-in-etcd/ for more details.
- mattkrea 11y agoWe're about to start running with it but I am definitely a little concerned at the failover table. I'm a lot more comfortable when I can achieve 2/3 failure and still be okay although the cost of adding a few more instances isn't too bad.
- philips 11y agoTo avoid split-brain a consensus system like etcd cannot make progress with less than 50% + 1 members operating. This is just a hard constraint of this type of distributed system. You can find resources that explain all of the constraints in depth at the Raft homepage: https://raft.github.io/ https://raft.github.io/
- mattkrea 11y agoThanks I'll check it out. Again, the cost of spinning up even a 9-node cluster these days is negligible but my initial plan was to separate etcd clusters by application rather than sharing between but we'll obviously be doing some serious experimentation as we move forward.
- ecnahc515 11y agoYou'll be glad to know that etcd has added basic ACL support to keys. If you were worried about isolation/access to the etcd cluster across apps, this might help you avoid running a cluster per-app.
- mattkrea 11y agoYep, that's exactly what I have some staff internally looking into. Very happy to see it.
- joshuak 11y agoIt's important to note, because it's often missed, that etcd can operate either as a participant in the high reliability / consensus portion of the cluster, but also as a proxy that does not participate. So you don't have to think in terms of 50%+1 of an entire cluster needing to stay up for consistency, only 50%+1 of the etcd cluster. This is nicely illustrated in the the cluster architecture docs at CoreOS.com. https://coreos.com/os/docs/latest/cluster-architectures.html https://coreos.com/os/docs/latest/cluster-architectures.html
- bkruse 11y agoEverything changed in 2.x - it's extremely stable and consistent. Never lost any data in the last 6 months and we beat the hell out of it. 2.x is really a completely different product
- edutechnion 11y agoI went down a similar road with etcd and fleet but abandoned it earlier this summer after testing failure scenarios with etcd. With a cluster of 5 etcd nodes in EC2, I started hard-killing etcd EC2 instances and noticed fleet inconsistency (e.g., nodes being restarted, not able to see the entire fleet). Can you expand on the etcd growth pains you've been through?
- icefall 11y agoI would definitely recommend that you reevaluate it with a newer version of etcd, as it has had some significant stability improvements post-2.0. I've been doing some fuzz testing of it lately and found that it has gotten much more reliable.
- bkruse 11y agoThe basis of this, was being pointed in the right direction by the community. etcd had a HUGE issue with the implementation of the raft consensus algorithm they were using. This was in version 0.x The tough part was that, even though etcd 2.0 was released in January [1], it was not put into CoreOS alpha until April [2] After moving to 2.x - all my problems went away. It had a small learning curve of setting up lots of nodes in the cluster vs proxies [3]. 2.x had a lot of functionality added, but the main one for us was it's reliability. Being able to query status of members, add/remove members from the cluster and monitoring. Before etcd 2.x, the whole etcd infrastructure would die (and consequently, fleet) if just ONE node restarted. Needless to say, it's come a long way. We've been running etcd 2.x since January in a container [4], then just doing export FLEETCTL_ENDPOINT=http://127.0.0.1:2379 http://127.0.0.1:2379 [1] - https://coreos.com/blog/etcd-2.0-release-first-major-stable-release/ https://coreos.com/blog/etcd-2.0-release-first-major-stable-... [2] - https://coreos.com/blog/coreos-alpha-with-etcd-2/ https://coreos.com/blog/coreos-alpha-with-etcd-2/ [3] - https://coreos.com/etcd/docs/latest/admin_guide.html https://coreos.com/etcd/docs/latest/admin_guide.html [4] - https://coreos.com/blog/Running-etcd-in-Containers/ https://coreos.com/blog/Running-etcd-in-Containers/