3 ms·
How many nodes can etcd handle without having noticeable decay in performance? their FAQ says 7 but did somebody use it in some other distributed app other than
by nif2ee 7y ago
How many nodes can etcd handle without having noticeable decay in performance? their FAQ says 7 but did somebody use it in some other distributed app other than k8s with more nodes? Assume that most of the operations are get and watch (i.e. write/read <<< 1.0), how big of a cluster in terms of number nodes can we scale up to?
- candiddevmike 7y agoNot sure why youre looking at so many nodes for a cluster. Most of the time scaling the number of etcd nodes does little to help performance, instead focus on giving the nodes plenty of IOPS.
- jzoch 7y agoIf you run a 3 availability zone architecture and want to survive AZ + 1 failures you need more nodes. 9 nodes is the minimum to guarentee that survival while still maintaing a quorum
- sethev 7y agoDo you think you're actually gaining availability with an AZ + 1 setup?
- danaur 7y agoComon, could you be constructive? Of course they do. Now whats your gripe with that?
- sethev 7y agoI was more curious on why they thought that would increase availability. It didn't seem like it would be constructive to speculate about why they chose that. AWS doesn't guarantee that failures within an AZ are independent of each other, so it's not clear how you would estimate what availability you'd gain with this. Losing everything in an AZ + 1 instance sounds like a very unusual and specific scenario to design for.
- danpalmer 7y agoThe use case here is an AZ going down (reasonable to guard against) and an individual machine failure (the first reason we use HA like this anyway). An AZ going down doesn’t make all other hardware reliable, and equally a machine going down from a cluster doesn’t mean that all AZs are going to be reliable. Many products have uptime requirements above what Amazon can provide at the AZ or machine level.
- ideal0227 7y agoTo improve read perf, you can add learner members or add a layer of cache proxy.
- nif2ee 7y agoThanks! never heard of the learner member feature until now. I was asking for an app where most of the nodes are exerting watch operations while only a few other nodes do PUTs due to human intervention. This means that I have a very low write/read ratio. Also I assume that the number of nodes is usually stable so it's not like a very dynamic system where nodes join and leave very frequently. Does this make it easier to have a cluster of 50-100 nodes in different datacenters without breaking etcd down?
- dilyevsky 7y agoProxy is deprecated in v2. V3 has grpc proxy but i think it’s mainly good for coalescing watches and discovery