3 ms·
Interesting timing. I was playing with Etcd just this morning. I'm glad there's some more options in this space. I haven't been happy with any setup so far. Do
by jandy 13y ago
Interesting timing. I was playing with Etcd just this morning. I'm glad there's some more options in this space. I haven't been happy with any setup so far.
Doozer (https://github.com/ha/doozerd https://github.com/ha/doozerd) got
me excited. It's small, fast, and written in Go. Unfortunately, its development seems quiet and fragmented. Its
lack of TTL-style values made it a pain to do a distributed lock service without having a sweeper for cleaning up dead locks.
Zookeeper (http://zookeeper.apache.org http://zookeeper.apache.org) is much more fully featured and
mature, but felt way too heavy compared to my nimble Go stack. Installing
and maintaining a JVM just for Zookeeper made me uncomfortable.
Etcd is interesting. It has TTLs, it's small and fast, easy to pick up/learn, and it's in active
development (and it's tied with CoreOS and Docker, so it's bound to get some
reflected love).
- philsnow 13y ago> Zookeeper (http://zookeeper.apache.org http://zookeeper.apache.org) is much more fully featured and mature, but felt way too heavy compared to my nimble Go stack. Installing and maintaining a JVM just for Zookeeper made me uncomfortable. I've spent almost a year dealing with a large, high-traffic zookeeper installation and I agree. Well, actually etcd seems to have some features which I would love to have in Zookeeper. I maintain several Zookeeper ensembles which I want to be highly available. Any time I need to swap out a node, increase or decrease the size of the quorum, change a node's port(s), or change a node between voting and observing, I have to draw out a diagram where I keep track of which nodes "know" what at which times. If I skip doing that, I run into situations where fewer than N/2 + 1 nodes agree on the current state of the ensemble, and they fail-stop and don't serve traffic. Here's a specific example of an issue that seems so blindingly obvious to me, but it clearly wasn't to whoever implemented it: if you want to specify that a zookeeper node is an observer (doesn't participate in leader elections), you have to put that in the config file in two places: once on the line for that node in the section where you tell all the nodes where all the other nodes are (like [0]), and you also have to have a separate line "peerType=observer". This last bit means you can't use the same zoo.cfg file for your observer nodes and your voter nodes, you have to keep two zoo.cfgs, make your init script or whatever use the correct one, and keep the files semantically in sync if you ever have to make further changes. What they should do is have each node look at [0] and say "oh I'm server.4 so I'm supposed to be an observer". It's piles and piles of little annoyances like that make me dislike Zookeeper. I'll be watching etcd. [0] zoo.cfg [snip] server.1=hostname1:port1:port2 server.2=hostname2:port3:port4 server.3=hostname3:port5:port6 server.4=hostname4:port7:port8:observer
- justinsb 13y agoThe next version of Zookeeper (3.5) allows for dynamic reconfiguration, so it'll be much easier to reconfigure your cluster online. Hopefully we'll also have self-repairing ZK clusters and you won't have to manage Zookeeper unless things are going seriously wrong (e.g. you only have 2 machines online across your whole cluster). Here's the bug: https://issues.apache.org/jira/browse/ZOOKEEPER-107 https://issues.apache.org/jira/browse/ZOOKEEPER-107 It may have taken 5 years, but this will be fixed. I love etcd & Zookeeper!
- lucian1900 13y agoZookeeper has a feature I particularly like that etcd does not (for now at least): it's possible to write to a node from a client such that the node will disappear if the client disconnects. This plus watching makes for a great liveness check between different machines.
- jandy 13y agoYes, that's true. It is pretty easy to simulate this with a short TTL and a heartbeat though.