4 ms·
I found this article [0] a couple weeks ago and found it informative. Basically Consul has service discovery built in, while Zookeeper et al. expect you to buil
by jake-low 9y ago
I found this article [0] a couple weeks ago and found it informative. Basically Consul has service discovery built in, while Zookeeper et al. expect you to build service discovery on top of the primitives they provide (assuming you need it).
[0]: https://www.consul.io/intro/vs/zookeeper.html https://www.consul.io/intro/vs/zookeeper.html
- kevinconaway 9y agoAlthough its not part of Zookeeper core, there is an Apache Curator module[1] for doing service discovery on Zookeeper [1] https://curator.apache.org/curator-x-discovery/ https://curator.apache.org/curator-x-discovery/
- takeda 9y agoThis article is biased because it was written by Consul author. I'm big fan of Zookeeper, and agree with parent that is not only industry standard, but it is very robust and battle tested. Having said that, this is not what Zookeeper was meant to. ZK is coordination service, if you have a distributed service that you need to coordinate and you don't want any single point of failure ZK is the choice. ZK can be used for service discovery, but that's not the proper use, you could equally set up a single database node, or MongoDB cluster to accomplish essentially the same thing. So specialized service is better, but I also don't like how Consul solves service discovery. If you think about it, SD doesn't require strong consistency, it is not a big deal if all nodes won't learn about each other at the same time, arguably that's actually preferred. By removing strong consistency (the only reason why ZK is even compared to Consul) we could have a better SD that's also more scalable.
- scarface74 9y agoIf you think about it, SD doesn't require strong consistency, it is not a big deal if all nodes won't learn about each other at the same time, arguably that's actually preferred. Can you go into more detail about this? Each computer has its own Consul agent that monitors the health of services running on it. You also get configuration from the agent on the individual computer. All of the agents communicate peer to peer and with the server don’t they? Isn’t that eventually consistent? I’m just starting to use Consul for configuration and haven’t started working on the service discovery part yet.
- justinsaccount 9y agohttps://envoyproxy.github.io/envoy/intro/arch_overview/service_discovery.html#on-eventually-consistent-service-discovery https://envoyproxy.github.io/envoy/intro/arch_overview/servi... has some on that
- takeda 9y agoYes, but that's only to detect whether node is up or down, the health checks are still strongly consistent and go through the servers. Because of this design, the don't get benefits of any of them. Side note, KV is actually strongly consistent, which is what is desired, the thing is that since KV comes with consul it encourages people to use it. One just need to keep in mind, that Raft and Paxos expect all nodes to be updated on every change (e.g. if you have 5 server node cluster all 5 nodes need to "sign on it" before it is accepted). This is counter-intuitive since adding nodes to cluster makes it actually slower (the goal is correctness and resilience not performance). Basically you just need to keep in mind to not use KV for some write heavy operations. Serving configuration files generally is fine.
- cube2222 9y agoIn both paxos and raft you only need a quorum to sign on it. In the case of raft/consul it's the majority. Which means, you need 3 of 5 nodes to sign on it. Also, key value store reads are usually done locally, from the agent, which goes over gossip and is eventually consistent. Health checks are also eventually consistent. They are done by the local agent, and through gossip are sent to the masters group.
- takeda 9y agoI'm aware that Paxos and Raft require odd number of nodes. My point was that the more nodes you have the slower the cluster is, because increasing number of the nodes doesn't increase performance but resilience. With 3 nodes you can afford to lose at most 1 node, with 5 -> 2, with 7 -> 3 etc. > Also, key value store reads are usually done locally, from the agent, which goes over gossip and is eventually consistent. That's not what the documentation says. The KV is using raft, you can relax some guarantees for reading by having clients use consistency=stale, but you're still communicating with the server, except in stale you only communicate with one instead all of them. > Health checks are also eventually consistent. They are done by the local agent, and through gossip are sent to the masters group. Also that's not what the documentation says. Gossip is used for adding/removing nodes and the overall health of the node (whether host is reachable) is also done through gossip, but the actual health checks (whether a service is running) is done through consensus protocol. I still want to emphasize that using consensus protocol for SD is not a good idea. Let say you're running in AWS and follow the best practices which is spreading your servers over multiple AZ. Let say you use us-east-1 which has 6 AZ. An AWS starts having network issues (like few months ago) and 2 AZs have network issues and can't communicate with the rest. The machines are still running, just can't communicate. It just happens that on 2 of those AZs you run 2 out of your 3 consul servers. Even though all the machines are healthy and running your Consul cluster is down, even for the remaining 4 AZs that don't have any problems.