4 ms·
I have seen some issues around Consul these days. As a person with no background in distributed systems, I am wondering why people choose Consul over alternati
by kbumsik 4y ago
I have seen some issues around Consul these days.
As a person with no background in distributed systems, I am wondering why people choose Consul over alternatives. Are there features that etcd doesn't offer?
- mrkurt 4y agoWe chose Nomad and adopted Consul as a result. Nomad and Consul work well together. I don't believe etcd would have been any better for us, though. Centralized service discovery that runs through raft consensus doesn't make a lot of sense for the things we need to do. And when I've had etcd blow up on me in the past, it's been similarly painful to recover from.
- aeyes 4y agoMost people only use etcd at small scale. If you try to store 10 or even 100GB in etcd you are going to run into uncommon problems. Most people don't even know that the Kubernetes control plane by default has a hard limit on etcd size. It used to be 2GB, not sure what it is now.
- kbumsik 4y agoDoesn't Consul have the similar storage limit btw? I have seen very few strongly consistent distributed KV store that scales beyond 10GB+
- dilyevsky 4y agoIt’s the underlying db limit (boltdb) which both etcd and consul use
- mdaniel 4y agoI would need a citation on the kubernetes control plane having any such hard limit. etcd is its own little snowflake, and I could very easily imagine it having some bad default value like that, or even kubeadm improperly configuring it However, related to that, for big-time clusters (q.v. https://news.ycombinator.com/item?id=35174655 https://news.ycombinator.com/item?id=35174655 and https://news.ycombinator.com/item?id=25907312 https://news.ycombinator.com/item?id=25907312) one should without question move events over into their own etcd cluster: https://openai.com/research/scaling-kubernetes-to-2500-nodes#etcd https://openai.com/research/scaling-kubernetes-to-2500-nodes...
- dilyevsky 4y agoIt’s etcd default https://etcd.io/docs/v3.5/dev-guide/limit/#storage-size-limit https://etcd.io/docs/v3.5/dev-guide/limit/#storage-size-limi... There’s also max object size of 1MB on the apiserserver side I believe
- ivzhh 4y agoByteDance replaces etcd with kubebrain [1], which is backed by their own KV store (TiKV seems also supported). The single-group raft is the hard limit. [1]: https://github.com/kubewharf/kubebrain https://github.com/kubewharf/kubebrain
- mdaniel 4y agoThat's interesting, thanks for the link. I held out high hopes for pluggable KV in kubernetes for the longest time, but since that issue was closed WONTFIX I resigned my hopes Heh, that kubebrain TODO is some "oh, really?" * Guarantee consistence in critical cases but I give them huge props for calling out Jepsen
- dilyevsky 4y agoI’ve run it with over 50G under heavy load (10k+ qps) and it was fine. It’s pretty sensitive to disk latency though
- kbumsik 4y agoThanks for the answer! AFAIK doesn't Consul also use Raft?
- grrdotcloud 4y agoRaft is amazing and totally frustrating. I think I understand how you're using it and curious if you've considered how AWS STS API manages their cross region syncing gets solved.
- pcthrowaway 4y agoEtcd is really only for basic config. If you want apps to discover each other and be able to communicate effortlessly, even across datacenters, Consul, in theory, enables this. I say in theory because I couldn't get federated Consul actually working.
- convolvatron 4y agodiscovery isn't that hard a problem that you should cede your agency to a external party like Hashicorp I used consul for a clustered service once, it was worth it for bringup. but I when I had problems I just wrote one in a couple days since I'd done so several times before. and it didn't fail for all the years that product was running.