3 ms·
I couldn't be more excited about this! Having finished a Kafka project about a year ago, we had so many Zookeeper production and test environment issues that it
by lchengify 6y ago
I couldn't be more excited about this! Having finished a Kafka project about a year ago, we had so many Zookeeper production and test environment issues that it was the running joke to check Zookeeper first if anything went wrong.
Honestly Zookeeper in theory is a great idea: Having a centralized service for maintaining config info saves a lot of heartache when dealing with an open source distributed systems project. But in practice, I've never had a smooth experience getting Zookeeper to run consistently, especially with Kafka.
For sure, part of it is that we treat it like oxygen, in that if it's gone for a few seconds everything just dies. But having dealt with similar systems both proprietary and open source, my opinion is that Zookeeper just hasn't risen to the challenges of its users in the past 5 years. If the next generation of software architects want to use open source streaming or distributed systems, Zookeeper needs to be rewritten or removed.
Also shout out to the confluent.io team: I never paid for your enterprise license, but without your blog posts, docker images, or slack room, I never would have been able to get Kafka working. Thanks again!
- skyde 6y agoi am curious to know why people expect a raft library to be more reliable if embedded inside the kafka controller versus running inside a service like Etcd. in the end broker will do RPC to a service (kafka controller/ etcd) and this service will use raft to replicate the state. It should be exactly the same. And if anything knowing which node are running the raft algorithm help you be more careful with rolling restart and upgrade.
- gen220 6y agoIt might sound kind of trite, but it has less to do with the replication algorithm and more to do with the fact that Zookeeper’s source is quite complicated vs, say etcd, so there are more opportunities for subtle bugs to appear. I encourage you to look at the bug section of the change log for zookeeper, and also the feature list of ZK vs etcd. Also, etcd powers many critical open source projects, so there are many institutional eyes that actively contribute to its improvement. IME if we ever encountered an issue at work with ZK, we found it impossible to trace it down to a bug that we could fix and upstream. Etcd’s been easier in this regard.
- lima 6y agoAgreed, etcd is rock solid. With regards to Kafka, it's probably easier and more robust to add their own consensus layer rather than switching to etcd - Kafka is already a distributed system built by a team of distributed systems engineers. It makes sense for them to build their own consensus, deeply integrated with the replication mechanism, rather than relying on an external database.
- ashtonkem 6y agoIt might be an effect of popularity, but I rarely hear complaints about reliability out of Etcd, while I hear about ZK issues on a consistent basis.