8 ms·
We tried to adopt this but found the documentation very lacking and a severe lack of quality client libraries for our language of choice (go).the "official" one
by mattboyle 7y ago
We tried to adopt this but found the documentation very lacking and a severe lack of quality client libraries for our language of choice (go).the "official" one had race conditions in the code as well as "todo" for key pieces littered throughout. There is another from comcast which is abandoned. We had a serious discussion about picking up ownership of the library or writing our own but as a small start up we didnt feel we could do it and still develop the product. I'll continue to keep an eye on pulsar but for now Kafka is the clear go to imo. It's well documented, great SAS offerings (confluent) and tons of books and training courses for it.
- ckdarby 7y ago> found the documentation very lacking Really? It is one of the few open source projects that we've felt has had modern documentation. How long ago was this? > As a small startup You'll spend more time & money on the OpEx cost with Kafka than picking up the client library for Pulsar.
- cpx86 7y ago> You'll spend more time & money on the OpEx cost with Kafka than picking up the client library for Pulsar. Could you elaborate why this would be the case?
- gen220 7y agoNot the OP, but I think they were exaggerating a bit. In practice, operating kafka is a major PITA, because it means you have to (1) choose a "flavor" wrapper (confluent seems to be a popular one), because the base project isn't easy to develop against (2) write your own wrappers of those wrappers, to keep your developers from shooting themselves in the foot with wacky defaults (3) suffer the immense pain that is authenticating topic write/reads, if that's even possible??? (4) stand up zookeeper... and probably lose some data along the way. (5) suffer zookeeper outages due to buggy code in kafka/zk (I've experienced lost production data due to unpredictable bugs in kafka/zk, but obviously YMMV). Based on my naive assessment, the kafka/zookeeper ecosystem is maybe 10x as complicated as the problem it's solving, and that shows up in the OpEx. I personally doubt that Pulsar is that much better, but it might be.
- ckdarby 7y agoThese are also valid. I wrote the reply explaining some of the OpEx here: https://news.ycombinator.com/item?id=21938463 https://news.ycombinator.com/item?id=21938463
- EdwardDiego 7y agoWhat do you mean by 1 and 2? I'm guessing you're referring to the kafka-clients API? The defaults for producer and consumer conf are quite sensible these days.
- gen220 7y agoI wasn’t around to make those decisions at my company, but I imagine that the “these days” component was the cause? There are a lot of configurations, new ones appear and old ones disappear or change names, etc. In this churny environment, where you want to keep on latest versions (necessitated by bugs mentioned in), you need abstractions to protect you somewhat from the churn. Confluent also seems to have a fair amount of churn, so you need wrappers for that, that you can update all at once for your developers.
- EdwardDiego 7y agoSorry, when I say these days, I mean >= Kafka 1.0. Things like auto commit offset in 0.8 days were something like 1 minute, as opposed to 5 seconds onwards, max fetch bytes was set significantly higher etc. My biggest problems with it were when developers who didn't really understand Kafka started setting properties that had promising names to bad values to "ensure throughput" - let's set max.poll.records to 1 to ensure we always get a record as soon as one is available! That might be my biggest issue with Kafka - it requires a decent amount of knowledge of Kafka to use it well as a developer. I'm not sure if Pulsar removes that cognitive burden for devs or not, but I'm interested in finding out. And yeah, the wrappers to remove that burden were written in our company too - but then proved quite limiting for the varying use cases for a Kafka client in our system. sigh
- mattboyle 7y agoIt was about 6 months ago. I completely disagree with the opex of picking up kafka vs developing a whole client library. Please could you try and explain how you came to this conclusion?
- ckdarby 7y ago> Please could you try and explain how you came to this conclusion? 1. Stateless brokers With Kafka any time a broker goes down you need to be aware of the kafka broker id. Yes, this can be fixed by creating your entire infrastructure as code and keeping track of state. This is something of great OpEx. I've seen few people successfully automate this, Netflix is one of the few. The rest just use manual process with tooling to get around, pager, Kafka tooling to spawn replacement node with the looked up broker id, etc. 2. Kafka MirrorMaker Granted I have not used v2 that recently came out in ~2.6 but dear gosh v1 was so bad that Uber wrote their own replacement from the ground up called uReplicator. The amount of time wasted on replication broken across regions is disgusting. 3. Optimization & Scaling Kafka bundles compute & storage. There's (maybe on a upcoming KIP) no way that I know of splitting this. This means you'll waste time on Ops side deciding on tradeoffs between your broker throughput and your broker space. Worse yet time & money will be wasted here. I'd just rather hire more people than waste time on silly things like this. This is where I justify taking on the expense of client libs. 4. Segments vs Partitions The major time wasters are where you end up in a situation with the cluster utterly getting destroyed. It will happen, it isn't a question of if but a question of when or the company goes belly up and nobody cares. It's 3 AM, the producer is getting back pressure, you get a page and now have to deal with adding on write capacity to avoid a hot spot. Don't forget you can't just simply do a rebalancement in Kafka or you'll break the contract with every developer who has developed under the golden rule of, "Your partition order will always be the same". You'll successfully pay the cost of upgrading the entire cluster and then spending 3 days coming up with a solution to rebalance without making all your devs riot against you when you break that golden contract. RIP Kafka Having spent a couple of years dealing with Kafka I'm sorry to burst people's bubbles but is dead. Even Confluent doesn't have a good enough story these days to not switch to Pulsar, they're going to sell you on the same consulting bs, "We're more mature", "We've got better tooling.", "Better suppott"... Yes, of course, it has been in the open source community 5 years longer and the company has been also around longer for that time. Kafka is dead, long live Pulsar.
- cbartholomew 7y agoWe provide a SaaS offering of Apache Pulsar in AWS, Azure, and GCP: https://kafkaesque.io/ https://kafkaesque.io/
- mattboyle 7y agoI didnt find this when looking, thanks will take a deeper look.
- jjeaff 7y agoCool name. That's one of those company names that almost seems like someone thought it would make a good company name first and thought it was so fitting, they should build a company around it.
- cbartholomew 7y agoThanks!
- matteomerli 7y agoWe're close to release a new "officially supported" native Go client library: https://github.com/apache/pulsar-client-go https://github.com/apache/pulsar-client-go
- jgraettinger1 7y agoIf you're a Go shop, Gazette is worth a look (https://gazette.dev https://gazette.dev).