3 ms·
>Kafka's publish behavior is even worse. The vast majority of Kafka publishers' default settings don't even send the data to the broker without waiting for a re
by synthc 6y ago
>Kafka's publish behavior is even worse. The vast majority of Kafka publishers' default settings don't even send the data to the broker without waiting for a response (like RabbitMQ without publish confirms); they send the data to a local buffer which is asynchronously flushed periodically/volumetrically.
When you publish a message to Kafka, it stores the message in a buffer and immediately returns a future that completes once the data is sent to the broker.
If you forget to wait on completion of the future that is on you, this behaviour is well documented.
- zbentley 6y ago> If you forget to wait on completion of the future that is on you Absolutely. However, data loss and edge cases emerge where process crashes and abrupt exits are concerned. With RabbitMQ, even with pub confirms disabled, the odds are that all of your publishes (sans, perhaps, the one you were in the middle of at the instant of exit) will make it to the broker. With Kafka, every message not yet flushed from the buffer will be dropped. I get that this is expected/documented, and I get why it exists (batchwise operations allow way higher net throughput into Kafka). However, I've encountered many groups of engineers from many companies who were surprised by and unable to effectively work around this behavior. Workarounds include performance sacrifices (flushing per-publish); complex signal handling semantics to coordinate flushes (complicated further by errors that occur if a signal arrives during flush); and atexit(3) flush hooks (terrifying, given the way some higher level programming languages' garbage collection interacts with atexit). At the end of the day, it matters how friendly a tool is when it's held by a novice. While there are many good reasons to use Kafka, that's not one of them.
- deleted 6y ago[deleted]