5 ms·
This is a fantastic comment, thank you for it. > Kafka publishers habit of lying to their clients I have no Kafka expertise but it feels like there ought to b
by soamv 6y ago
This is a fantastic comment, thank you for it.
> Kafka publishers habit of lying to their clients
I have no Kafka expertise but it feels like there ought to be a Kafka configuration option somewhere that tells it not to do that? (And in fact changing the batch size to 1 won't help if it continues ack'ing messages without syncing to disk)
- zbentley 6y ago> changing the batch size to 1 won't help if it continues ack'ing messages without syncing to disk Sorry, I may not have been clear. For many Kafka clients, setting the batchsize to 1/interval to 0 causes every publish operation to block on the batch being flushed internally. Flushes are acknowledged by Kafka (and, like any RPC, can succeed or fail), and their success indicates that the persistence system was engaged, indicating disk persistence to a configured number of disks. Other Kafka clients provide a manual "flush" operation which acts similarly. More information about, and discussion of this phenomenon in the Python Kafka client can be found here: https://github.com/confluentinc/confluent-kafka-python/issues/137 https://github.com/confluentinc/confluent-kafka-python/issue... The deceptive behavior only happens when a publish operation does not entail a flush. At that point, unsuspecting client code may have assume that data left over the network when it did not. RabbitMQ's behavior with publisher confirms disabled is similar but not identical. Even with pub confirms off, most RabbitMQ clients "publish" operations are synchronous (as in they return after data has been written out over a socket). Publisher confirms are an added layer of resilience that enables clients to listen to RabbitMQ saying "I have successfully routed (and, depending on settings, persisted to disk) these messages". That acknowledgement may never come (rabbit may be overloaded, crashing, or may reject a publish), but on most happy-path systems data loss is not severe even without publisher confirms--turning them on is there to fill the non-happy-path case, and also to provide a very primitive system of backpressure to publishers, forcing them to wait on an overloaded broker rather than sending it even more data when e.g. its persister is slowing down. Because of that behavior, I'd argue that even without publisher confirms, RabbitMQ's publish behavior is still less data-loss prone than Kafka's. Regardless, the only way to run either system in truly reliable "publish means my data is on the broker" mode is to set the bufsize/flush interval to the minimums (in Kafka's case) or to wait for a confirmation after every single publish (in Rabbit's). As with all strategies to maximize reliability, those both come with a performance cost, and shouldn't be blindly adopted unless you have a good understanding of how much data and performance loss is acceptable.