3 ms·
This is an excellent point, there IS a fundamental tradeoff between latency and throughput. Computers are much better a processing larger chunks of data linearl
by boredandroid 12y ago
This is an excellent point, there IS a fundamental tradeoff between latency and throughput. Computers are much better a processing larger chunks of data linearly rather than small bits of data. However the question to ask is how much data do you have to batch together to get good throughput? There are diminishing returns and you stop getting much benefit after about 1MB in my experience. So this line of reasoning will not get you to additional latency of more than a few seconds, it definitely won't get you to a 24 hour batch data cycle.
Actually because we think this tradeoff is important we allow consumers from Kafka to specify how much data they want accumulated on the server before the server should complete their fetch request. The default is 1 byte, which means respond as soon as there is anything new for me, which optimizes latency. However setting this to a higher number will reduce round-trips and optimize throughput by avoiding lots of small fetches. This allows you to trade a small amount of latency for throughput. (In either case you can bound the waiting with a timeout so the latency is never worse than some maximum delay even if sufficient data hasn't arrived).
In any case when doing re-processing, as described in this article, there will always be lots of data accumulated and you will always fetch chunks of the maximum size you have configured (say 1MB). So in the reprocessing case you always get the "batch"-like throughput.