2 ms·
> why Flink and not Kafka ? I assume you mean "Kafka Streams" when you say this, as Kafka is just a event bus and Flink can be used to read messages from it.
by PhoenixReborn 5y ago
> why Flink and not Kafka ?
I assume you mean "Kafka Streams" when you say this, as Kafka is just a event bus and Flink can be used to read messages from it.
The biggest advantage of Flink (IMO) is that you can write the Flink logic once, and reuse it for both batch and stream processing. So if I write a Flink job that consumes a Kafka stream and produces some aggregated outputs, that same job can be run against my data lake in S3/GCS/Azure Blob Storage, etc.
Kafka Streams does not support batch processing, or working on top of anything other than Kafka. Flink supports building on top of other message buses like Pulsar as well: https://flink.apache.org/news/2019/11/25/query-pulsar-streams-using-apache-flink.html https://flink.apache.org/news/2019/11/25/query-pulsar-stream...
- dikei 5y agoActually, unified batch and stream processing in Flink is a bit of false advertisement. In Flink, stream and batch have different API (DataStream vs DataSet) so reusing logic is not really practical, unless you abstract out everything, in which case you might as well use Spark which is faster than DataSet API. Flink's developers is trying to get rid of DataSet API and move to DataStream and Table API for everything, but it's not done yet.