3 ms·
The question whether we need distributed stream processing can hardly be answered unambiguously. In general, everything depends on what features we actually nee
by asavinov 8y ago
The question whether we need distributed stream processing can hardly be answered unambiguously. In general, everything depends on what features we actually need: scalability, fault tolerance, high availability, performance (latency or throughput) etc.
In any case, there are two ways to distribute stream processing tasks:
* [~Heavyweight] Use a cluster with multiple runners like [Spark Streaming](https://spark.apache.org/streaming/ https://spark.apache.org/streaming/) or [Flink](https://flink.apache.org/ https://flink.apache.org/)
* [~Lightweight] Do stream processing in applications, at the edge, on gateways or devices like [Bistro Streams](https://github.com/asavinov/bistro/tree/master/server https://github.com/asavinov/bistro/tree/master/server) or [Kafka Streams](https://kafka.apache.org/documentation/streams/ https://kafka.apache.org/documentation/streams/)
Normally, distributed stream processing requires also partitioning the data, e.g., by user or session ids, so that these isolated streams can be processed independently at different nodes.