3 ms·
The new Structured Streaming API looks pretty interesting. I have the impression that many Apache projects are trying to address the problems that arise with th
by graffitici 10y ago
The new Structured Streaming API looks pretty interesting. I have the impression that many Apache projects are trying to address the problems that arise with the lambda architecture. When implementing such a system, you have to worry about dealing with two separate systems, one for low-latency stream processing, and the other is the batch-style processing of large amounts of data.
Samza and Storm mostly focus on streaming, while Spark and MapReduce traditionally deal with batch. Spark leverages its core competency of dealing with batch data, and treats streams like mini-batches, effectively treating everything as batch.
And I imagine in the following snippet, the author is referring to Apache Flink, among other projects:
> One school of thought is to treat everything like a stream; that is, adopt a single programming model integrating both batch and streaming data.
My understanding of Structured Streaming also treats everything like batch, but can recognize that the code is being applied to a stream, and do some optimizations for low-latency processing. Is this what's going on?
- rxin 10y agotl;dr is yes (to your last question). The longer answer is that this is about how to logically think about the semantics of computation using a declarative API, and the actual physical execution (e.g. incrementalization, record at a time processing, batching) is then handled by the optimizer.