3 ms·
No mention of Apache Spark? Thinking about where Spark would fit in here, because it has both structured and unstructured APIs--I think it would fall under mos
by munro 5y ago
No mention of Apache Spark? Thinking about where Spark would fit in here, because it has both structured and unstructured APIs--I think it would fall under most of these categories depending how you use, going back to the comment in the article:
> Most of these systems are equally expressive and can emulate each other
But a quick stab at how I think it could fall under this taxonomy:
[structured]
If using the DataFrame API
[unstructured]
If using the DStream API
---
[low temporal locality]
If you don't watermark and let the state grow--which I don't have much experience with,
but I think could possibly run into scalability issues, you'd have to write some
imperative logic to cull state--but it's very necessary for low temporal locality.
[high temporal locality]
If using watermarks, it will close the and take the output for you
---
[internally consistent / internally inconsistent]
Not really sure how Spark falls here, I think if you use the DataFrame API, you can have
it be consistent, but a lot of operations aren't supported, so you may end up having to
switch to DStreams and write code imperatively.
And then further, what if one of the streams is processed faster than each the other in
an aggregation? I'm not sure if there's a way to specific business rules around making
things atomic--but I'm sure you could hack it if you dropped low level enough.
I think this is touched upon in the other article:
> When combining multiple streams it's important to synchronize them so that the outputs
> of each reflect the same set of inputs.
If someone has more experience with Spark [Structured] Streaming, I would love to hear your thoughts. I stick mostly with Spark's batch jobs, playing around with datasets in a REPL/Notebook--which it really shines for.
That said, I'm really excited by the future of Spark Structured Streaming--I want to write declarative code and say how my data gets from A to B--not what needs to happen to get from A to B.
- jamii 5y agoSpark structured streaming is in there under structured, high temporal locality. It didn't make it into https://scattered-thoughts.net/writing/internal-consistency-in-streaming-systems/ https://scattered-thoughts.net/writing/internal-consistency-... because it has severe limitations for low temporal locality operations: * As of Spark 2.4, you can use joins only when the query is in Append output mode. Other output modes are not yet supported. * As of Spark 2.4, you cannot use other non-map-like operations before joins. Here are a few examples of what cannot be used. * Cannot use streaming aggregations before joins. * Cannot use mapGroupsWithState and flatMapGroupsWithState in Update mode before joins. * There are a few DataFrame/Dataset operations that are not supported with streaming DataFrames/Datasets. Some of them are as follows. * Multiple streaming aggregations (i.e. a chain of aggregations on a streaming DF) are not yet supported on streaming Datasets. * Limit and take the first N rows are not supported on streaming Datasets. * Distinct operations on streaming Datasets are not supported. * Sorting operations are supported on streaming Datasets only after an aggregation and in Complete Output Mode. * Few types of outer joins on streaming Datasets are not supported. I haven't looked into the implementation but I'm guessing they just don't have good support for retractions in aggregates/joins. I also think it's likely to at least fall afoul of early emission and confusing changes with corrections. The rest of spark doesn't really fit the post at all afaict - there is no streaming or incremental update, just batch stuff.