4 ms·
Spark structured streaming is in there under structured, high temporal locality. It didn't make it into https://scattered-thoughts.net/writing/internal-consist
by jamii 5y ago
Spark structured streaming is in there under structured, high temporal locality.
It didn't make it into https://scattered-thoughts.net/writing/internal-consistency-in-streaming-systems/ https://scattered-thoughts.net/writing/internal-consistency-... because it has severe limitations for low temporal locality operations:
* As of Spark 2.4, you can use joins only when the query is in Append output mode. Other output modes are not yet supported.
* As of Spark 2.4, you cannot use other non-map-like operations before joins. Here are a few examples of what cannot be used.
* Cannot use streaming aggregations before joins.
* Cannot use mapGroupsWithState and flatMapGroupsWithState in Update mode before joins.
* There are a few DataFrame/Dataset operations that are not supported with streaming DataFrames/Datasets. Some of them are as follows.
* Multiple streaming aggregations (i.e. a chain of aggregations on a streaming DF) are not yet supported on streaming Datasets.
* Limit and take the first N rows are not supported on streaming Datasets.
* Distinct operations on streaming Datasets are not supported.
* Sorting operations are supported on streaming Datasets only after an aggregation and in Complete Output Mode.
* Few types of outer joins on streaming Datasets are not supported.
I haven't looked into the implementation but I'm guessing they just don't have good support for retractions in aggregates/joins. I also think it's likely to at least fall afoul of early emission and confusing changes with corrections.
The rest of spark doesn't really fit the post at all afaict - there is no streaming or incremental update, just batch stuff.