3 ms·
I really wonder what people actually use stream processing for, like very concrete examples. My best examples would only go far filtering a stream over a time w
by CSDude 5y ago
I really wonder what people actually use stream processing for, like very concrete examples. My best examples would only go far filtering a stream over a time window to compute an aggregate. My job does not require anything more, it's always basic ETL, but I really need to hear specific examples where it's useful for others. Been a long time fan of Apache Flink.
- thinkharderdev 5y agoStandard use cases for stream processing are: 1. Enriching event streams. Say you have a stream of log records with an IP address field. You want to enrich with a geo-location before sending the logs to Elasticsearch. 2. Windowed aggregation. Maybe you have an application that is emitting "login" events and you want to to detect login attempts from different IP addresses within X minutes of each other. 3. Joining multiple event streams. You have multiple different event streams and you want to join them together using some common join key (maybe session ID or something like that) to compute a metric that aggregates all of them. There are plenty of more esoteric use cases as well.
- CSDude 5y agoThey are the obvious ones, I'm looking for more advanced & specific ones.
- infinite8s 5y agoOperational analytics when dealing with real world systems (transportation logistics, tracking machine states on factory floors, sensor data fusion, real-time operational dashboards for capital markets, etc).
- Mehdi2277 5y agoStreaming model training for recommendation systems like tiktok, facebook, youtube, etc. If you want models to learn quickly to new events allowing them to do better on very new content/users you will want to use a streaming pipeline. I have seen flink used for model training at large scale here. Model evaluation can also be done in a streaming manner with training and needs to be done parallel for model monitoring. The lowest latency model refresh time I've come across is ~5 minutes. Going lower is likely to be a lot of data transfers for model synchronization (sync the serving model with training model) for little value.