3 ms·
I have written Flink pipelines that solved similar problems to this post: sessionization of time-skew data points from multiple sources with variable and possib
by dikei 4y ago
I have written Flink pipelines that solved similar problems to this post: sessionization of time-skew data points from multiple sources with variable and possibly large delays; simple and economical is not how I would describe it :) .
If one of your datasource get lag behind, Flink would buffer a huge amount of data waiting for it to catch up due to how watermark work when joining 2 streams, and you would still encounter out-of-memory error even with RocksDB from time to time if your session window get too large. In addition, with our state size frequently reached hundred of GBs, recovering from failure was not exactly fast either.
- DeathArrow 4y agoAs I understood, they wanted to eliminate the MQ layer. If they used Flink, they would still need to keep Kafka and they would be introducing another layer of complexity with Flink. So, instead of simplifying they would make the stack more complex.
- skinnyarms 4y agoThis, especially when some records can be "several megabytes"