4 ms·
This depends on use case. SQL is the king for batching process - queries are declarative, decades of effort put into optimization. For real-time / streaming us
by ptrik 4y ago
This depends on use case. SQL is the king for batching process - queries are declarative, decades of effort put into optimization.
For real-time / streaming use cases, however, there is yet a mature solution in SQL yet. Flink SQL / Materialize is getting there, but the state-of-the-art approach is still Flink / Kafka Streams approach - put your state in memory / on local disk, and mutate it as you consume messages.
This actually echoes the "Operate on data where it resides" principle in the article.
- fifilura 4y agoDo you find flink SQL immature? For me it looks a lot like syntactical sugar on top of the datastream api. Same thing, less code?
- dagss 4y agoWe do mini-batch processing in SQL. Some hundred milliseconds latency, some hundred events consumed from to the inbound event table per iteration. Paginate through using a (Shard, EventSequenceNumber) key; writers to table synchronize/lock so that this is safe. Kafka-in-SQL if you wish. Or, homegrown Flink. (There are many different uses for the events inside our SQL processing pipelines, and have to store the ingested events in SQL anyway) I am sure real Kafka+Flink has some advantages, but...what we do works really well, is simple, and feels right for our scale. It is enough batching in SQL to real speed/CPU benefits on inserts/updates into SQL (vs e.g. hitting SQL once per consumed event which would be way worse). And with Azure SQL the infra is extremely simple vs getting a Kafka cluster in our context.