3 ms·
The point is, you can’t do synchronization reliably using any message/event-based approach alone. You always need a reconciliation mechanism, as you’ve correctl
by halfcat 1y ago
The point is, you can’t do synchronization reliably using any message/event-based approach alone. You always need a reconciliation mechanism, as you’ve correctly noticed.
It’s always shocking to me how many FAANG people will say, “we want an event driven solution with a message bus, that the right way to do it, we don’t want batch, that can’t scale”, and then need to bolt on a validation/reconciliation step for it to be reliable. Which of course is a batch job.
Unless you control a system end to end (which is rare, there’s usually some data from a system you don’t control the schema of), or are highly incentivized to make sync happen reliably (e.g. bitcoin), you’re always limited by a batch job somewhere in the system.
You could even say the batch job is often doing the heavy lifting, and the message bus is just an optimization (that’s often not adding the value needed to justify the complexity).
- UltraSane 1y agoUsing something like Kafka you can get reliable at least once messaging and then you just need to make the CDC updates idempotent. I'm not sure exactly what you mean by batch job but if n bytes change at the source you shouldn't have to copy more than n bytes to the destination.
- halfcat 1y agoIf you can get all of your data efficiently (and transactionally consistent) into Kafka, that’s the scenario I mention where you have control of all your systems. Inevitably, even if you achieve this at some point, it never lasts. Your company acquires another, or someone pushes for a different HR/CRM/whatever system and gets it. You mention if n bytes change in the source, but many systems have no mechanism of determining that n bytes have changed without scanning the entire data set. So we’re back to a batch job (cron, or similar).
- UltraSane 1y ago"many systems have no mechanism of determining that n bytes have changed without scanning the entire data set." This is so insanely inefficient it can't scale to very large amounts of data. If you can't do data syncing at the application layer you can do it at the storage layer with high end storage arrays that duplicate all writes to a second storage array, either synchronously or asynchronously. Or duplicate snapshots to a remote array. They work really well.