4 ms·
I like the idea of using “fat” events as a means of data integration between systems. Calling them RESTful events is a good term as I’ve pitched in much the sa
by adamkl 5y ago
I like the idea of using “fat” events as a means of data integration between systems.
Calling them RESTful events is a good term as I’ve pitched in much the same fashion. Every time a system responds to a create/update/execute request, drop an equivalent message into an event stream. Initially these can form the basis of a streaming ETL to whatever systems you use for analytics, but as new internal systems are brought online (e.g a new CRM system to help sales manage all the leads signing up for your successful product) the integration is already there, waiting. Just add a new consumer to the existing stream.
Overtime, you end up with a “data mesh”. If you look past the marketing of the term, it’s basically just means each application in your company should publish both sync (for real-time usage) and async (for data sharing between systems of record) interfaces.
- hashimotonomora 5y agoCouldn’t that bring inconsistencies eventually? If one commits and the other doesn’t.
- themoop 5y agoThis is where you might want to use the outbox pattern. If you depend on dual writes you will definitly have consistency issues at some point
- Gwypaas 5y agoThe question at hand will always reduce down to the "Two Generals problem" [0]. The outbox pattern is a nice separation of concerns and gives at least once consistency as long as the forwarding happen before writing the event as processed in the outbox. Both side effects happening atomically through all failure cases is impossible. To solve that issue you need a cooperating destination for your writes which handles de-duplication. Then you get into the weeds of causality, ordering and idempotence. Ugh. For some real world examples see the AWS Kinesis documentation which says that any application using Kinesis must be able to handle duplicate records [1]. > There are two primary reasons why records may be delivered more than one time to your Amazon Kinesis Data Streams application: producer retries and consumer retries. Your application must anticipate and appropriately handle processing individual records multiple times. [0]: https://en.wikipedia.org/wiki/Two_Generals%27_Problem https://en.wikipedia.org/wiki/Two_Generals%27_Problem [1]: https://docs.aws.amazon.com/streams/latest/dev/kinesis-record-processor-duplicates.html https://docs.aws.amazon.com/streams/latest/dev/kinesis-recor...
- frankdejonge 5y agoFor anyone who's interested in the outbox pattern, I also blogged about that: https://blog.frankdejonge.nl/reliable-event-dispatching-using-a-transactional-outbox/ https://blog.frankdejonge.nl/reliable-event-dispatching-usin...
- adamkl 5y agoIt is possibly that you could get inconsistencies between systems, but you mitigate that with durable event streams (e.g. Kafka) and by ensuring that each data entity in your enterprise has a system of record that can be used to resolve conflicts. This is really not that different than using ETL jobs to move data about, but taking a streaming approach vs a batch approach.
- majormajor 5y ago> It is possibly that you could get inconsistencies between systems, but you mitigate that with durable event streams (e.g. Kafka) and by ensuring that each data entity in your enterprise has a system of record that can be used to resolve conflicts. The event stream layer isn't where the sync problems have arised in the systems I've worked on. It's the "commit transactionally both to your database and the event stream" part. Not a lot of systems are built to be ready to roll back the database change if the event publish fails. Or to be able to handle duplicate events if you error the other way.
- adamkl 5y agoI guess if the concern is around atomically committing to the DB and event stream at the same time, you could hang the event stream off the DB using change-data-capture to populate the events. Ultimately, whenever you are pushing data between systems you can end up with inconsistencies which why it’s important to clearly define systems of record.
- majormajor 5y agoThat's basically what the database would do anyway, so generally, yes (e.g. even with a RDBMS it's internally got some sort of write log, in most cases). That's gonna help a lot! It doesn't give you any guarantees about "at a given time" consistency, though. Maybe your queue or your event processor got backed up. So how real-time do you need both sides? These are, of course, problems that people have solved to varying level of satisfaction for most use cases! But I've seen systems built by people who saw the blog post about "have both!" and then didn't think about all these cases and then... it gets messy.
- pc86 5y agoTypically you see these suggested as part of a "change data capture" type process where the event is only published after the action is committed to the data store. The downside (IMO) is this required directly integrating with the data store itself which isn't always easy to do, or obvious from a git/CICD perspective.
- jhgb 5y ago> Every time a system responds to a create/update/execute request, drop an equivalent message into an event stream. Don't we call that just "a WAL"?
- klysm 5y agoWrite ahead log is something with more specific and semantics. You write something in a WAL _ahead_ of some process for durability sake.
- jhgb 5y agoPossibly, but it records the complete information about actions to be performed on some system, which presumably is what the "fat events" were supposed to be about (unless I misinterpreted it somehow).
- morelisp 5y agoTo the extent a fat event is a "RESTful" event, it should definitely not be about the actions but about the current state of things. Now, maybe it only emits when that state has changed, but that's distinct from what's in the message. What exactly a WAL contains depends on the specific storage, but it's commands at least as often as it's state - probably more often since you want your WALs to be tiny and they often also facilitate rollback. In general you wouldn't want to emit a fat event until the data is firmly committed to its system of record which is usually not the event system itself; in that use case it will be a kind of 'write behind' log.