3 ms·
This is essentially the idea behind CDC (Change Data Capture) [0]. Martin Kleppmann has some great blogs about this as well [1]. [0] https://en.wikipedia.org/
by cfors 5y ago
This is essentially the idea behind CDC (Change Data Capture) [0].
Martin Kleppmann has some great blogs about this as well [1].
[0] https://en.wikipedia.org/wiki/Change_data_capture https://en.wikipedia.org/wiki/Change_data_capture
[1] https://www.confluent.io/blog/turning-the-database-inside-out-with-apache-samza/ https://www.confluent.io/blog/turning-the-database-inside-ou...
- kmdupree 5y agoThanks for sharing these links!
- gunnarmorling 5y agoIf you look for a ready-to-use open-source implementation of CDC for Postgres (and other databases), take a look at Debezium [1]. On audit logs in particular, we have a post on our blog, which discusses how to use CDC for that, focusing in particular on enriching change events with additional metadata like business user performing a given change by means of stream processing [2]. One advantage of log-based CDC over trigger-based approaches is that it doesn't impact latency of transactions, as it runs fully asynchronously from writing transactions, reading changes from the WAL of the database. Disclaimer: I work on Debezium [1] debezium.io [2] https://debezium.io/blog/2019/10/01/audit-logs-with-change-data-capture-and-stream-processing/ https://debezium.io/blog/2019/10/01/audit-logs-with-change-d...
- djbusby 5y agoThat is very cool.
- tomnipotent 5y agoDebezium is awesome, thanks for the great work! It's in my toolbox for when embulk [1] batch processing doesn't cut it (or even in combination with). https://github.com/embulk/embulk https://github.com/embulk/embulk
- gunnarmorling 5y agoThank you so much, it's always awesome to hear that kind of feedback!
- tomhallett 5y agoGunnar, I've been thinking about your "transaction" pattern more, and I was wondering - Why don't you stream the transaction metadata directly to a kafka "transaction_context_data" topic? Would writing some of the data to the db while writing some of the data directly to kafka make it less consistent during faults? The reason I ask: I'm curious what it would look like to use this pattern for an entire application, I think it could be a very powerful way todo "event sourcing lite" while still working with traditional ORMs. Would writing an additional transaction_metadata row for most of the application insert/updates slow things down? Too many writes to that table?
- gunnarmorling 5y ago> Would writing some of the data to the db while writing some of the data directly to kafka make it less consistent during faults? Yes, this kind of "dual writes" are prone to inconsistencies. If either the DB transaction rolls back, or the Kafka write fails, you'll end up with inconsistent data. Discussing this in a larger context in the outbox pattern post [1]. This sort of issue is avoided when writing the metadata to a separate table as part of the DB transaction, which either will be committed or rolled back as one atomic unit. > Would writing an additional transaction_metadata row for most of the application insert/updates slow things down? You'd just do one insert into the metadata table per transaction. As change events themselves contain the transaction id, and that metadata table is keyed by transaction id, you can correlate the events downstream. So the overhead depends on how many operations you in your transactions already. Assuming you don't do just a single insert or update, but rather some select(s) and then some writes, the additional insert into the transaction metadata table typically shouldn't make a substantial difference. Another option, exclusive to Postgres, would be to use write an event for the transaction metadata solely to the WAL using pg_logical_emit_message(), i.e. it won't be materialized in any table. It still can be picked up via logical decoding, we still need to add support for that record type to Debezium though (contributions welcome :). [1] https://debezium.io/blog/2019/02/19/reliable-microservices-data-exchange-with-the-outbox-pattern/ https://debezium.io/blog/2019/02/19/reliable-microservices-d...
- krageon 5y ago> kafka Except you already have postgres, why add another thing to it?
- deleted 5y ago[deleted]