3 ms·
Disclaimer: I work for trivago and I am partially responsible for the architecture behind trivago's Kafka use. I'd like to invite you to watch this talk and ge
by xenji_fm 8y ago
Disclaimer: I work for trivago and I am partially responsible for the architecture behind trivago's Kafka use.
I'd like to invite you to watch this talk and get a few more insights why we use Kafka in the ways we do: https://www.youtube.com/watch?v=cU0BCVl4bjo https://www.youtube.com/watch?v=cU0BCVl4bjo
Let me try to come up with a TL;DR here:
trivago comes from a complete on-premise, central database point of view. Change Data Capture via Debezium into Kafka enables a lot of migration strategies into different directions (e.g. Cloud) in the first place, while not having the need to change everything on the spot.
It seems like a common pattern to compare Kafka with a pure MQ technolgy. Kafka can also serve as a persistent data storage and a source of truth for data.
I hope this makes the picture a bit more clear to you. Feel free to ask if I missed something.
- deleted 8y ago[deleted]
- wenc 8y agoThanks for heads up about Debezium. That is very cool. Debezium seems to be a production version of Martin Kleppmann's CDC-to-Kafka POC, Bottled Water [1]. Database replication is the killer app for CDC, but CDC can be used for so much more than replication, like event-based alerting, triggering, etc. [1] https://www.confluent.io/blog/bottled-water-real-time-integration-of-postgresql-and-kafka/ https://www.confluent.io/blog/bottled-water-real-time-integr...
- gunnarmorling 8y ago(Disclaimer: Debezium lead here) While the basic idea of using PG logical decoding for CDC is the same, Debezium is a completely different code base than Bottled Water. Also we provide connectors for a variety of databases (MySQL, Postgres, MongoDB; Oracle and MongoDB connectors are in the workings). If you like, you can also use Debezium independently of Kafka by embedding it as a library into your own application, e.g. if you don't need to persist change events or want to connect it to other streaming solutions than Kafka. In terms of CDC use cases, I keep seeing more and more the longer I work on it. Besides replication e.g. updates of full-text search indexes and caches, propagatating data between microservices, facilitating the extraction of microservices from monoliths (by streaming changes from writes to the old monoliths to new microservices), maintaining read models in CQRS architectures, life-updating UIs (by streaming data changes to Web Sockets clients) etc. I touch on a few in my Debezium talk (https://speakerdeck.com/gunnarmorling/data-streaming-for-microservices-using-debezium?slide=5 https://speakerdeck.com/gunnarmorling/data-streaming-for-mic...).
- abledon 8y agoThis is why Hacker news is epic ... engineers reply includes technology, child reply is author of said technology
- wenc 8y agoThanks for the explanation. Any plans to support SQL Server? (SQL Server is prevalent in the enterprise world)
- gunnarmorling 8y agoYes, we're working on it right now (meant to say "Oracle and SQL Server are in the workings" above).
- ec109685 8y agoDid you consider running Kafka in both on-premise and in the cloud given it can handle the "persist data forever" use case.
- xenji_fm 8y agoWe have cases where we do this, yes. We create a cluster in the cloud mostly for traffic optimization between the cloud zones and the on-premise datacenters. The cost to value efficiency needs to be reconsidered on a case-by-case basis.