10 ms·
I've had the question for a while so I'll ask it here, maybe someone can help me. Suppose you modeled your domain with events and your stack is build on top of
by daddykotex 9y ago
I've had the question for a while so I'll ask it here, maybe someone can help me.
Suppose you modeled your domain with events and your stack is build on top of it. As stuff happens in your application, events are generated and appended to the stream. The stream is consumed by any number of consumers and awesome stuff is produced with it. The stream is persisted and you have all events starting from day 1.
Over time, things have changed and you have evolved your events to include some fields and deprecate others. You could do this without any downtime whatsoever by changing your events in a way that is backward compatible way.
What is the good approach to what I'd call a `replay`?
When you want to replay all events, the version of your apps that will consume the events may not know about the fields that were in the event for day one.
- zjaffee 9y agoThe best way I've seen it done, is to version all of your schemas and have a database that signals all the transformations needed to be done for any given schema version. That way, when your reading a particular event, you can query for all the operations needed to be done on such event and perform them.
- daddykotex 9y agoI'm also wondering: how one deal with changes to events when using KSQL?
- nehanarkhede 9y agoMind clarifying what you mean by changes to events? If you create a STREAM or TABLE using KSQL, it makes sure that they are kept updated with every single event that arrives on the source Kafka topics. That's what you'd expect in a true event-at-a-time Streaming SQL engine, which is what KSQL is.
- daddykotex 9y agoSuppose you build a STREAM or TABLE from a topic and assume a field in the event is `id`. Later on, you introduce an update to this event where where your replace `id` by `user_id`, how is KSQL reacting?
- coolio222 9y agoI encountered that problem. The ad hoc fix, was to have a version field in each event and functions that translate the old event into new event(s). The code that processes the events only processes events of the current version. If your old events had been denormalized this might result into repetition of events when splitted.
- daddykotex 9y agoOk, thanks for the pointer
- Joeri 9y agoTo add to that, you can treat it like you would schema migration on databases: implement v1 to v2, v2 to v3, etc... and replay the migrations in order to migrate from whatever version of the event is to the latest version. This allows keeping event migration code as immutable as the event versions it migrates between.
- stingraycharles 9y agoAs always in these types of the scenarios, the answer is: it depends. It depends on the amount of data you have. It depends upon how big the diversion from the original schema is. Etcetera. My personal philosophy is to always leave event data at rest alone: data is immutable, you don't convert it, and you treat it like a historical artifact. You version each event, but never convert it into a new version in the actual event store. Any version upgrades that should be applied are done when the event is read; this requires automated procedures to convery any event version N to another version N + 1, but having these kind of procedures in place is good practice anyway. Some might argue that doing this every time an event is read is a waste of CPU cycles, but in my experience this far outweights possible downsides of losing the actual event stored at that time in the past, and this type of data is accessed far less frequently than new event data.
- daddykotex 9y agoOk this makes sense. It matches what some others have been saying as well. Thanks
- stingraycharles 9y agoThere's a whole world out there about this kind of stuff. Take a look at CQRS and some of the posts by Greg Young; they're highly informative and one of the first people to really capture this way of dealing with data properly.
- _pmf_ 9y agoI wonder what happened to Greg Young. I think he had a book in the pipeline (which I was looking forward to), but as far as I can glean from social media, some burnout related stuff happened.
- dkersten 9y agoI suppose you can always trade those CPU cycles off against storage and cache the N+1 version (in a separate Kafka topic or elsewhere), so now reading the latest-version data is fast, yet you still retain the original data intact, at the expense of more storage. This does complicate the storage though, as you now have multiple days a sources, but nothing that can't be solved.
- assface 9y agoWhat you are referring to is called a "temporal database". Your specific example is called a "bitemporal database". https://en.wikipedia.org/wiki/Temporal_database https://en.wikipedia.org/wiki/Temporal_database https://en.wikipedia.org/wiki/Bitemporal_Modeling https://en.wikipedia.org/wiki/Bitemporal_Modeling
- sixdimensional 9y agoI've often wondered the same. In data warehousing, particularly the Kimball methodology, if descriptive attributes are missing from dimensions, for example, it is common to represent them using some standard value, like, "Not Applicable" or "Unknown" for string values. For integers, one might use -1 or similar. For dates it might be a specially chosen token date which means "Unknown Date" or "Missing Date". It doesn't solve the problem of truly unknown missing information, but it at least gives a standard practice for labeling it consistently. Think of trying to do analytics on something where the value is unknown?? Not too easy, but at least it is all in one bucket. Certainly, if past values can be derived, even if they were not created at the time the data was originally created, that is one way of "deriving" the past when it was missing. But, otherwise, I don't think there is any other way to make up for the lack of data/information.
- dm3 9y agoHere's a pretty good, even if a bit too verbose, explanation of various issues and solutions related to event versioning: https://leanpub.com/esversioning/read https://leanpub.com/esversioning/read. The text is written by Greg Young - the lead on the EventStore [1] project. [1] https://geteventstore.com/ https://geteventstore.com/
- daddykotex 9y agothank you
- bamazizi 9y agoI highly recommend you look into gRPC. Building apps using event sourcing, CQRS and microservices can easily become hell if the data models are not thought through.
- _pmf_ 9y agoI thought RPC and CQRS are diametrically opposite patterns. (Although you can use RPC in a CQRS context, but only as a transport/encapsulation layer (so the response says "request queued" or "error", but does not divulge domain specific information ("item created", "item not created")).
- linkmotif 9y agoRecently asked on the Kafka users mailing list https://lists.apache.org/thread.html/82692004eb2292e1240c33968e2fdbf1fde3bb8a2dd411c92049e1bd@%3Cusers.kafka.apache.org%3E https://lists.apache.org/thread.html/82692004eb2292e1240c339...
- daddykotex 9y agoThanks for sharing, it makes me realize that our messages are not independent.
- 1_2__4 9y agoDon't persist the stream. The problem gets a lot easier if you stop thinking of a message bus as a data store.
- daddykotex 9y agoHow do you do a replay if you don't keep the event somewhere? I did not mean that messages were to be stored in the bus?
- manigandham 9y agoIf the fields are changing then you effectively have DDL and migrations in your code already... so decouple them and version the schema officially. Then record these schema changes as events in the same event stream. Build a view on against these schema change events as a table of schema version by timestamp to allow for parsing any arbitrary event.
- vdm 9y agohttps://martin.kleppmann.com/2012/12/05/schema-evolution-in-avro-protocol-buffers-thrift.html https://martin.kleppmann.com/2012/12/05/schema-evolution-in-... https://www.safaribooksonline.com/library/view/designing-data-intensive-applications/9781491903063/ch04.html#ch_encoding https://www.safaribooksonline.com/library/view/designing-dat...