3 ms·
> The Iceberg table serves as the only schema definition for the ingestion process. In various alternative architectures, the schema is often stored in multiple
by hashhar 4y ago
> The Iceberg table serves as the only schema definition for the ingestion process. In various alternative architectures, the schema is often stored in multiple systems, requiring these disparate systems to be kept in sync with each other.
> Here, the continuously running Flink job parses the JSON data source directly following the Iceberg table schema definition. This means future schema changes can be facilitated through just one source of truth.
This is the most interesting point. I'd love to take this implementation out for a spin - could solve a lot of pain I know other people deal with too.
Most of the pain (or even grunt-work) of managing data pipelines is updating, validating and managing schemas.
In the past the common approach people suggested was to have each application write the data with the same schema but in practice it's never possible unless it's a greenfield project and the services don't need to talk to other external or pre-existing systems. What ends up happening is that a translator (or validator) service comes up whose job it is to translate across the various schemas. 2x storage in Kafka, 2x compute for the consumers and probably 20x more maintenance and ways things can go wrong.
- erichwang 4y agoMaking the schema management simpler was definitely one of the things we wanted to do here. We wanted it to be easy to set up and maintain, and schema inconsistency is one of those things that has been a continued sore spot from my past experiences. The only caveat here is that our post only discusses data that is already in Kafka for the most part. To have a complete solution, ideally the client submitting the structured event to Kafka would also recognize and enforce that same serialization schema -- which is a bit of harder problem to solve in a general sense, but there are ways to do this.