3 ms·
Yes, we are using it in production (somewhat) - for a a few tables. Our workflow with it is: 1. Stream new data into a partitioned S3 bucket available as a ta
by hashhar 6y ago
Yes, we are using it in production (somewhat) - for a a few tables.
Our workflow with it is:
1. Stream new data into a partitioned S3 bucket available as a table in metastore.
2. Use Hudi to create a copy-on-write compacted version of the table (think SELECT * FROM (SELECT *, row_number() over (partition by uuid order by ts desc) as rnum) WHERE rnum=1).
Now people who are interested on the snapshot of the table rather than the event log can query the Hudi table instead of the raw event log tables.
This has improved perf for the common use-cases while allowing us to use a common pipeline. Earlier we would have to dump such data into an RDS (Aurora Postgres) with UPSERT but that meant table bloat was an issue and cost was high depending on the date-range we wanted to store compacted.