3 ms·
I'm working on an alternative Iceberg client to work better in write heavy use cases. Instead of many smaller files it writes on the same file until it's 1mb in
by simlevesque 2y ago
I'm working on an alternative Iceberg client to work better in write heavy use cases. Instead of many smaller files it writes on the same file until it's 1mb in size but it gives it a new name. Then I update the manifest to the new filename and checksum. I keep old files on disk for 60 seconds to allow pending queries. I'm also working on auto compaction, when I have ten 1mb files I compact them, same with ten 10mb files, etc...
I feel like this could be a game changer for the ecosystem. It's more cpu and network heavy for writes but the reads are always fast. And the writes are still faster than pyiceberg.
I want to hear opinions or how this could never work.
- mritchie712 2y agonice! anywhere we can follow your progress?
- simlevesque 2y agoNot right now sadly I have some work obligations taking my time but I can't wait to share more. I'm using a basic implementation that's not backed by iceberg, just Parquet files in hive partitions that I can query using DuckDB.
- thom 2y agoInteresting. My personal feeling is that we're slowly headed to a world where we can have our cake and eat it: fast bulk ingestion, fast OLAP, fast OLTP, low latency, all together in the same datastore. I'm hoping we just get to collapse whole complex data platforms into a single consistent store with great developer experience, and never look back.
- simlevesque 2y agoI think it's possible too and the Iceberg spec allows it but the implementations are not suited for every use case.
- ndm000 2y agoI’ve felt the same way. It’s so inefficient to have two patterns - OLAP and OLTP - both using SQL interfaces but requiring syncing between systems. There are some physical limits at play though. OLAP will always take less processing and disk usage if the data it needs is all right next to each other (columnar storage) where as OLTP’s need for fast writes usually means row based storage is more efficient. I think the solution would be one system that stores data consistently both ways and knows when to use which method for a given query.
- thom 2y agoIn a sense, OLAP is just a series of indexing strategies that takes OLTP data and formats it for particular use cases (sometimes with eventual consistency). Some of these indexing strategies in enterprises today involve building out entire bespoke platforms to extract and transform the data. Incremental view maintenance is a step in the right direction - tools like Materialize give you good performance to keep calculated data up to date, and also break out of the streaming world of only paying attention to recent data. But you need to close the loop and also be able to do massive crunchy queries on top of that. I have no doubt we'll get there, really exciting times.
- ndm000 2y agoCompletely agree. All of the pieces are there and it's just waiting to be acted upon. I haven't seen any of the major players really doubling down on this, but would be so compelling.
- gregw2 2y agoI'd love it, but I feel like there is another horizon that Iceberg hasn't tackled to truly get us there. Iceberg (and Delta Table format) is really OLAP-optimized, being built on a columnar datastore, Parquet. This means it will be slow to do writes compared to a traditional row-based datastore and doesn't really have normal/optimal OLTP indexing. Fast OLTP + Fast OLAP + low latency is best done via HTAP-type databases which store data in both row and columnar form and give you ability in the SELECT clause to pick your latency tolerance and the query engine will pick the OLTP engine if it knows there are still some OLAP writes queued up that entered the system more than <latency-timeframe> ago but aren't fully on disk yet. Various vendors do have HTAP, but all with proprietary storage engines and query engines. But Iceberg alone doesn't get you there. I haven't seen discussion of this I don't know if anyone has tried to write both Hudi and Iceberg/Delta in parallel so they could do HTAP; maybe they use pure Hudi instead? I'd have to re-look at Hudi to see if it's deferred compaction is more like this. XTable doesn't seem to target this issue.
- FuriouslyAdrift 2y agoso... sharding?
- chehai 2y agoThis approach reminds me of ClickHouse's MergeTree. Also, https://paimon.apache.org/ https://paimon.apache.org/ seems to be better for streaming use cases.