4 ms·
This Google article was nice as a high level overview of Iceberg V3. I wish that the V3 spec (and Iceberg specs in general) were more readable. For now the best
by hodgesrm 1y ago
This Google article was nice as a high level overview of Iceberg V3. I wish that the V3 spec (and Iceberg specs in general) were more readable. For now the best approach seems to be read the Javadoc for the Iceberg Java API. [0]
[0] https://javadoc.io/doc/org.apache.iceberg/iceberg-api/latest/index.html https://javadoc.io/doc/org.apache.iceberg/iceberg-api/latest...
- twoodfin 1y agoThe Iceberg spec is a model of clarity and simplicity compared to the (constantly in flux via Databricks commits…) Delta protocol spec: https://github.com/delta-io/delta/blob/master/PROTOCOL.md https://github.com/delta-io/delta/blob/master/PROTOCOL.md
- eatonphil 1y agoTo the contrary, the Delta Lake paper is extremely easy to read and implement the basics of (I did) and Iceberg has nothing so concise and clear.
- twoodfin 1y agoIf I implement what’s described in the Delta Lake paper, will I be able to query and update arbitrary Delta Lake tables as populated by Databricks in 2025? (Would be genuinely excited if the answer is yes.)
- eatonphil 1y agoNot sure (probably not). But it's definitely much easier to immediately understand IMO.
- twoodfin 1y agoOK, but at least from my perspective, the point of OTF’s is to allow ongoing interoperability between query and update engines. A “standard” getting semi-monthly updates via random Databricks-affiliated GitHub accounts doesn’t really fit that bill. Look at something like this: https://github.com/delta-io/delta/blob/master/PROTOCOL.md#writer-requirements-for-clustered-table https://github.com/delta-io/delta/blob/master/PROTOCOL.md#wr... Ouch.