3 ms·
To state the obvious, Delta is an open standards format which should be widely supported. Databricks have also bought into Iceberg and will probably lead with
by benjaminwootton 2y ago
To state the obvious, Delta is an open standards format which should be widely supported.
Databricks have also bought into Iceberg and will probably lead with that or unify the two in future.
- kianN 2y agoThere is an open source Delta that is a very good library. This is not the same as Databricks' implementation and there are at times compatability issues. For example, by default if you write a delta table using Databricks' dbr runtimes, that table is not readable by the open source Delta because due to the "deletion vectors" optimization that is only accessible within Databricks. That aside, I was more pointing out that Delta, particularly via a commercial offering, is a data format biased towards Spark in terms of performance, since it is being developed primarily by Databricks as a part of the spark ecosystem. If you are plan to use Delta regardless of your compute engine, it makes perfect sense as a benchmark. However, for certain circumstances, the performance wins could be (in my case was) worth it to switch data formats.
- mwc360 2y agoOSS Delta support deletion vectors. The problem is that the OSS Deltalake (based on Delta-rs) python library does not and this prevents engines like DuckDB and Polars from writing to suck tables. I'm pretty sure DV is in OSS since 3.1