2 ms·
What was the motivation for building your own instead of using something like DuckDB which support parquet out of the box? What are the differences?
by bfm 4y ago
What was the motivation for building your own instead of using something like DuckDB which support parquet out of the box? What are the differences?
- eatonphil 4y agoIn addition to other reasons they may have, an embedded database for Go being built in Go means they don't need to require CGO, which DuckDB in Go would require.
- brancz 4y agoThis was definitely part of the motivation, but even more importantly (and it's possible that we missed it in the duckdb documentation when we explored doing exactly this) we needed the ability to add columns dynamically when we see a new label-name. This is sort of an analogy of wide-columns in Cassandra, but forcing it into the columnar layout to allow it to be searched and aggregated by efficiently. From our research all the open source column databases at best support a map type, through which we loose the columnar layout since the values of the map are all stored together giving us row-based database characteristics. (all databases except InfluxDB IOx, whose developers we talked to extensively and who highly inspired this design)
- bfm 4y agoIt makes sense, DuckDB's documentation has significantly improved, but it is still lacking when it comes to the limitations of using parquet. We have also hit some roadblocks when updating schemas backed by parquet files, so we now only use DuckDB for querying parquet via SQL.