4 ms·
The current system diagram implies using duckdb's default storage format directly. I wonder how well this would actually work with the proposed zero-ETL design
by robertclaus 2y ago
The current system diagram implies using duckdb's default storage format directly. I wonder how well this would actually work with the proposed zero-ETL design of basically treating this as a live replica. I was under the impression that as an OLAP solution DuckDB makes performance compromises when it comes to writes - so wouldn't live replication become problematic?
- fanyang01 2y agoGreat question! Improving the speed of writing updates to DuckDB has been a significant focus for us. Early in the project, we identified that DuckDB is quite slow for single-row writes, as discussed in this issue: https://github.com/apecloud/myduckserver/issues/55 https://github.com/apecloud/myduckserver/issues/55 To address this, we implemented an Arrow-based columnar buffer that accumulates updates from the primary server. This buffer is flushed into DuckDB at fixed intervals (currently every 200ms) or when it exceeds a certain size threshold. This approach significantly reduces DuckDB's write overhead. Additionally, we developed dedicated replication message parsers that write directly to the Arrow buffer, minimizing allocations. A community contributor has validated & enhanced the effectiveness of our approach: https://github.com/apecloud/myduckserver/pull/207 https://github.com/apecloud/myduckserver/pull/207 We plan to publish benchmarks on replication latency and query performance in the coming weeks. Stay tuned!