3 ms·
[2] should be https://www.cockroachlabs.com/blog/how-we-built-a-vectorized-execution-engine/ https://www.cockroachlabs.com/blog/how-we-built-a-vectorized... Go
by vvern 6y ago
[2] should be https://www.cockroachlabs.com/blog/how-we-built-a-vectorized-execution-engine/ https://www.cockroachlabs.com/blog/how-we-built-a-vectorized...
Good note on the use of MVCC in Postgres. That on-access cleanup dramatically helps with the explosion of the data size and seems generally nice. Cockroach promises to keep data in a time window rather than just what's needed for concurrent operations. The user controls the time window but often people don't change it and even if they do, the build-up can be dramatic.
The sequential nature of the data storage and its implications for performance are pretty different in postgres and cockroach as I understand it. In cockroach we have fewer opportunities to exploit parallelism of write operations acting on the same "range" (a cockroach level concept for a raft group). The sequential nature of the workload is a problem not because of how the data ultimately gets laid out on disk but rather on how it gets processed and sequenced for replication. In particular, all of the writes will go to the same "range" which owns the tail of the log. Postgres, if anything, is happy with sequential workloads as they touch the fewest interior blocks of the B+-Tree.
Cockroach effectively can't offload the work of GC (as we call it) to foreground or already existing tasks mainly because it mains maintaining consistency stats between replicas hard. See an attempt here: https://github.com/cockroachdb/cockroach/pull/42514 https://github.com/cockroachdb/cockroach/pull/42514.