4 ms·
I'd be curious to hear if anyone can speak to how this compares to: - Citus columnar storage for Postgres, now owned by Microsoft - Amazon Aurora PostgreSQL c
by drewda 4y ago
I'd be curious to hear if anyone can speak to how this compares to:
- Citus columnar storage for Postgres, now owned by Microsoft
- Amazon Aurora PostgreSQL compatibility layer
- Timescale hypertables
- mattashii 4y agoboth of the mentioned features from Citus and Timescale are extensions on PostgreSQL, so they do not alter the way storage works in PostgreSQL in any meaningful manner. AlloyDB Seems better compared to Aurora, and they seem to share a significant amount of design principles with log-shipping, but now with an extra caching layer at or near the database instance. As for the 'columnar' features, from the release post it seems like AlloyDB does not specifically store the data at rest in a columnar format, but transforms the rows into columnar for analytical queries (but I'm not sure about that). That indeed improves performance of certain classes of queries.
- manigandham 4y agoBoth Citus and Timescale offer different storage (columnar) storage layouts.
- manigandham 4y agoCitus and Timescale are postgres extensions that focus on partitioning data primarily. Citus focused on scaling out across multiple nodes based on whatever primary key you wanted. Timescale focused on single-node and then added multiple nodes later, focused on time (or other integer-based values) and more utilities around time-based analysis. The Citus team also had a columnar data storage extension that's finally more production ready and Timescale created their own implementation by using columnar data but still storing it in the default rowstore and handling the differences in the query layer. Aurora Postgres and AlloyDB are fundamentally the same thing and involve taking the "top" portion of actual PostgreSQL (wire protocol, parser, query planner, etc) and attaching it to their own rebuilt storage layer. Since storage is the bottleneck, they can scale that out using their cloud architecture and make it seamless to the DB compute layer on top. Other open-source databases like Yugabyte also follow this approach with their own data layer implementation to add distribution and replication. There are still some fundamental limitations with this approach which is why you still have the concept of instances and primary/replicas instead of one massive distributed instance like Spanner, but most customers want faster/scalable Postgres rather than shifting to a proprietary DB.