4 ms·
This sounds very exciting, but I would have greatly appreciated some actual technical details. From this statement > Sharded tables – These tables are distribu
by dkhenry 3y ago
This sounds very exciting, but I would have greatly appreciated some actual technical details. From this statement
> Sharded tables – These tables are distributed across multiple shards. Data is split among the shards based on the values of designated columns in the table, called shard keys.
It sounds like this is very much managed CitusDB on top of Aurora, but without any details about the implementation its impossible to know if its just repackaged Citus, or some new novel technology built specifically for Aurora.
- bdcravens 3y agoLikely it'll be talked about in-depth this week at Reinvent. This intro article is par for AWS "product announcements".
- mjb 3y agoAurora Limitless Database is based on our own investments in database-optimized virtualization (Caspian), in scale-out log-first database storage (Grover)[2], and in a custom approach to cross-shard transactions that makes use of the high-quality hardware clocks available in EC2[1]. We'll be talking more about the internals over the coming months. [1] https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-time-sync-service-microsecond-accurate-time/ https://aws.amazon.com/about-aws/whats-new/2023/11/amazon-ti... [2] https://assets.amazon.science/dc/2b/4ef2b89649f9a393d37d3e042f4e/amazon-aurora-design-considerations-for-high-throughput-cloud-native-relational-databases.pdf https://assets.amazon.science/dc/2b/4ef2b89649f9a393d37d3e04...
- tristor 3y ago> makes use of the high-quality hardware clocks available in EC2 So are you using vector clocks in the backend to handle transaction ordering and maintain consensus?
- alex_127 3y agovector (or matrix) clocks are awfully slow in real life. have shipped product with this, will not recommend for future.
- dkhenry 3y agoI appreciate the context, but this just adds so many more questions. How are you integrating hardware clocks into the Postgres transaction model? Is this just postgres wire compatible like Cockroach or Yugabyte, or have you modified Postgres MVCC implementation to not use the standard TID. Does this support all of postgres at the Transaction coordinator ? Super exciting announcement, and I am really looking forwards to learning more!
- jitl 3y agoThey’re probably using it similar to how Spanner uses TrueTime: https://cloud.google.com/spanner/docs/true-time-external-consistency https://cloud.google.com/spanner/docs/true-time-external-con...
- dkhenry 3y agobut Spanner isn't Postgres compatible, they have their own transaction processing layer that is built for TrueTime. If this is using the Postgres frontend how are they incorporating TrueTime into the normal Postgres MVCC model that uses the monotonic xid. edit: More details make it seem like this _isn't_ Postgres on the frontend, its just mostly postgres compatible, which would explain how they got away from the xid based MVCC. So it seems like this is an entirely new distributed database, rather then a modification to existing postgres.
- eightnoteight 3y agoI think they are internally using postgres only. aurora uses a custom storage process on host which internally supports the custom storage engine (s3) given the mentions of "log as a database", I believe that depending on the transaction the storage responds differently. like how mysql mvcc uses undolog to essentially rollback database internally so that transaction sees data consistency. they could be doing something similar i.e regardless of whatever postgres uses, if they can get a reference of transaction and its start time then they can use the custom storage engine to rollback the log and respond in that way
- eightnoteight 3y ago
- twotwotwo 3y agoI'm glad someone from AWS is here! Some general info on the sharding approach would be interesting, e.g. whether secondary indices on sharded tables are scatter-gather, or separately sharded on the index columns, or you configure it per index, or what. From what I hear from folks using things like Vitess, if you're used to a monolithic SQL database there's often some things to learn to mentally model your query cost well after you move to a sharded world, and understanding more up front can save heartburn later. Writing up those details well is a good thing that AWS could do.
- vasco 3y agoThis does look exactly like CitusDB even in the naming of some things (reference / distributed tables) and how they describe it being added to an existing cluster as extra nodes with an extra manager node that will redirect requests!