6 ms·
ScyllaDB is Moving to a New Replication Algorithm: Tablets
- carpintech 3y agoMoving from Vnode-based replication to tablets to dynamically distribute data across the cluster
- jzelinskie 3y agoThis sounds a lot like ranges in CockroachDB. Anyone familiar with the deep details to highlight the differences?
- Nican 3y agoI thought of the same thing, so I am trying to find information on the documentation about things that CockroachDB does very well: 1. Consistent backups/transactions. When a backup is made, is that a single point in time, or best-effort by individual tablet. For example, backing up an inventory and orders table, the backup could have an older version of inventory, where orders have already been completed for some of them. It looks like Scylla backsup per node, so it could mean that data might have a slight time offset from one another. 2. Replicate reads. Like CockroachDB, it looks like Scylla will redirect the read to the range lead, but CRDB also provides the option to get stale reads from a non-lead. This is usually good for cross-region databases where reading off the lead can be big increases in latency. I do not have a lot of time, but I am having a hard time finding much information about the architecture on Scylla's documentation. My personal guess is that Scylla optimized their code for performance, and less worry about data integrity.
- acevedoorafael 3y ago> My personal guess is that Scylla optimized their code for performance, and less worry about data integrity. Definitely, but they are implementing Raft-based transactions which provide higher consistency. That should enable a higher variety of use cases.
- geenat 3y agoWould be nice if the deployment story became a bit more like CockroachDB too.
- eatonphil 3y agoHow so? Mind saying more?
- aeyes 3y agoI suppose that they are saying that CockroachDB is a single binary which you just drop on a machine and you are good to go. For ScyllaDB you need to install Java, Python and several ScyllaDB related packages.
- hobofan 3y agoJava? I thought the whole raison d'etre for ScyllaDB was "Cassandra without Java"?
- taywrobel 3y agoThe server implementation is, but administering it still requires the Java based Cassandra tooling like nodetool and cqlsh
- heipei 3y agocqlsh is written in Python. Which doesn't mean it's less of a pain in the ass ;)
- taywrobel 3y agoSorry, that was phrased poorly; was in reference to the parent comment’s “For ScyllaDB you need to install Java, Python and several ScyllaDB related packages”. Just meant to say it does have tooling which requires other languages/environment specifics.
- dmagda 3y ago[flagged]
- dangoodmanUT 3y ago"can i copy your homework" "yeah but change it a bit" (compared to the other comment that looks nearly identical, looks lik yugabyte sent some people in here)
- intelVISA 3y agoYeah sry about that, they gave us an expired GPT4 key so we had to reuse the same lines a few times.
- bbss 3y agoVery similar to how BigTable[1] works under the hood which was built ~20 years ago. [1] https://static.googleusercontent.com/media/research.google.com/en//archive/bigtable-osdi06.pdf https://static.googleusercontent.com/media/research.google.c...
- jeffbee 3y agoThe load shifting part is similar to the way BigTable splits, merges, and assigns tablets. But the rest of it is not related, because BigTable does not try to offer mutation consistency across replicas. If you write to one replica of a BigTable, your mutation may be read at some other replica, after an undefined delay. Applications that need stronger consistency features must layer their own replication scheme atop BigTable (such as Megastore). What this post is describing for replication seems more comparable to Spanner.
- axiak 3y agoI don't understand this comment. Bigtable requires that each tablet is only assigned to one tablet server at a time, enforced in Chubby. There's no risk of inconsistent reads. Of course this means that there can be downtime when a tablet server goes down, until a replacement tablet server is ready to serve requests.
- jeffbee 3y agoRight, the contrast I was trying to draw was between what they depict, where multiple nodes are holding a replica of the tablet and performing synchronous replication between themselves, and what BigTable would do, which is to have the entire table copied elsewhere, with mutation log shipping. What they are doing is more analogous to how Spanner would do replication.
- dikei 3y agoUnless you're doing multi-cluster replication, there is no log shipping in BigTable: the data replication within a cluster is taken care of by the underlying filesystems. Single-cluster BigTable is strongly consistent.
- bsdnoob 3y agoTablets remind me of vitess
- magden 3y agoThis is the right move for Scylla. Overall, looks similar to YugabyteDB that distirbutes data by sharding tables into tablets as well. The cluster monitors the cluster size (number of nodes) and the size of each tablet (data volume), and adds new tablets or re-shards large ones automatically:https://docs.yugabyte.com/preview/architecture/docdb-sharding/tablet-splitting/#automatic-tablet-splitting https://docs.yugabyte.com/preview/architecture/docdb-shardin...
- dangoodmanUT 3y ago"can i copy your homework" "yeah but change it a bit" (compared to the other comment that looks nearly identical, looks lik yugabyte sent some people in here)
- collinc777 3y agoThis sounds like partitions in DynamoDB