7 ms·
YugaByteDB – A Transactional Database with Cassandra, Redis and PostgreSQL APIs
- lukeqsee 8y agoDoes anyone know a comparison between this and CockroachDB? Or have experience running either in production? They seem to compare against other databases, but not against Cockroach which seems to be the biggest competitor. I'm looking at implementing a global-scale database cluster with very specific requirements, and YugaByte seems to meet those but a comparison against CockroachDB seems warranted. Edit: just noticed they do compare features against CockroachDB: https://docs.yugabyte.com/latest/comparisons/#distributed-sql-databases https://docs.yugabyte.com/latest/comparisons/#distributed-sq..., but they don't have an in-depth comparison.
- ddorian43 8y agoPostgres layer is still extremly beta (has no update support etc.) But it should become good fast since they use fdw api (storage engine api in pg12) and reuse most of pg code.
- rkarthik007 8y agoHi @lukeqsee, we are working on just this, please stay tuned!
- atombender 8y agoYugaByte only added distributed transactions in the new 1.1 release (previously, it only supported single-row transactions). The architecture is a variation on 2PC and seems sound, but I think it's fair to say that it's early days. Meanwhile, Cockroach and TiDB both have battle-tested implementations. YugaByte is pretty quirky. Rather than settle on a native data model, they offer several "personalities" that mimic other products (Cassandra, Redis and PostgreSQL), and you can mix and match them. But they've implemented each API with their own set of weird warts. For example, their CQL implementation has CREATE INDEX, but it does not index existing data [1]. You have to either create the secondary index before inserting data, or force a reindex of everything with a dummy UPDATE statement. Who would ship such a product? More warts/Cassandraisms: UPDATE is actually an upsert; an update that doesn't match a row will insert it. And SELECT ... WHERE expressions can only use AND expressions (!) [2]. Hopefully the forthcoming SQL API should be saner, but it's very limited at the moment. It does not seem to support joins, transactions or indexes, for example. Meanwhile, CockroachDB and TiDB both have rich SQL implementations, including joins and aggregations, with cost-based query optimizers that can take advantage of multiple indexes and table statistics. It seems more appropriate to compare YugaByte with distributed key/value stores like FoundationDB, Cassandra, Scylla and Redis. [1] https://docs.yugabyte.com/latest/explore/transactional/secondary-indexes/ https://docs.yugabyte.com/latest/explore/transactional/secon... [2] https://docs.yugabyte.com/v1.0/api/cassandra/dml_select/#root https://docs.yugabyte.com/v1.0/api/cassandra/dml_select/#roo...
- lukeqsee 8y agoThanks atombender! This is helpful.
- manigandham 8y agoThose things you mentioned are just how Cassandra works. It's a wide-column key/value store and everything is an upsert because that's how the write path is designed, there is no read-then-write (other than LWTs). If you're using an application that expects Cassandra then this is normal so why would Yugabyte change the semantics? It's true that indexing is unfinished but they've also long been problematic in the CQL data model so Yugabyte is moving faster and actually releasing something for those who can work with it. Cassandra took years to come up with several types of indexes that all have problems and Scylla is overdue by 18 months on their implementation.
- atombender 8y agoThat's why I referred to the behaviour as Cassandraisms. The question was posed in the context of CockroachDB, so my answer stands -- CDB vs. YugaByte is an apples to oranges comparison, seeing as YugaByte isn't an RDBMS by any useful measure, and has a very long way to go to get there.
- zenithm 8y agoIn the NoSQL space, FaunaDB indexes are very powerful. They are term partitioned and sharded and have compound terms, covered values, transformations, etc.
- manigandham 8y agoCockroach treats the entire installation as one giant cluster, with optional 'locality' feature for each node to distribute data. Yugabyte has regional clusters, which can replicate from a single master write cluster. CRDB is working on ways to get fast local reads for different regions. Other than that, Yugabyte also supports more protocols like Redis and Cassandra, and enhancements within them. It's really more of a distributed key/value store like FoundationDB but with protocol layers on top included out of the box.
- deleted 8y ago[deleted]
- lykr0n 8y agoWhen I see databases like this pop up, a part of me wonders why they don't devote effort to develop a plugin for MariaDB or PostgreSQL? Follow in the steps of Citus Data.
- deleted 8y ago[deleted]
- rkarthik007 8y agoCTO of YugaByte here... this is indeed our plan. We are working with the community on pluggable storage coming out next year, and in the slightly longer term want to make YB a plugin (extension in the case of PG) model. While a reasonable amount of work is needed in order to get there, this is definitely the direction!
- Xorlev 8y agoHere's a question -- what makes YugaByte unique? There's already an early but strong market here in the form of Google Spanner on GCP, Cockroach DB, and TiDB -- what does YugaByte bring that none of them do?
- manigandham 8y agoMultiple data models.
- manigandham 8y agoJust to be clear, this would mean a PG extension that works with the new PG pluggable storage to use Yugabyte as the backing store? Does that mean still running a separate Yugabyte cluster or is everything self-contained in the extension?
- kmuthukk 8y agoYugaByte DB's design is that a YB cluster supports Postgres in a native, self-contained & scale-out manner (much like YugaByte's Cassandra and Redis flavored offerings). At a high-level, the upper half of the Postgres DB is being largely reused. The lower-half, i.e. the distributed table storage layer, uses YugaByte's underlying core-- a transactional and distributed document-based storage engine. For the DB to be scalable, the lower-half being distributed is necessary but NOT sufficient. The upper-half also needs to be extended to be made aware of other nodes executing DDL/DML statements and dealing with related concurrency while still allowing for linear scale. Also, making the optimizer aware of the distributed nature of table storage is the other major piece of work in the upper-half. These changes required in the upper half is what makes the "100% pure extension" model a bit harder... but that's something we intend to explore jointly with the Postgres community.
- sunnycpp 8y agoDesign is very similar to Hbase minus the dependency on Zookeeper and HDFS. RocksDB has been very smartly modified to use it as an optimized storage layer.
- ralfn 8y agoThose are a lot of claims. Databases that make half those claim, turn out to be misrepresenting some of them. So far its empty promises on a box. When can we expect independent 3rd party evaluation of all these claims? The most famous one is Jespen: https://jepsen.io/analyses https://jepsen.io/analyses Because this all sounds way too good to be true. What's the catch? How stable is this? Are the claims tested other than in theory?
- manigandham 8y agoThey made a post about testing with Jepsen: https://blog.yugabyte.com/jepsen-testing-on-yugabyte-db-database/ https://blog.yugabyte.com/jepsen-testing-on-yugabyte-db-data...
- ralfn 8y agoAwesome. Thanks. This is very very encouraging!