5 ms·
CockroachDB's Consistency Model
- bithavoc 8y agoI just learned a new thing: Time Travel Queries[0]. I haven't used it but it seems queries are able to go back 25 hours by default, freaking cool. [0]https://www.cockroachlabs.com/docs/stable/select-clause.html#select-historical-data-time-travel https://www.cockroachlabs.com/docs/stable/select-clause.html...
- tyingq 8y agoIt seems like there should be a specialty database that lets you branch any point in time for not just queries, but also inserts, updates and deletes. I recall a fair amount of buzz around something called "thingamy" that was going to deliver that. Then it went quiet, closed source, and "call for pricing".
- SamReidHughes 8y agoThat is certainly possible. What would you use it for?
- tyingq 8y agoInstant what-if for end users that doesn't require IT help or lots of space/time, and doesn't interfere with the current prod data, etc. Think data savvy, but not IT savvy end users. Like financial analysts.
- SamReidHughes 8y agoThanks for answering. Myself, I'd only thought of performance for creating and using test environments in software development. Which, I guess, is another kind of what-if scenario.
- foobiekr 8y agoFork() the process, fork the database. That would be pretty interesting. Long running jobs where transforming a consistent snapshot is required.
- keymone 8y agoBeing able to deterministically look at what database looked like at any point in time is priceless for debugging/investigating sessions.
- tehlike 8y agoEventsourcing/eventstore gets close.
- NightMKoder 8y agoDatomic does this. There’s (Clojure syntax) d/with that given a database at any point in time runs a transaction “virtually” (there isn’t a current db per se - you can use (d/db connection) to get a reasonably current one though) . You can then query that virtual database as you would a real database. If you do want to durably commit, you have to use (d/transact) which takes a transaction in the same format as d/with, but operates on a connection rather than a point in time database (the transaction can have special functions that read the active db at the point in time of the transaction, to implement things like cas).
- atombender 8y agoThis "branching" could be thought of as a generalized kind of ACID transaction. After all, a transaction mimics having an isolated snapshot of your database, a one-level-deep temporary branch that is merged back on commit, and which can (like version control merges) fail on conflict. I'd love to see a database that supported branching in this way. The difficulty here isn't really the technical implementation of the branching at the data store level, which is rather trivial, but the merging. The challenge is exactly the same as a version control system like git. git can leave conflict merging to a human operator, but a database system can't, unless you bubble the conflict up to the client, which will work in some cases (interactive UIs) and not in others (automated sysrems). To support any kind of conflict resolution at all, values and updates need to be reformulated as CRDTs or some similar mechanism that can provide conflict-solving semantics. Instead of recording tuple changes as simple value assignments, they need to be merged as operations -- if x starts out as 42, then doing "update ... set x = x + 1" is the equivalent of "set x = 43" in a typical modern RDBMS transaction, but to preserve the semantics in the face of a conflicting merge, the transaction log needs to be recorded as "x = x + 1".
- toomim 8y agoI'm working on something like this! I have been having the exact same idea.
- ryanworl 8y agoThe implementation of these kinds of features is a lot easier than one might think if you use multi-versioned storage. You delay garbage collection of old versions of rows until the configured cut off (or never if you want to go to any time) and ignore rows with a higher timestamp when reading. Where you can hit trouble is if you can’t garbage collect fast enough and you run out of IOPS or disk bandwidth and the system wasn’t smart enough to issue back pressure before the new data was accepted into the system.
- grogers 8y agoHowever, the garbage collection is typically _heavily_ optimized for short lived garbage. Try pausing GC by keeping a transaction open for an hour or more and see how the DB performance changes. So I think there's still probably a little extra smarts going into something like mariadb's system versioned tables.
- ryanworl 8y agoI may be misunderstanding the documentation, but it appears MariaDB system versioned tables don't do any garbage collection at all (it stores all versions forever), and requires reading over (skipping) old versions at runtime.
- atombender 8y agoPostgres used to have time travel -- it is relatively easy feature to build on a MVCC storage -- but it was deprecated in 6.2 and then removed [1]. Performance was one reason listed, as was space usage. It's easy to see why. For example, modern Postgres versions will in some cases reclaim dead tuples for new data, thus "cheating" by bypassing MVCC. If Postgres had to support time travel, such optimizations wouldn't be possible. I seem to remember developers talking about adding time travel back now that the codebase is better (and better understood) and it could be redesigned around current techniques. It's a cool feature to have. [1] https://www.postgresql.org/docs/6.3/c0503.htm https://www.postgresql.org/docs/6.3/c0503.htm
- evanweaver 8y agoFaunaDB is temporal as well and allows for unlimited historical retention. You can also compare two different times in history in a single transaction.
- shaklee3 8y agoThis is a great article. I read the jepsen crdb analysis long ago, but never understood exactly what was wrong. This describes it (and defends) really well.
- ccmonnett 8y agoI don't use it (yet) but every interaction I've had with Cockroach as a company has been great. Even to the point where I had a marketer for a not-quite-direct competitor to Cockroach whisper to me "You know, for your use case, you should probably just use Cockroach..."
- evanweaver 8y agoThe suggestion that this anomaly can occur in FaunaDB is wrong: > We do not allow the bad outcome in the Hacker News commenting scenario. Other > distributed databases that claim serializability do allow this to happen. FaunaDB's consistency levels are as follows: | No indices | Serializable indices | Other indices -----------+--------------+----------------------+-------------- Read-Write | Strict-1SR | Strict-1SR | Snapshot Read-Only | Serializable | Serializable | Snapshot In FaunaDB, all read-write transactions execute at external consistency, including all their row read intents, and all index intents when requested. Any read-only transaction, even if uncoordinated, will see a consistent prefix view of all its priors. Not just causal priors, so there is no "causal reversal". All physical, linearizable, multi-partition priors. In FaunaDB, it is not possible to see the later comment without seeing the previous comment, even when executing at snapshot instead of externally consistent isolation. FaunaDB can serve serializable/snapshot reads out of any datacenter without any global coordination, and can serve externally consistent reads with coordination whenever requested. CockroachDB doesn't offer global reads at all, to work around the clock skew issues, but can only serve reads from partition leaders. Comparing these models fairly requires executing all transactions in FaunaDB at external consistency, which is both more consistent, and has lower tail latency in a global context, than CockroachDB.
- andreimatei1 8y ago> The suggestion that this anomaly can occur in FaunaDB is wrong: You're right. I've removed the reference to FaunaDB. Sorry about that. What I've meant to reference is the snapshot isolation of FaunaDB's read-only transactions, but the comment was in the wrong context for that.
- andydb 8y agoFaunaDB has to make painful (to applications) tradeoffs between latency and consistency in global scenarios: 1. All read-write transactions pay global latency to a central sequencer. So, yes, FaunaDB is strictly serializable, but at the cost of high latency of read-write transactions. 2. Read transactions have to choose between: a. Strict serializability but high latency b. Stale reads but low latency Anomalies absolutely can happen in FaunaDB if applications use option 2b. However, most users will not appreciate the subtlety here, and some will unwittingly go into production with consistency bugs in their application that only manifest under stress conditions (like data centers going down and clogged network connections). Their only other option is 2a, and that is just a no-go for global scenarios. You can't route your regional reads through a central sequencer that might be located on the other side of the world. Your argument reduces to: "FaunaDB has no anomalies in global scenarios! That is, as long as you're OK with a global round-trip for every read-write and read-only transaction...". FaunaDB has not solved the consistency vs. latency tradeoff problem, but has simply given the application tools to manage it. A heavy burden still rests on the application. By contrast, Spanner users get both low latency and strict serializability for partitioned reads and writes (that's and, not or, like FaunaDB). CockroachDB users get low latency and "no stale reads" for partitioned reads and writes. Partitioning your tables/indexes to get both high consistency and low latency is a requirement, but it's not difficult to do this in a way that gives these benefits to the majority of your latency-sensitive queries, if not all of them. After all, this is what virtually every global company does today - they partition data by region, so that each region gets low latency and high consistency. You only pay the global round-trip cost when you want to query data located across multiple regions, which is rare by DBA design. The main point of the article is that the "no stale reads" isolation level is almost as strong as "strict serializability", and is identical in virtually every real-world application scenario. This means CockroachDB is equivalent to Spanner for all intensive purposes.
- andrewflnr 8y agoYou might want to come up with a better example for how your causal reverse anomaly isn't that big a deal. My first reaction was approximately "dafuq? Nathan saw an inconsistent database! That's terrible!" It took a long time before you got around to explaining that with a foreign key check (which I pretty much assumed as part of the scenario), it wouldn't happen.
- andreimatei1 8y agoPerhaps I could emphasize things differently. FWIW, besides the fact that having a foreign key constraint in that schema would prevent the badness from happening, the even bigger reason why scenarios like that are unlikely is that, realistically, for Tobi to reply to a comment, he must have seen that comment he's about to reply to, and it's very hard to imagine a scenario where he'd see it but still have the response transaction not read it (cause if the txn read it, the two transactions wouldn't be independent any more and so they'd be well ordered). The foreign key is just one way of ensuring that the read happens.
- andrewflnr 8y agoTo me, that still makes it sound like a bad example. Adding all those qualifiers is a distraction from the actual point you're trying to make.