6 ms·
Document Database Transaction Models
- evanweaver 5y agoI’m around to answer questions and discuss. Anybody who has used Couchbase transactions, or sharded Mongo transactions, and can corroborate our analysis?
- jimsimmons 5y agoI like the presentation of ideas. Can you expand on use cases of Firebase’s nested document model? Seems powerful as a file system but not sure how that will play with the complexities of distributed applications.
- evanweaver 5y agoFirebase was originally designed more as a realtime communication mechanism than an operational database. The idea was that clients would subscribe to different nodes in a data hierarchy to receive realtime notifications from other clients that were publishing to those nodes. Depending on what was in the client view, sometimes you wanted to subscribe to a leaf, sometimes to a subtree, sometimes to everything. As these things tend to go, when there is a place to store arbitrary data, all kinds of things get shoved into it, so the mixed model in Firestore is a compromise between the original tree-of-nodes data model and a more conventional document data model. My assumption is the Firestore-to-Spanner mapping creates subcollections as shared tables with foreign keys to the parent documents, but I don't actually know. However, that would match the mandatory 1-to-many-to-1-to-many data layout, and makes more sense than shoving all the dependent data into the document itself or creating multiple millions of SQL tables for millions of documents.
- jimsimmons 5y agoThis is classic HN speak but it seems like it should be easy to build a Raft or Paxos consensus based sharded database. We have had plenty of Spanner esque implementations since it came out. Would it be too hard to take something like TiDB and modify it to support nested document model? Of course testing and bringing it up to production levels is its own feat but nonetheless I’m surprised such an interesting DB avenue gets so little attention.
- sargun 5y agoHaving played with the Firebase Python client, outstanding writes within a transaction are "kind of" exposed because you fetch the document, and make alterations to it before saving it (within a "transaction"). That object, when you make alterations to it, bubbles up the changes -- AFAICT, this is a bit of client side hack, but it's ergonomically wonderful.
- evanweaver 5y agoI guess most ORMs are like that. Is the object shared by reference across the entire runtime or do you end up with divergent objects?
- sargun 5y agoIt's only valid in the context of that function invocation. The docs say not to write impure functions or do anything with concurrency -- if you start two simultaneous transactions, the client isn't "smart" about it unfortunately.
- sargun 5y agoGiven that most application developers started with something like a postgresql, or a MySQL that has pessimistic locking, and thus transactions rarely end in read / write conflicts (aborts), how have people (and potentially programming languages) adapted to optimistic concurrency control? Also, when you say: > On the other hand, server-side transactions use a more typical pessimistic relational lock. They open a transaction, do some work, and then commit it. What do you mean? What are server-side transactions?
- evanweaver 5y agoWe haven't seen much difference in practice between frequent aborts and frequent timeouts. Both are better than deadlocks. I meant transactions issued from a server ("cloud") client to the database, as opposed to a mobile client.
- jhgb 5y ago> how have people (and potentially programming languages) adapted to optimistic concurrency control If they started with PostgreSQL, then they're already adapted to optimistic concurrency control, right?
- vvern 5y ago> Additionally, out of band coordination, most likely via human beings, is required to make sure that all potential readers of the transactional writes are also transaction-aware. Subtle burn on the mongo and its client-coordinated session causality model.
- Pamar 5y agoDocument databases are very convenient for modern application development, because modern application frameworks and languages are also based on semi-structured objects instead of tabular data. Citation needed.
- mjburgess 5y agoAll object-oriented programs have a data model which is "semi-structured" (ie., non-tabular). Almost no software application data models are tabular.
- jhgb 5y agoWait, you need a citation to show that commonly used languages have structs/records/classes with pointers/OOPs? OK, in that case, I'm citing Java spec, C++ spec, C# spec, Javascript spec, Smalltalk spec, Common Lisp spec...
- Pamar 5y agoAll the languages you listed worked fine to develop complex applications using "tabular data". I find ... debatable that building yet again one more FB or IG challenger is significantly more important than maintaining or creating new Banking, ERP, Inventory Management solutions and for that stuff "tabular data" doesn't seem to be a bad match, or at the very least, using json or xml documents will not provide much of an improvement (all IMHO, of course...).
- jhgb 5y ago> All the languages you listed worked fine to develop complex applications using "tabular data". Point is, it's not a native representation for them. If you just wanted to serialize a portion of a program's heap, independently of the program (which is presumably the "very convenient" part of it), it would be a graph of records or objects, not a bunch of tables. For example, no word processor to my knowledge represents the document being edited as a bunch of tables outside of the language's native data model. That's why document databases don't use tables for documents.
- jhgb 5y ago> In the distant past of the 2010’s, document databases didn’t offer transaction support, instead implementing various forms of eventual consistency. Vendors and open-source maintainers promoted the idea that transactionality was an unnecessary, complexifying feature that damaged scalability and availability—and many claimed that adding it to their systems was impossible. I find it interesting that the proposed explanation for lack of transactions in document databases is the CAP theorem which covers distributed systems, NOT document databases. Clearly such thing as a document database with transactions is not impossible if, for example, GemStone/S can support transactions just fine.
- evanweaver 5y agoGemStone/GemFire use a transactional protocol akin to Tuxedo. Open a bunch of locks, write a bunch of updates, release the locks. As per the docs (https://gemfire82.docs.pivotal.io/docs-gemfire/latest/developing/transactions/how_cache_transactions_work.html https://gemfire82.docs.pivotal.io/docs-gemfire/latest/develo...) this does not offer isolation or even atomicity, so it doesn't give you the C in CAP at all. These are exactly the kind of "transactions" you get when you try to implement everything at the application level rather than the database level. Couchbase transactions (in the article) are the same. And it's not that different from Vitess cross-shard transactions either, which are not isolated (https://vitess.io/docs/reference/features/two-phase-commit/ https://vitess.io/docs/reference/features/two-phase-commit/). Tandem SQL used the same scheme as well I believe. Prior to Spanner, there were no production databases that offered ACID transactions across distributed, disjoint shards.
- jhgb 5y agoI'm sorry, what does Gemfire have to do with Gemstone/S? That seems like a completely different software from a different vendor. > Open a bunch of locks, write a bunch of updates, release the locks. That's how transactional databases using two-phase locking generally work, isn't it?
- evanweaver 5y agoI thought GemFire was directly derived from GemStone, via numerous acquisitions. If GemStone has a different transaction model I don't know it. The point is, no distributed database with a naive two-phase lock is truly transactional.