7 ms·
Immutability, MVCC, and garbage collection
- jasonwatkinspdx 13y agoPretty disappointing critique. There are manifold differences between couchdb, datomic, rethinkdb and more traditional sql databases, but the author can't see past his pet issue. He doesn't seem to understand the use cases, or the infrastructure differences. In terms of use cases, there are plenty of analytically datasets that are strictly monotonic. There is no opportunity to reclaim overwritten storage in this case. But CoW index write amplification you say? Well now let's talk about infrastructure differences. Datomic can use s3 as a storage layer. Cold data storage is then $10/month/terabyte. And compaction is trivially non-blocking in a CoW dataset. It can be done by EC2 spec instances whenver is convenient with zero impact on the production system. Don't generalize the experiences of a few folks running couchdb on a couple in house servers to people working with a fundamentally different approach on fundamentally different infrastructure. Back to the cheap cost of cold storage, for most businesses that's well into "who cares" cost territory, but having the data also allows them to preserve full lineage in the advent they discover a data loss bug in their application. That could potentially be quite valuable. One way to think of these systems is that they have integrated backup data. MySQL folks generally don't count the size of backup sets as part of the database footprint, and common backup schedules (eg father, grandfather) provide only a limited ability to dig back into the past. By way of simple argument: would anyone advise using a source control system that only kept the last few hundreds of commits? Why do we treat our data differently? The answer is largely one of implementation constraints: text is cheap, vcs doesn't index the way databases do, etc. But as storage and processing densities continue to grow more datasets shift to the trivial category. As much as many HNews posters dream of being in a 'big data' shop, the vast majority of startups have maybe a couple GB in MySQL or Postgres. Our status quo toolset and approach is discarding value for these folks. Maybe you think that's fine, but some of us are going to keep pushing to enable new capabilities and advantages. Increasingly these implementation constraints will disappear as people attack them. It'd be nice to not suffer scornful surface criticisms from people who base their career around the status quo.
- jeffdavis 13y agoI think you are making a mistake by reading it as a critique of immutability in general. Rather, I think the author is criticizing an overly-simplisitic approach to using immutability. When he describes traditional databases: "What’s happening in such a database is a combination of short-term immutability, read and write optimizations to save and/or coalesce redundant work, and continuous “compaction” and reuse of disk space to stabilize disk usage and avoid infinite growth." He's saying that's what people arrive at after working out the issues with immutability. When you say: "Increasingly these implementation constraints will disappear as people attack them." the author would probably respond: "Yes, and people will attack them with a lot of the same strategies that SQL databases are already using today".
- jasonwatkinspdx 13y agoI strongly disagree with the idea that temporal databases will be forced to revise themselves to the implementation designs of Postgres, MySQL, etc. In fact, I'd make the firm prediction that the paged window CoW strategy of LMDB will become dominant for these more traditional database designs, and that a new class of temporal databases will improve upon that architecture.
- jeffdavis 13y agoIt's not so much that everything will look exactly like postgres. It's that the biggest difference right now between the designs is complexity (the kind that comes with maturity). Every mature general-purpose database system makes use of both CoW/immutable design principles as well as locking and update-in-place in various contexts. It remains to be seen how "temporal databases" (as you call them) will adapt to real-world demands, but it might not look as radically different as you think. Also, for anyone who thinks that CoW is the be-all-end-all design principle, I have my doubts. The main reason is that it doesn't compose very cleanly -- a design that involves CoW over CoW over CoW starts to look wasteful after a while. That being said, I would not be surprised if some of these new databases found some new trade-offs that worked very nicely and got adopted by more traditional systems. Storage is not nearly as much of a concern for many kinds of problems, and it seems natural that such a shift would lead to some design changes.
- rdtsc 13y agoWell they are all tools with trade-offs. It is good to have something like Datomic. The trick with Datomic (from a 10 minutes summary I listened to last year) is that because of immutability interesting things can happen in respect to caching. You can cache things locally and have smart clients. It simplifies and adds interesting performance properties. Sometimes immutability is not right for you, sometimes it is. He criticizes CouchDB as well. Yeah you need to have double the disk space. But CouchDB now has auto-compaction triggers based on fragmentation threshold or based on times of day. Immutability allows completely lock-less reading. Crash only stopping (can just kill the power or the service anytime without corrupting the database). Master to master replication lets you build custom cluster topologies. Well defined conflict models and explicit conflict handling is invaluable (as opposed to other databases that hide and paper over that, sometimes losing data in the process -- read Aphyr's blog http://aphyr.com/ http://aphyr.com/ for an thorough study on that). Do you need those features? Maybe you don't, maybe you do. It is good to be aware of it and have it your toolbox if you need it one day.
- carry_bit 13y ago"If entities are made of related facts, and some facts are updated but others aren’t, then as the database grows and new facts are recorded, an entity ends up being scattered widely over a lot of storage." As far as I understand this isn't so. All data in Datomic is stored in the indices, so the facts associated with an entity should be stored together.
- gfodor 13y agoI'm going to go out on a limb here and use a logical fallacy, but I have a feeling that Rich Hickey probably has some understanding of the history of database theory and didn't just build Datomic out of ignorance.
- jeffdavis 13y agoBut he also does not have infinite engineering resources at his disposal. I expect that a lot of the designs are probably somewhat simplistic compared to the battle-hardened implementations in common use today. It doesn't mean Rich Hickey was wrong; merely that we are seeing only the ultra-clean design now. After Datomic takes off and has a lot of production users in demanding environments for 5-10 years, I doubt the design will be quite so clean.
- DigitalJack 13y agoI thought he codesigned it with a company called Relevance
- moomin 13y agoWell, there used to a firm called Datomic, and one called Relevance, but they merged to Cognitect.
- jeltz 13y agoIf we are allowing the appeal to authority fallacy then neither would Baron Schwartz write this article out of ignorance. We have two different authorities disagreeing here.
- vardump 13y agoAppend-only b-tree/skiplist/whatever immutability can be combined with "checkpoint hashes" (think git) using for example SHA512, that can be transmitted elsewhere at a low cost and used later to validate system integrity. Any data store tampering would be detectable. In a mutable database you'd need to transmit whole transaction log for the same integrity guarantee. Thus immutable store makes it trivial to implement systems that need to have full history and audit trail. When there's only a trivial amount of data involved, immutable state can solve a lot of headaches. No programming error can lose or corrupt data for instance - every case can be traced back after the fact. I hope some immutable store databases take off that cover those use cases. Make it a graph database, bonus points if the underlying engine implements directed hypergraph. And artificial intelligence reasoning engine. :)
- greatsuccess 13y agoId like to keep immutability and databases separate in this comment. Immutability the supposed "big advantage" of functional languages, is a discussion which is one-way. No one ever discusses the implications of a system which is constantly, needlessly, insanely doing nothing but MAKING COPIES OF DATA. This is not how systems are/should be designed and Im sure as hell not going to use a functional language until functional understands basic common sense. A system that is constantly copying data for no reason will do nothing else.
- catnaroek 13y ago> Immutability the supposed "big advantage" of functional languages No, immutability is just the means. The objective is writing programs that are easier to reason about. > No one ever discusses the implications of a system which is constantly, needlessly, insanely doing nothing but MAKING COPIES OF DATA. Just because copy assignment requires, well, copying in languages like C++, it does not mean things work the same way in functional languages. Precisely because data structures are immutable, you can internally just pass around a pointer to the data structure and pretend you have copied it (e.g., when you pass it as parameter to another function).
- greatsuccess 13y agoYou are correct, thats easy to reason about unless you are writing the garbage collector, in which case its pointless to write a garbage collector because there will never be anything to collect and memory will expand infinitely. As for the persistent data structure comments, that has nothing to do with the fact that every function in a functional program has to copy data at such a ridiculously fined grained level that all you are doing is writing a garbage creator. This data that every function creates in the name of immutability is meaningless, void of any purpose and is a bug.
- talaketu 13y ago> No one ever discusses the implications of a system which is constantly, needlessly, insanely doing nothing but MAKING COPIES OF DATA. see persistent data structure
- yresnob 13y agoThe author needs to read up way more on datomic... and clojure for that matter then try again...
- dschiptsov 13y agoVery rare example of an old-school thinking and reasoning. Indeed, industrial-strength databases are hard, and as an Informix DBA, I could only agree, that there is no silver bullet, and that just "check-point early, check-point often" strategy has its trade-offs in terms of concurrent query performance. And, of course, unpredictable stop-the-world GC pauses are unacceptable. Informix have polished their Dynamic server for a decade and it is still one of the best (if not the best) solutions available. The more subtle notion is that one could never avoid compactiom/GC pauses, so the strategy should be to partition the data and perform "separate" compaction/GC, blocking/freezing only some "snapshot" of the data, without stopping the whole world, like some modern FS do. Of course, implementation is hard, and here the simplicity of persistent, append-only data-structures could be beneficial for the implementer of "partitioned/partial GC". There must be some literature, because these ideas were well researched in old times, when design of Lisp's GC, or CPU caches were hot topics.
- noelwelsh 13y agoThe author ends with "Also, if I’ve gone too far, missed something important, gotten anything wrong, or otherwise need some education myself, please let me know so I can a) learn and b) correct my error." so allow me to offer the following: When you spend a substantial portion of your post sniping at the competition you come across as an arrogant arse. For example, second sentence in the post is "My overall impressions of Datomic were pretty negative, but this blog post isn’t about that." Then remove this line! It adds nothing to the point you claim you want to make. Same thing in the next section regarding append-only B-trees, and then we have snark directed at RethinkDB, and so on. I got about half-way through the post before I just skimmed to the end as I was sick of the arrogance and insults, and I was beginning to doubt your authority to comment on this area as a result -- if you have substantial points to make you don't need to cover them up with this crap.
- webmaven 13y agoTo me, what was most striking about the article (once I got past the pot-shots), was how much the descriptions of Datomic and RethinkDB reminded me of the ZODB, the object database built into the Zope framework: http://www.zodb.org http://www.zodb.org ZODB is an append-only object database, and one of the data structures commonly used for persistence in it are BTrees: http://www.zodb.org/en/latest/documentation/guide/modules.html#btrees-package http://www.zodb.org/en/latest/documentation/guide/modules.ht... Several disadvantages noted in the OP (infinitely growing data, needing to pack the DB to discard old object versions, needing 2x disk space to do the packing) definitely apply to ZODB, but others do not, as it is ACID-compliant and implements MVCC.
- richhickey 13y agoI'd ignored this, trying to have a nice holiday with my family (without having to contend with wrongness on the internet :), but would like to introduce some facts. Datomic is not an append-only b-tree, and never has been. There's a presumption in the article that "immutable" implies "append-only b-tree", specifically, an argument against append-only b-trees is used to disparage immutability. That's pretty naive. One only need to look at e.g. the BigTable paper for an example of how non-append immutability can be used in a DB. Such architectures are now pervasive. Datomic works similarly. One can't make a technical argument against keeping information. It's obviously desirable (git) and often legally necessary. Our customers want and need it. Yeah, it's also hard, but update-in-place systems (including MVCC ones) that discard information can't possibly be considered to have "solved" the same problem. Prefixing a polemic with "apparently" doesn't get one off the hook for spreading a bunch of misinformation.