5 ms·
I'm a big fan of MongoDB, and I think its replicated durability characteristics are good enough for many classes of applications. If you could lose a minute of
by jamwt 16y ago
I'm a big fan of MongoDB, and I think its replicated durability characteristics are good enough for many classes of applications. If you could lose a minute of data and have it just be "bad" instead of "customer enraging and business threatening", then it's a very nice database system.
However--I do think the decision to make writes "fire and forget" is just a mistake. If you use abstractions like a connection pool under heavy concurrency, you can get unpredictable behavior in terms of when the data is "actually there."
For example, all in one thread:
with connection pool: do_write operation A
do other things...
with connection pool: read something, assuming A has been applied
Specifically, the fact that you get an arbitrary connection out of the pool means you cannot be sure that the database has completely processed operation A before executing your new query.
MongoDB has a "safe=" flag in their Python bindings that implements the project's official (http://www.mongodb.org/display/DOCS/Last+Error+Commands#LastErrorCommands-UseCases http://www.mongodb.org/display/DOCS/Last+Error+Commands#Last...) recommendation for "it is written" consistency. It's a bit of a hack, but it calls "getLastError()" on the connection that does the write before returning from the update()/save()/insert()/delete() call. I think it's astonishing this behavior isn't the default.
In fact, I recently updated diesel's bindings to be safe=True, so getLastError() is always executed on write operations to make sure they succeeded before the call returns:
http://github.com/jamwt/diesel/commit/95cd71d82ffc5c308060b651b6d1056dd1908b45 http://github.com/jamwt/diesel/commit/95cd71d82ffc5c308060b6...
- ericflo 16y agoIt's not really about losing a minute of data though--the whole thing could become corrupted, and repairing it will be extremely difficult.
- kristina 16y ago...which is why you have a slave. Master corrupt? Promote the slave to master, get it its own slave. Repairing a corrupted database, even a relational database, often takes too long for a production app.
- janl 16y agoSo you run naked during the time of the master rebuild? — The only sensible solution is to run two slaves at least, IMHO. — I haven't looked, but do you promote that?
- mathias_10gen 16y agoActually, that will be a recommended config for replica sets. Most of our slides already show 3 replicas per set. Also, most people have backups so its not really "running naked". See http://www.mongodb.org/display/DOCS/Backups http://www.mongodb.org/display/DOCS/Backups for a few ways to backup mongodb. With LVM/EBS/ZFS or any other snapshotable filesystem, backups can be done almost instantly. With EBS you can even get an insta-slave from the snapshot.
- mikealrogers 16y agowait, seriously? the suggested default step 1 for MongoDB is to acquire 3 servers? i mean, no other database suggests such a huge default configuration. even knowing that their datacenter can get hit by lightening and all a lot of large production sites don't even run with this kind of redundancy. this seems like a pretty taxing workaround for not keeping an append-only transaction log.
- mathias_10gen 16y agoNo, it is a suggested configuration, not the suggested one. And recommending three nodes is not that uncommon for distributed systems, because you cant have a quorum with only two nodes. Some users are more concerned with handling 10s or 100s of servers than having single-server durability. That said, for most users two servers are fine. Other users don't need any replicas at all since they do a nightly dump from a stored source into mongodb, or they just take regular backups. There are many ways to achieve system-wide durability, not all of them require the database to be durable.
- didip 16y agoI believe Cassandra also suggest 3 nodes configuration.