5 ms·
I'm hoping this'll be a viable replacement for MongoDB. (Sparse/Schema-free is incredibly useful for me, as is JSON-centric modeling) jedberg already asked for
by codewright 14y ago
I'm hoping this'll be a viable replacement for MongoDB. (Sparse/Schema-free is incredibly useful for me, as is JSON-centric modeling)
jedberg already asked for a compare/contrast, but let me provide some specifics I care about that you might be able to answer.
1. Is it fair to say that thanks to MVCC, running an aggregation or map-reduce job isn't going to lock the whole damn thing up like it does on MongoDB?
2. You've got a distributed system that is seemingly CP, do the availability/consistency semantics compare with HBase? Master-slave? Replication? Sharding?
3. Latency is a big one for us and is a large part of why we use ElasticSearch. How does the read-latency on RethinkDB compare with Mongo/MySQL/Redis/et al ?
- coffeemug 14y ago1. Yes -- that was the main motivation for MVCC. We wanted to allow people to use rethinkdb for analytics and map/reduce on top of the realtime system without dealing with having to replicate data into something else. 2. Short answer: we favor consistency (via master/slave under the hood). It allows for much easier API, much fewer issues in production, etc. The user experience is just better. If you're ok with out of date results, you can do that too without paying the price of consistency guarantees. The downsite of our design is that you might lose write availability in case of netsplits (if the client is on the wrong side of the split). Longer answer: checkout the FAQ at http://www.rethinkdb.com/docs/advanced-faq/ http://www.rethinkdb.com/docs/advanced-faq/ 3. Read latency should be equivalent to other comparable master/slave systems. We don't do quorums, so latency will be much better than quorum/dynamo-based designs.
- rbranson 14y agoI want to preface my comment: this is impressive work, congratulations on shipping, and this is what MongoDB should have been from the start. In reality, most transactional database deployments are heavily skewed towards read workload, so reading from hot slaves is basically a requirement for master/slave databases. So, in most real world applications at scale, apps already deal with inconsistencies between slaves and the master and are making the "difficult" choice of dealing with CAP trade-offs. Asynchronous replication also creates a potential for difficult or impossible to recover from data loss in the sense that masters & slaves always have a continuous possibility for split-brain. RethinkDB does not provide multi-shard transaction atomicity and/or isolation, which in my experience is the biggest difficulty thrown up in front of developers coming from single-node databases. I feel like the difficulty of dealing with inconsistencies across multiple versions of a single object is far more familiar as most developers have at least dealt with cache invalidation in some form. It's really having to ensure and deal with potentially out of order operations (inconsistency in the ACID sense) across a "graph" of data that's more insidious.
- coffeemug 14y agoI mostly agree with what you're saying, but I also think there's enormous value in making easy things be really easy. Even with today's state of the art adding a shard, dealing with consistency issues, adding replicas, etc. is relatively hard. Perhaps not in a computer-sciency sense (all the problems are fairly well understood), but in an operational sense. Lots and lots of work needs to be done even with systems like MongoDB, let alone with MySQL. And once you're done with that, you can't really run complicated queries easily, so you have to solve that problem. We set out to make these things be really easy (whether we succeed or not remains to be seen). We want the users not to have to deal with these issues at all whenever possible. You should be able to set up a cluster, add shards, and run cross-shard joins and aggregation in five minutes. Of course once that problem is solved, there are tougher problems like high-performance cross-document distributed ACID, but I think the industry as a whole is relatively far away from that right now. (there are some solutions to this - e.g. Clustrix, but they require specialized hardware which makes it out of reach for most developers)
- nlavezzo 14y agoCongratulations on shipping - looks like a very well thought out product. Regarding your last comment on high-performance distributed ACID, that's what we've built at FoundationDB, although FoundationDB is a key value store so transactions are multi/cross-key instead of cross-document.
- Nitramp 14y agothere are tougher problems like high-performance cross-document distributed ACID, but I think the industry as a whole is relatively far away from that right now Megastore and Spanner solve that problem, with varying tradeoffs: http://research.google.com/pubs/pub36971.html http://research.google.com/pubs/pub36971.html http://research.google.com/archive/spanner.html http://research.google.com/archive/spanner.html
- erichocean 14y agoOur internal database does too (with a different design than Spanner, but stuff still comes "online" atomically for everyone across the globe at the same time, with similar latency). Unlike FoundationDB, and like Spanner, we're doing it with complex object graphs, not just key-values, and we also do it with consistent secondary indexing (I'm not sure if Spanner supports this or not). This isn't "the future", this is now. People are doing it, and have been for awhile. If you're going to "rethink the database", distributed global consistency should be at the top of your list today. RethinkDB seems like its merely "rethinking" Mongo. The main benefit of global consistency, of course, is ease of use. Global consistency is so much easier to reason about and write code for!
- StavrosK 14y agoBy the way, how safe is the JS interpreter? Can you get into trouble by running untrusted code in map/reduce queries?
- coffeemug 14y agoJS interpreter (V8 under the hood) runs in a process pool -- similar to a thread pool, but outside of the core rethinkdb process. The code running in the JS interpreter cannot corrupt memory or crash the rethinkdb process (if it crashes, rethinkdb will simply start another v8 process). You also can't write from js executed on the server, so the data is safe (though I think it's more of a limitation than a feature). Currently if you write an infinite loop in js, or write code in a way where it starts eating up memory we don't do anything to restart the js process, but it would be relatively easy to implement.
- StavrosK 14y agoI see, thanks. I'm definitely interested in that, as I'm developing http://www.instahero.com http://www.instahero.com and the current approach isn't very scalable, so I'm evaluating alternatives. RethinkDB looks like a good candidate so far.