5 ms·
They're doing what works for them and good for them for that. But...I think a LOT of people are really missing out by passing over Riak. Many of the issues the
by nirvana 14y ago
They're doing what works for them and good for them for that. But...I think a LOT of people are really missing out by passing over Riak.
Many of the issues they found with CouchDB have been resolved with Riak. I think the sync API for CouchDB is really cool, but Riak has the auto-sharding thing down cold.
Riak runs map reduce queries across multiple nodes, so performance and capability can grow as you add nodes.
CouchDB's views are neat but they impose some constraints that Riak's more dynamic approach resolves (at the cost of possibly running more queries, but these results can be cached easily giving Riak a form of "views" for often run queries.)
I believe Riak's choices for backend are superior to CouchDB's. Further, Riak supports multiple backends so you can choose the one appropriate for your service (including InnoDB, LevelDB and Basho's Bitcask, as well as a super secret hidden gem of a Caching RAM backend.)
Riak now has indexing of data, and queries on these indexes, but I can't compare it to CouchDB. I can say that the feature is close enough for me to not miss SQL.
I think Riak's "view performance" compared to CouchDB should be good, but may not compare to MySQL, but then, we're talking single node performance. Riak is distributed- you need more performance, you just add nodes and point them at the cluster. MySQL requires you to architect a (from my perspective) brittle configuration of servers that can run into SPF issues.
For instance they talk about having a single write master. What happens when a meteor crashes to earth and takes out that machine? Really unlikely, sure, but I have had enough machines have failures (and failures are often really weird) that I don't trust ANY machine to be a single point of failure. ... and when I'm forced to, like being in a single datacenter or having a single network switch, I don't like it, so I avoid it when I can.
Riak has automatic sharding and automatic rebalancing. It loses a node and keeps running. You add nodes and it redistributes around. Riak is an operational dream.
Not to bash CouchDB at all (or MySQL). I think CouchDB is a great product for certain use cases.
I just think a LOT of people are really missing out by passing over Riak.
- dgrnbrg 14y agoDoes riak support range queries?
- smn 14y agoIt does on secondary indexes http://wiki.basho.com/Secondary-Indexes.html http://wiki.basho.com/Secondary-Indexes.html
- aphyr 14y agoI should mention that 2I performance may be a little slow, depending on what kind of indices and queries you need. It's not hard to try it out and benchmark, though.
- nirvana 14y agoYes, it does range queries based on key, if the key follows a defined format. It also has secondary indexes.
- rdtsc 14y agoCouchDB has some features that other databases don't have: a continuous changes feed, REST interface, master-to-master replication, a web interface to the data and management (Futon). We need those features (Yes including Futon. It is a feature because it lets us quickly prototype and debug. It makes a black box that you drop your data in transparent). But we not using to for large datasets. We are using it mostly for configuration, and setup. There is a custom clustering setup built in its m-m replication and changes feeds. But I agree that for IO scaling and large data sets Riak would be a top choice. But there are other contenders to look at as well: Cassandra, BigCouch and the upcoming Couchbase Server 2.0
- nirvana 14y agoEvery database is different, but that doesn't mean there are a lot of really unique features. The unique feature of CouchDB is the way it does views, but that doesn't mean you can't' do views (in fact, in my opinion, better) in other databases. Continuous changes feed- you can get this with Riak and more importantly you can get a feed of just the relevant changes. Plus you don't need this in Riak the way you do in CouchDB because Riak already has distribution built in. Master-Master replication as done in CouchDB is inferior to the turely distributed database that Riak is. (Eg: its not replication, it is distributed itself.) Web interface to data management-- futon was a lead here but there are several tools for Riak that cover these bases in my opinion. I think CouchDB as a configuration database is an excellent job. I looked at BigCouch which is taking the Dynamo Ring concept and applying it to CouchDB which is a good solution for couchDB (in fact, they should build it in to the core) ... but that's also what Riak is built from the ground up to be (a dynamo ring.) Cassandra is a different animal and I can never figure out what Couchbase is going to become.
- mpd 14y agoI investigated using Riak for dealing with our metrics a few months ago, but with the data sizes we are dealing with, even the Riak people told us that Hadoop was likely a better solution. Once you are dealing with more than 500k keys or so, Riak starts to fall over. EDIT: The 500k key limit pertains to mapreduce jobs, not the overall data size.
- astrodust 14y agoThat doesn't seem like a very large number. Are you sure?
- mpd 14y agoYes. I should clarify that I meant 500k keys used in a single m/r job. We needed to be able to run m/r over roughly 200 million keys at the time.
- nirvana 14y agoAnd it turns out you are misrepresenting the situation completely. You can run M/R over key sets in the billions of keys. It sounds like you've not organized your data at all. You're bashing a product here based on your lack of knowledge, not the products lack of capabilities.
- dmpk2k 14y agoBased on what I've seen for some internal things, that's one claim I'd like to see support for. m/r on Riak has been an unmitigated disaster here for anything beyond incredibly trivial working sets.
- aphyr 14y agoI've never seen a Riak MR job over more than 3 million keys complete, on a 6-node SSD cluster. It might be possible, but you'd have to throw a lot more HW at it than the comparable Hadoop setup.
- heretohelp 14y ago
- dochtman 14y agoI looked at Riak a bit (we use CouchDB and Redis at our startup). It seemed to not scale down as nicely as CouchDB: no Futon, binary API's and a lot of emphasis in the documentation on sharding stuff (where we just run CouchDB on a single server with a backup server where we continuously-replicate to).
- pixelcort 14y agoOne of the nice features of CouchDB is that the views are incrementally updated, ideal if you have a large dataset that changes frequently with little changes and you want frequently get the most up to date transformed data. Looking into Riak, I am under the impression there is some caching done on parts of their map reduce system, but I couldn't really find a lot of advice on the performance characteristics of frequently running a map reduce over a large, slowly-changing dataset. Does anyone else have experience using Riak for this kind of thing?
- aphyr 14y agoIn short, don't. In Riak you would aim to update those views at write time, to minimize the number of reads required.
- mattbriggs 14y agoI think the eventual consistency is what narrows riaks use case. Only heard good things about it though