5 ms·
I investigated using Riak for dealing with our metrics a few months ago, but with the data sizes we are dealing with, even the Riak people told us that Hadoop w
by mpd 14y ago
I investigated using Riak for dealing with our metrics a few months ago, but with the data sizes we are dealing with, even the Riak people told us that Hadoop was likely a better solution.
Once you are dealing with more than 500k keys or so, Riak starts to fall over.
EDIT: The 500k key limit pertains to mapreduce jobs, not the overall data size.
- astrodust 14y agoThat doesn't seem like a very large number. Are you sure?
- mpd 14y agoYes. I should clarify that I meant 500k keys used in a single m/r job. We needed to be able to run m/r over roughly 200 million keys at the time.
- nirvana 14y agoAnd it turns out you are misrepresenting the situation completely. You can run M/R over key sets in the billions of keys. It sounds like you've not organized your data at all. You're bashing a product here based on your lack of knowledge, not the products lack of capabilities.
- dmpk2k 14y agoBased on what I've seen for some internal things, that's one claim I'd like to see support for. m/r on Riak has been an unmitigated disaster here for anything beyond incredibly trivial working sets.
- aphyr 14y agoI've never seen a Riak MR job over more than 3 million keys complete, on a 6-node SSD cluster. It might be possible, but you'd have to throw a lot more HW at it than the comparable Hadoop setup.
- heretohelp 14y agoI see you're irrationally proselytizing again. We just saw you over in the programming languages thread, now you're here, refusing to confront the reality of how broken M/R is in Riak. What monkey crawls around on your back to make you so confrontational and irrational?
- nirvana 14y agoI think your statement is both out of date and quite broad. Plus you imply that a database would be limited to 500k keys which is silly, when you really mean a map-reduce job. And further, are you really doing MR over your entire dataset, all the time, or would key filtering, ranges or secondary indexes be a better fit? Its easy to do M/R in Riak over only the correct amount of data. Hadoop may have been better for what you are doing, and logging metrics is a particular use case where a specialized database is most appropriate. But it is incorrect to imply that Riak falls over at any specific key limit. This is simply untrue. With Riak you can always add more nodes if you need more capacity, and Map Reduce is done in a distributed fashion so adding more nodes adds map reduce capacity. Its not perfect but it is not brittle.
- mpd 14y agoYou're welcome to peruse the mailing list thread. http://lists.basho.com/pipermail/riak-users_lists.basho.com/2011-November/006668.html http://lists.basho.com/pipermail/riak-users_lists.basho.com/...
- nirvana 14y agoThat thread shows that all of the particulars of your claims about Riak are actually false. Further it seems you didn't bother to understand how Riak can solve your problem and thus decided that it cannot.
- cynicalkane 14y agoVerbatim, from the mailing list: "If large-scale mapreduce (more than a few hundred thousand keys) is important, or listing keys is critical, you might consider HBase." "Riak can also collapse in horrible ways when asked to list huge numbers of keys. Some people say it just gets slow on their large installations. We've actually seen it hang the cluster altogether. Try it and find out!"
- deleted 14y ago[deleted]
- aphyr 14y agoI feel somewhat responsible for this confusion, as the guy being quoted here... :-( Riak will handle billions of keys just fine. We had, I dunno, a half a billion in a six node bitcask-backed cluster and were only at half capacity. Much much bigger installs exist. The limit I was referring to is for a single mapreduce job; Riak MR just isn't well-suited to operations over millions of keys at a time. It can do it, but Riak MR isn't really designed for bulk processing: and I wouldn't be surprised to see MR become unusably slow over millions of keys. You'll get better performance out of Hadoop, generally, for bulk analytics. The other tough point is key-listing. Listing buckets, listing keys, key filters, MR over buckets, all those features are essentially useless in production. Where the number of keys is large and unguessable it can become a logistical nightmare to keep track of them. 2I key indexes can help, though.
- ithkuil 14y agoI have a use case, which I don't know if it's common or not. I want to put millions of items in riak, play with it, and then throw then away. I might want to do that because I'm testing out something, or because it's the result of some periodic batch processing in production, which I want to get by key later. Unfortunately, riak doesn't seem to have the notion of a "db", "keyspace" or whatever you want to call it; i.e. something which you could "drop" and that will simply delete a directory with a dozen of files in it (should be quite cheap). The only thing I can do is to drop the whole riak db, which has the following problems: 1) I have to do it manually on all nodes (stopping the cluster, deleting the files etc) 2) I cannot share a riak cluster between several users/team, so that each user/team can play with a portion of it but there is only a central installtion of the whole cluster. Every application (which I want to be able to drop all the db and recreate it) has to run it's own riak cluster. Initially I thought that "buckets" were intended to solve this problem, but buckets don't map to a separate storage location, it's just a way to group items. Even listing all buckets present in the db requires scanning all keys and, as the doc says, "Similar to the list keys operation, this requires traversing all keys stored in the cluster and should not be used in production." Although I've been told that "riak is not designed to do this and that", I'm not sure if these limitations are really technical, or just because the product development effort was targeted at some of the aspects, and these issues could be addressed in a later stage. Any idea?