6 ms·
Riak! We're ( http://bu.mp http://bu.mp ) using more Riak every day. So far so good.
by jamwt 15y ago
Riak! We're ( http://bu.mp http://bu.mp ) using more Riak every day. So far so good.
- kuviaq 15y ago+1 for Riak, it's suiting our needs very well so far.
- aonic 15y agoWhat kind of data are you storing in Riak? And is it write or read heavy usage?
- jamwt 15y agoAll kinds of stuff, mostly ~1-4k values. Read:Write is some low integer, ~2.
- devongall 15y agoLoving Riak as well! Our balance is definitely on the write-heavy side of things.
- grourk 15y agoDitto (http://dropc.am http://dropc.am). Very write heavy load for us, which Riak handles without blinking. Fault tolerant, robust, and easy to administer. Every machine is identical, no special "master" nodes or anything like that.
- grncdr 15y agoI've heard this about Riak and I was quite excited to test it out for a new project, but in the limited testing I've done Cassandra and HBase both absolutely smoke Riak in terms of write performance. Not really apples to apples I suppose, but I was really surprised at how slow Riak was when handling many (millions) of small writes. We haven't finished our testing/profiling phase yet, so any hints on how to optimize a large number of small writes (on the order of ~dozen bytes each) would be appreciated.
- jamwt 15y agoWithout going too much into specifics and picking on individual databases (which I could do, boy do I have the scars...) When you hit a certain traffic level, scalability, latency and robustness become far more important than single-node ops/s. I need to be able to add nodes and repair failed nodes while under load--I need the 99.9% latency mark to stay ~100ms while doing so. I don't really care how many bajillions of ops a second your database can do in some concocted scenario, b/c you're not going to do that many in the real world anyway (trust me, we tried). The disk subsystem is going to give you a few hundred, maybe a few thousand if you're lucky, IOPS, then your latency will spike to hell and your phone will wake you up at night. Maybe in the world where 99% of ops are reads, you will put up impressive numbers, but now you're just showing you are pretty good at using the disk cache. That's a relatively easy problem. The riak guys seem to get all this better than most: http://blog.basho.com/2011/05/11/Lies-Damn-Lies-And-NoSQL/ http://blog.basho.com/2011/05/11/Lies-Damn-Lies-And-NoSQL/ So, to give you a short answer to your direct question: Use SLC SSDs + md + RAID-0. Have at least 5 nodes. Use bitcask, but realize that your keys will need to fit in memory. Also, realize that really small values aren't a great fit for Riak in some ways b/c the overhead per value is at least a few hundred bytes. Also, it's important to note this is where I'm at right now, but maybe not where you (generally) are at. Riak may not make you happy at server #1, but it will make you pretty happy at server 10 and server 100. Riak's sweet spot is people with scaling pains. If you only need a server or two to try some stuff, and you don't have any users yet, you might cause yourself more headaches than you need. Sometimes you don't need a locomotive, you need a motorcycle. (These guys have a pretty great motorcycle: http://rethinkdb.com/ http://rethinkdb.com/ )
- bretthoerner 15y agoWas your experience with Cassandra different? Happy at server 10 and 100, that is?
- jamwt 15y agoTBH, we didn't seriously pursue Cassandra when we considered distributed database systems b/c the vast majority of "back-reference checks" we did on the YC network and other area startups was "stay away." We got some very frank advice from some people whose opinions on databases I take very seriously to stay away, including reports from within FB. Having said that, I cannot claim to have firsthand proven or disproven anything about Cassandra.
- jarin 15y agoI haven't had a chance to use Riak in a production application yet but am definitely looking for the first possible excuse to use it.
- duffomelia 15y agoWe're using it in production and loving it.
- dsl 15y agoI love Riak. It's become my go to for "this just has to work" (and I actually work on problems that need to scale, not ones I hope will have to scale). The only improvement you could make to it would be adding some of the fancier bits that make Redis really nice, like sets and lists.
- StavrosK 15y agoWhich datastore is it closer to? Mongo, Redis, Postgres? I haven't looked much into it, but I hear so many good things that maybe I should.
- jamwt 15y agoCassandra. It's an eventually consistent, fully-distributed database, in the Dynamo mold: http://www.allthingsdistributed.com/files/amazon-dynamo-sosp2007.pdf http://www.allthingsdistributed.com/files/amazon-dynamo-sosp...
- StavrosK 15y agoThank you.
- dsl 15y agoIt's Amazon S3, but without the outrageous per query costs.
- sjs 15y agoHear hear! Mongo but no Riak? Come on.
- rgnitz 15y agoRiak isn't in the same ballpark as MongoDB. They are closer to Hadoop.
- sjs 15y agoMySQL and PostgreSQL aren't in the same ball park as Mongo. It's irrelevant. Based on merits, architecture, and implementation Riak handily beats several of the DBs listed. It's a glaring omission and not the only one. As many chose Other as did Oracle.