8 ms·
RocksDB – A persistent key-value store for fast storage environments
- gfodor 13y agothis is cool, though I'd wonder how it compares to Kyoto Cabinet. another big issue I've run into personally is the fact that both LevelDB and KC don't explicitly support multiple processes reading the db at once. (KC's API allows this but advises against it, LevelDB afaik doesn't even allow it.) I wonder if RocksDB gets past this.
- maaku 13y agoHyperDex?
- stass 13y agoIf you just need concurrent reading and update the database from a single process, you can use CDB. It always served me well.
- dhruba_b 13y agoRocksDB allows only one process to open the DB in read-write mode but other process can access the DB read-only (with a few configuration settings)
- hyc_symas 13y agoKyoto Cabinet will self-corrupt if you use it that way. LMDB supports multi-process explicitly.
- gfodor 13y agocan you explain how this happens? if it's just a read-only process, how can it corrupt anything?
- rdtsc 13y agoWell LevelDB is already good. And if this improves on it, that's great. I was looking at embedded key value stores and also found -- HyperLevelDB (from creators of Hyperdex database). They also improved on LevelDB in respect to compaction and locking: http://hyperdex.org/performance/leveldb/ http://hyperdex.org/performance/leveldb/ So now I am curios how it would compare. Another interesting case optimized for reads is LMDB. That is a small but very fast embedded database at sits at the core of OpenLDAP. That one has impressive benchmarks. http://symas.com/mdb/microbench/ http://symas.com/mdb/microbench/ (Note: LMDB used to be called MDB, you might know it by that name).
- hyc_symas 13y agoLSMs have a long long way to go to catch up to LMDB. http://symas.com/mdb/hyperdex/ http://symas.com/mdb/hyperdex/
- rdtsc 13y agoHoward, is LMDB effectively limited to 128T (on 64bit machines and 2GB on 32bit ones, not that one should be running large databases on 32bit machines)? Also what about concurrent writes? Does it have a database wide writer lock or is it per key (per page?) ?
- hyc_symas 13y agoIt is limited to the logical address space. Since most current x86-64 machines have only 48bit address space, 256TB, and assuming the kernel keeps half of the space for itself, then yes, the current limit is 128TB. But I suspect we'll be seeing 56bit address spaces fairly soon. It is a single-writer DB, one DB-wide writer lock. Fine-grained locking is a tar pit.
- rdtsc 13y agoMakes sense. Most impressive about LMDB to me is the zero-copy model for readers, with is no extra memcpy needed, maybe that is something obvious for database gurus but it is pretty clever trick I think.
- hyc_symas 13y agoIt's pretty significant, yes. Eliminating multiple copies of everything got us a 4:1 reduction in memory footprint in OpenLDAP slapd (compared to our BerkeleyDB-based backend). This is another reason we don't spend too much time worrying about data compression and I/O bound workloads - when you've essentially expanded your available space by a factor of 4, you get the same benefits of compression, without wasting any of the memory or CPU time. And when you can fit a 4x larger working set into your space, you find that you need a lot less actual I/Os.
- _kst_ 13y agoA very minor point: The illustrative code snippet on the home page has a spurious semicolon on the first line: #include <assert>;
- jamesgpearce 13y agofixed! - thanks
- dhruba_b 13y agoHi guys, I am Dhruba and I work in the Database Engineering team at Facebook. We just released RocksDB as an open source project. If anybody has any technical questions about RocksDB, please feel free to ask. Thanks.
- mml 13y agoAny replication support? (Or any sort of distribution?)
- jonstewart 13y agoThe primary enhancements over LevelDB seem to be parallel compactions of disjoint ranges, to take advantage of cheap seeks on flash storage, and the ability to parameterize core algorithms and data structures to suit a particular anticipated workload. All very cool; anything else major? Also, there aren't JNI bindings... are there? Thanks for the contribution. Just started using LevelDB on a project, but deployment will involve fast flash storage and rocksdb looks like a worthy successor. Jon
- dhruba_b 13y agoThanks for your comments Jon. RocksDB shares some of its genes with LevelDB.. something like a parent-child relationship. Please check out Universal Comaction Style, multi-threaded-compaction, pipelined memtables. I used to have JNI bindings that I pulled in from https://github.com/fusesource/leveldbjni https://github.com/fusesource/leveldbjni but it was difficult for me to update the JNI everytime we added new apis to RocksDB. It would be great if somebody who needs Java support can implement JNI bindings for RocksDB.
- ashwinmurthy 13y agoCan it be configured as a distributed No SQL database like Cassandra?
- techtalsky 13y agoCan this be used as a drop-in replacement for LevelDB on queue technologies like ApolloMQ that use LevelDB as the default?
- deleted 13y ago[deleted]
- snewman 13y agoVery nice work, and the wiki is also quite nice -- I wish more projects had a page like https://github.com/facebook/rocksdb/wiki/Rocksdb-Architecture-Guide https://github.com/facebook/rocksdb/wiki/Rocksdb-Architectur.... It's really nice to see a clear, terse summary of what makes this project interesting relative to its predecessors. At my company (scalyr.com), we've built a more-or-less clone of LevelDB in Java, with a similar goal of extracting more performance on high-powered servers (and better integration with our Java codebase). I'll be digging through rocksdb to see what ideas we might borrow. A few things we've implemented that might be interesting for rocksdb: * The application can force segments to be split at specified keys. This is very helpful if you write a block of data all at once and then don't touch it for a long time. The initial memtable compaction places this data in its own segment and then we can push that segment down to the deepest level without ever compacting it again. It can also eliminate the need for bloom filters for many use cases, as you often wind up with only one segment overlapping a particular key range. * The application can specify different compression schemes for different parts of the keyspace. This is useful if you are storing different kinds of data in the same database. * We don't use timestamps anywhere other than the memtable. This puts some constraints on snapshot management, but streamlines get/scan operations and reduces file size for small values. Do you have benchmarks for scan performance? This is an important area for us. I don't have exact figures handy, but we get something like 2GB/second (using 8 threads) on an EC2 h1.4xlarge, uncached (reading from SSD) and decompressing on the fly. This is an area we've focused on. I'd enjoy getting together to compare notes -- send me an e-mail if you're interested. steve @ (the domain mentioned above).
- hyc_symas 13y agoSkyDB using LMDB gets 3GB/sec on a standalone PC. https://groups.google.com/forum/#!msg/skydb/CMKQSLf2WAw/zBO1X35alxcJ https://groups.google.com/forum/#!msg/skydb/CMKQSLf2WAw/zBO1...
- bjconlan 13y agoWow, Awesome link, LMDB always seems to fly under the radar, SkyDB+LMDB. Genius. (and written in go! I'm sold... well will at least give it a bash)
- parshap 13y agoNode.js bindings (compatible with levelup) have already been released by rvagg: https://npmjs.org/package/rocksdb https://npmjs.org/package/rocksdb
- Patient0 13y agoI'm surprised that the C++ code is not using the RAII idiom in some obvious places. For example: https://github.com/facebook/rocksdb/blob/master/db/db_impl.c https://github.com/facebook/rocksdb/blob/master/db/db_impl.c There are many places with bracketed calls to mutex_.Lock and mutex_.Unlock(). An example: mutex_.Unlock(); LogFlush(options_.info_log); env_->SleepForMicroseconds(1000000); mutex_.Lock() Why didn't the authors use the RAII idiom here? Even if there are no exceptions expected, the code would still be simpler and less error prone by using a guard object.
- tsewlliw 13y agofixed your link: https://github.com/facebook/rocksdb/blob/master/db/db_impl.cc#L1665 https://github.com/facebook/rocksdb/blob/master/db/db_impl.c... Take another look! There's a guard object used at the function scope to ensure the lock is released, and this block is bracketed to release and reacquire the lock, not acquire and release. There may be a case for a guard object that does the release/reacquire, but its definitely not a slam dunk like acquire/release
- phunge 13y agoStill, that's not exception-safe, correct? If LogFlush or SleepForMicroseconds throws an exception the mutex will be unlocked twice, which pthreads disallows for normal mutexes...
- deleted 13y ago[deleted]
- cbsmith 13y agoYou know, for a second I thought you were wrong, but I changed my mind. This does look like a bug, and a simple on to avoid at that. It's tough, because Rocks is still highly based on LevelDB, which conforms to Google's coding style guideline, which makes RAII more than a bit tricky to do right.
- e12e 13y ago
- wbolster 13y agoThe benchmark at https://github.com/facebook/rocksdb/wiki/Performance-Benchmarks#2-bulk-load-of-keys-in-random-order https://github.com/facebook/rocksdb/wiki/Performance-Benchma... states that for LevelDB, "in 24 hours it inserted only 2 million key-values", and that "each key is of size 10 bytes, each value is of size 800 bytes". I might be missing something, but that took just a few minutes on my ~2 year old desktop machine. Sample code: https://gist.github.com/wbolster/7487225 https://gist.github.com/wbolster/7487225
- dhruba_b 13y agoThere was a typo, the 2 million should have been 200 million keys. I fixed the wiki page. Thanks again for pointing it out.
- arthursilva 13y agoLooking forward to see this in Riak
- canadi 13y agoTnx for all the comments! Feel free to continue the discussion at https://www.facebook.com/groups/rocksdb.dev/ https://www.facebook.com/groups/rocksdb.dev/