13 ms·
MDBM – High-speed database
- justin66 12y agoThis looks interesting. At this stage of the game a more meaningful benchmark might involve LMDB, Wiredtiger, and, yes, LevelDB.
- hendzen 12y agoI don't think it's comparable to benchmark MDBM against LMDB or WiredTiger as keys are not kept in sorted order (no range queries), there is no support for transactions, and MDBM does not offer durability in the event of power loss. MDBM is pretty much an optimized persistent hash table. LMDB and WiredTiger aim to be full-fledged ACID compliant database storage engines with functionality similar to that of BerkeleyDB or InnoDB.
- justin66 12y agoIf you look at the MDBM page you'll notice that they currently benchmark against LevelDB, BerkeleyDB, and Kyoto Cabinet. What I was suggesting involves the same theme but better, newer competition. I agree that it's an apples to oranges comparison in any case.
- deleted 12y ago[deleted]
- beagle3 12y agoThe timings they give there can only make sense on a fast SSD or when the database they benchmark on is completely cached. It's an apples-to-oranges comparison only of MDBM wins significantly against LMDB. If they are comparable in timing, or e.g. MDBM is 20% faster, then it would be an apples-to-apples comparison, MDBM having 20% speed advantage, and LMDB having every other possible advantage (memory safety, ACIDity, ordered retrieval, multiple databases, etc.) LMDB is truly, incredibly, really marvelous. On 64-bit it comes close to being the end-all-be-all local KV-store. If your databases are not more than a few tens of megs each, the same is true for 32-bit processors as well.
- justin66 12y agoYes, I am quite impressed by it as well.
- rakoo 12y ago> If your databases are not more than a few tens of megs each, the same is true for 32-bit processors as well. Why is that ? Shouldn't 32-bit processors give you enough space in the range of hundreds of MiB ?
- hyc_symas 12y agoSure. But you only have 2-3GB of usable address space, and if your app has a lot of other variables to work with, there's not much space left over. I should note that in LMDB 1.0 we'll have dynamic unmapping and remapping, to allow 32-bit machines to work with larger DBs. (There's still a significant performance cost for this. It's only being done to allow folks to use the same code on 32 and 64.)
- jetm9 12y agoOff-topic: will you have something like networked KV store aka redis. what do you think of Redis or Cassandra?
- hyc_symas 12y agoFor redis, there's already ardb, ledisdb, redis-NDS, and a few other redis clones that support LMDB. We also already have memcacheDB on LMDB, as well as HyperDex with LMDB. I haven't looked closely at Cassandra since it's in java, and after I didn't find a simple backend plugin API I didn't look any further.
- jbooth 12y agoIf you're interested in building a distributed DB on top of LMDB, I took care of part of the problem (replication/consistency) in my flotilla library at https://github.com/jbooth/flotilla https://github.com/jbooth/flotilla Basically I'm just layering the raft consistency algorithm on top of LMDB. Both systems single-thread write transactions for consistency, so there's some mechanical sympathy. Doesn't mandate any specific data model or even a client-server network transport, it's basically a replicated embedded DB. Anyone could build replicated redis on top of it or a Cassandra clone if they want to get into managing shards/rings. Sample app (still WIP) at https://github.com/jbooth/merchdb https://github.com/jbooth/merchdb
- hyc_symas 12y agoYou make some good points. We benchmark LMDB against LevelDB and its derivatives even though none of the LevelDB family offer ACID transactions. (http://symas.com/mdb/ondisk/ http://symas.com/mdb/ondisk/ ) Despite this fact, people will ask the question and try to make the comparison, so we run those tests. It's silly, but most people seem to pay attention to performance more than to safety/reliability. From my totally biased perspective, MDBM is utter garbage. They use mmap but make absolutely zero effort to use it safely. This was the biggest obstacle to overcome in developing LMDB; I had a few lengthy conversations with the SleepyCat guys about it as well. It's the reason it took 2 years (from 2009 when we first started talking about it, to 2011 first code release) to get LMDB implemented. If you want to call something a "database" you have to do more than just mmap a file and start shoving data into it - you have to exert some kind of control over how and when the mapped data gets persisted to disk. Otherwise, if you just let the OS randomly flush things, you'll wind up with garbage. As Keith Bostic said to me (private email): "The most significant problem with building an mmap'd back-end is implementing write-ahead-logging (WAL). (You probably know this, but just in case: the way databases usually guarantee consistency is by ensuring that log records describing each change are written to disk before their transaction commits, and before the database page that was changed. In other words, log record X must hit disk before the database page containing the change described by log record X.) In Berkeley DB WAL is done by maintaining a relationship between the database pages and the log records. If a database page is being written to disk, there's a look-aside into the logging system to make sure the right log records have already been written. In a memory-mapped system, you would do this by locking modified pages into memory (mlock), and flushing them at specific times (msync), otherwise the VM might just push a database page with modifications to disk before its log record is written, and if you crash at that point it's all over but the screaming." The harsh realities of working with mmap are what dictated LMDB's copy-on-write design - it's the only way to ensure consistency with an mmap without losing performance (due to multiple mlock/msync syscalls). None of these design considerations are evident in MDBM. LMDB's mmap is read-only by default, because otherwise it's trivial to permanently corrupt a database by overwriting a record, writing past the end, etc. MDBM's mmap is read-write, and the only "protection" you get is a doc that tells you "be Vewwy vewwy careful!" Ridiculously sloppy. LMDB's design and implementation are proven incorruptible. MDBM (and LevelDB and all its derivatives) are proven to be quite fragile. https://www.usenix.org/conference/osdi14/technical-sessions/presentation/pillai https://www.usenix.org/conference/osdi14/technical-sessions/... Leaving reliability aside for a moment, there's also the issue of performance and efficiency. We used to use DBM-style hashes for the indexes in OpenLDAP, up to release 2.1. We abandoned them in favor of B-trees in OpenLDAP 2.2 because extensive benchmarking showed that BDB's B-trees were faster than its hash implementation at very large data sizes. The fundamental problem is that hash data structures are only fast when they are sparsely populated. When the number of data records you need to work with increases to fill the table, you start getting more and more hash collisions that result in lots of linear probes (or whatever other hash recovery strategy you're using). The other problem is that the very sparse/unordered nature of hashes makes them extremely cache unfriendly - you get zero locality-of-reference for groups of related queries. So as your data volumes increase, you get less and less benefit from the amount of RAM you have available. When the data exceeds the size of RAM, the number of disk seeks required for an arbitrary lookup is enormous, and every read is a random access. Using a hash for a large-scale data store is just horrible. (We tested this extensively a decade ago http://www.openldap.org/lists/openldap-devel/200401/msg00077.html http://www.openldap.org/lists/openldap-devel/200401/msg00077... )
- extralam 12y agointeresting. follow
- philliphaydon 12y agoDo people get annoyed by all the JavaScript frameworks and Databases coming out in regards to adoption from a company point of view? I mean every other day a new database comes out and claims to be better in one way or another than something else and then its like "fuck I picked X when now there's Y" It seems over the last year technology has been growing more rapidly than any other period. Fun times but so hard to keep track of everything!
- akbar501 12y ago> It seems over the last year technology has been growing more rapidly than any other period. I tend to agree with this statement. The entire stack appears to be going through a revolution. The data layer in particular is seeing very rapid change after being largely (not entirely) static for decades.
- tacos 12y agoFor those old timers who did distributed systems work there's not that much new under the sun. I look at something like this and say "ah, a quirky and somewhat dangerous cache layer." What's different is bloggers promoting it as a "database." While I'm sure someone out there will see this and say "wow, that's exactly what I need!" chances are that if you have these sorts of scale issues you're going to have to figure it out on your own. I'd rather see a write-up of how they arrived at this particular conclusion than another non-database.
- optimusclimb 12y agoWhile I understand that feeling, I've come to realize it's an inevitability, that shouldn't affect your work. At any given time, you either have a need/problem, or you don't. If you DO, you evaluate the current tech available, and hopefully select something that fits your needs. You build out around said tech, and if your choice was correct, that means it's either solving your problem, or on it's way to. If something comes along while you're implementing with your chosen solution, that looks similar, but better, it's only noise - because hey, you found a solution. Just as we don't all re-write all of our code whenever a new language comes along (unless the thing in question was desperately in need of a re-write anyway) even if newer languages are nicer, we needn't switch DBs or frameworks for the same reasons.
- coreymgilmore 12y agoThoughts on using this as a cache instead of memcache or redis? Yes, it does not have nearly as many features or functions but when raw performance is needed I could see this working (given an api for using this via Node.JS, PHP, etc.).
- hendzen 12y agowhy even pay the cost of memory mapping if its a transient embedded cache not shared between servers? just use a std::unordered_map, or better yet a tbb::concurrent_unordered_map or whatever the equivalent is for your language
- clutchski 12y agoBecause it will persist between restarts?
- hendzen 12y agothen it seems strange to call it a cache.
- yummyfajitas 12y agoIt could be a cache to a much slower backend. I pull a lot of stock price and other data to my server which is is slow - 200-1000ms per series. I cache it in postgres which allows me to load it nearly instantaneously. It also allows me access to data while I'm disconnected. Another reason to cache to disk is that you want to store more data than you have ram.
- otterley 12y agoBecause it's shared between processes on the same server.
- nly 12y agoIn theory STL implementations, if used with a custom allocator, should be able to pull this off... that's why the STL containers all have internal 'pointer' typedefs. Practically speaking, Boost.Interprocess includes a shared memory hash table implementation. Boost Multi Index, which is a further generalisation of containers to allow the construction of database-like indexes, is also Interprocess compatible. http://www.boost.org/doc/libs/1_57_0/doc/html/interprocess/allocators_containers.html#interprocess.allocators_containers.additional_containers.multi_index http://www.boost.org/doc/libs/1_57_0/doc/html/interprocess/a...
- otterley 12y agoI'm so excited that they finally open-sourced this. It's relatively old tech at Yahoo, stuff folks outside never got to see. It was difficult to explain to later colleagues the stuff I knew about shared-memory databases because I couldn't give them a frame of reference. mdbm performance is even better on FreeBSD than Linux because FreeBSD supports MAP_NOSYNC, which causes the kernel not to flush dirty pages to disk until the region is unmapped. Perhaps mdbm's release will finally get the Linux kernel team to provide support for that flag.
- jzawodn 12y agoSame here. I remember wishing we could Open Source it back in the early 2000s. Good to see this coming out so people can take a little credit for work they did back in the day.
- deleted 12y ago[deleted]
- deleted 12y ago[deleted]
- extralam 12y agoyahoo back to IT company ?
- EGreg 12y agoHow is this different than memcache?
- deleted 12y ago[deleted]
- pjscott 12y agoMemcache is an in-memory cache. This is an on-disk key-value store.
- polskibus 12y agoCan anyone say whether it would be hard to port it to Windows? Maybe there already is something for Windows that is as good as this ?
- luckydude 12y agoI can't speak to the yahoo version, they've wacked it, but the base mdbm that we still use today works fine on windows, has for years.
- chatman 12y agoLet the horrors of MDBM not get to you. I've used it when I worked at Yahoo, and the client support for Java etc. sucks.
- jwr 12y agoThis is a very big deal, especially because of the BSD licensing.
- qwerta 12y agoI dont want to brag. But there is also DBM inspired Java port. And in-memory mode outperforms java heap collections such as j.u.HashMap.
- remon 12y agoI'm not very comfortable with storage engines that directly build on memory mapped files. MongoDB's current storage engine is mmap based and it's sub optimal at best which is undoubtedly part of the reason they're building a completely new storage engine now (WiredTiger).
- hyc_symas 12y agoUsing mmap well takes great care. MongoDB was careless. There's good reason to believe the MDBM designers were careless too.
- deleted 12y ago[deleted]
- luckydude 12y agoMDBM guy here. Care to elaborate on what we got wrong?
- hyc_symas 12y agoSee below. https://news.ycombinator.com/item?id=8734356 https://news.ycombinator.com/item?id=8734356
- cbsmith 12y agoUmm... the lack of transactional integrity is part of the mdbm design. So it's only "careless" in the strictest sense of the term (the designers explicitly wanted to exploit not having to care about it). mdbm is certainly not without limitations, but is careful about its use of mmap to an extent that comparisons with MongoDB are laughable.
- hyc_symas 12y agoI have already admitted my obvious biases, but seriously - when you design such a trivially corruptible system, you can't call it a persistent database; it's at most a cache. It won't survive a system crash at all, it probably won't survive an application crash intact either. To call it persistent is laughable.
- PhuFighter 12y agoI'm curious to see what the total timings would be like to get the data in a useable form - as opposed to just fetching a record from a data store. As noted - these data stores just store and retrieve data and don't do things like joins or ordering, etc. Could there be a comparison between these datastores and the traditional ACID compliant databases when it comes to retrieving actual data in a useful format? E.g. perhaps doing a join or an ordering of some sort? I don't expect databases (e.g. Oracle, MS SQL Server, DB2) to be faster in raw performance, but I do expect them to be faster in terms of total development time and bug fixing since the application developer wouldn't have to do the locking, page pinning/unpinning, etc. manually.
- swah 12y agoWhere does it say that this database is persistent?
- t1m 12y agoIt is memory mapped, which means that it is persisted to disk, perhaps confusingly if you aren't familiar with mmap.
- mbrzusto 12y agoHow similar in performance is MDBM to GDBM (the GNU DBM)? They appear to be similar (if not identical) in functionality.
- api 12y agoNot sure, but in my experience GDBM is a bit on the slow side. MDBM uses mmap(), so for that reason alone it should be faster.
- luckydude 12y agoMDBM was designed to be fast with special care taken on the lookup path. The goal was to do lookups with as few cache misses as possible. You can get to any key with at most two page faults.
- cbsmith 12y agoThey might appear similar, but that's just because they share the same DBM interface heritage.
- discardorama 12y agoHow is MDBM for concurrent access? How does it handle locking (i.e., one big lock that blocks everyone else, or key-level locking)?
- luckydude 12y agoSo the SGI owned code, that I don't have, did page level locking. There two kinds of locks, rd/wr on the directory, and rd/wr on a page. If you are inserting a key you get a read lock on the directory and a write lock on the page. If it fits in the page then you are done. So you can have lots of concurrent writers until a page is full and you have to split it. Bob Mende did that work I think, you might track him down for details.
- i_am_ralpht 12y agoWhere is the original open source release from Silicon Graphics which Yahoo based this work on? Did they ever make one?
- luckydude 12y agoNah, they didn't care and I didn't want to piss them off so I just handed the code to anyone who asked for it.
- swah 12y agoCould not install this in Ubuntu 12.04 - basic commands are failing. I think they tested only in BSD? ln -s -f -r /tmp/install/lib64/libmdbm.so.4 /tmp/install/lib64/libmdbm.so ln: invalid option -- 'r' Try `ln --help' for more informatio