6 ms·
Very nice! If you wanted it to save to disk instead of living wholly in memory, how would you do that in C?
by flunhat 9y ago
Very nice! If you wanted it to save to disk instead of living wholly in memory, how would you do that in C?
- kevingadd 9y agommap! I mean, not really, but it's a surprisingly viable starting point for simple problems. I've shipped it. Some production-grade software still uses mmap, albeit with a bunch of additional complexity to make sure it works okay in weird cases. See this old comment thread I dug up that might be relevant: https://news.ycombinator.com/item?id=3982514 https://news.ycombinator.com/item?id=3982514 It seems like Redis might currently make use of mmap for some of its data, but I couldn't find an up-to-date source for that, just some old blog posts by the developer.
- pjc50 9y agoThe trouble is that naive mmap doesn't give you the control over write ordering that you need to implement reliability. Sadly, about 50% of database design is trying to ensure the thing has a decent chance of starting up again if it crashes or loses power. This mostly consists of fighting file systems and disk caches.
- lsllc 9y agoThere's `msync`
- valarauca1 9y agoThis is why msync is for. You can sync page by page and ensure ordering yourself.
- drfuchs 9y agoYou think you can, but you can't! Linux always sync's the whole file; see https://lwn.net/Articles/502612/ https://lwn.net/Articles/502612/ for a discussion of why fixing this horrible bug is too dangerous.
- icedchai 9y agoI used to work on a system that used mmap for basically everything, including a proprietary database that processed financial transactions. The original production system ran on a commercial Unix. (It's been about 15+ years, but I think it was AIX.) At one point, we ported it to a different Unix platform which had slightly different mmap semantics. It looked like everything worked, except for one minor detail: the data was never actually synced to disk. Ever. Until shutdown. Unfortunately, since the system ran 24/7, it effectively never synced. First time the system had a power failure, there was massive data loss on startup. Oops. (We were able to recover through log replay...)
- tomcam 9y ago> (We were able to recover through log replay...) I call that a win! What was the logging mechanism--something bespoke? What was your solution?
- imtringued 9y agoThat sounds like a terrible idea even if it were 100% reliable. When you shut down it will have to flush the entire memory at once. If you're syncing to a hdd it can mean several minutes of downtime.
- pjc50 9y agoA correctly working system will "flush behind", so there should be no difference between a clean shutdown and a power loss. If there's several minutes of unsynced data that's data that will be lost on power off.
- zero_iq 9y agoJust going to chip in here to mention LMDB, which essentially provides an mmap btree database with single writer + multi-reader snapshot concurrency, ACID transactions, and is extremely robust. If you think something may benefit from a shared memory data store, lmdb may be worth considering as a fast, reliable, high concurrency alternative to reinventing wheels with raw mmap + manual sync.
- switchbak 9y agoNot the OP here, but why the downvotes? Is this person off-base, or is that a poor solution?
- ovao 9y agoThe poster's comment reads a bit like an advertisement, but LMDB would be a good solution for persisting data, in my opinion. It has a pretty simple C API and it's been described by its authors as "crash-proof".
- pletnes 9y agoWhy not really? From the sqlite docs: Beginning with version 3.7.17 (2013-05-20), SQLite has the option of accessing disk content directly using memory-mapped I/O and the new xFetch() and xUnfetch() methods on sqlite3_io_methods.
- dkersten 9y agoI would say its less that you can’t or shouldn’t use mmap, and more that its not as simple as throwing mmap at it - if you care about crash recovery/not corrupting data (which most databases should care about), but that certainly doesn’t mean that mmap can’t be a part of the solution.
- convolvatron 9y agoi think that topic is larger than the presentation thus far. hopefully memory mapped flash will obviate the need for alot of this but: o you need to make an on-disk structure (i.e. btree) to let you search your tables (indices) o that has to be fronted by an in memory cache o your on-disk structure should maintain consistency even if the power fails and some writes get lost - often but not always this is facilitated by keeping a separate write-optimized structure called a write-ahead-log - if you use a WAL, you'll also need to implement the replay mechanisms to get the primary indices back up to date on a failure o because the disk operations are expensive and high latency you'll need to start managing concurrency explicitly back ends are alot of work. unfortunately this is a place where the lack of decent concurrency mechanisms in your language and OS interface can really cause alot of headache. pedagogically, i guess you would start with a completely synchronous non fault tolerant btree? or a maybe just the log, and introduce trees as a read accelerator?
- noam87 9y agoWhat are some good readings / places to start on this topic?
- convolvatron 9y agothats a really good question, unfortunately everything I learned was by working with people who knew more than I did. its pretty sad that alot of systems work is that way. this book looks to be pretty comprehensive, and explicitly discusses on-disk structures, write ahead logging, and recovery. http://infolab.stanford.edu/~ullman/dscb.html relational languages and transactions get alot more playin the academy..i think its because you can make sense of them in some abstract way without getting sucked into involved discussions about caching heuristics and OS write guarentees, scheduling and fussy performance characteristics
- jwhitlark 9y agoI like http://dataintensive.net/ http://dataintensive.net/, though it's a slightly different focus.
- flavio81 9y ago> Very nice! If you wanted it to save to disk instead of living wholly in memory, how would you do that in C? The idea would be to map your memory space to disk space (using random access as provided by C stdlib); the problem would then be: 1. Caching - i guess you should then provide your cache implementation. And i would guess this opens a can of sync problems... 2. Finding a way to efficiently write the data on disk (i.e. which data should stay contiguous?; how much "slack" space should i leave on "extents" (oracle slang for a container for many data blocks)? Should you do column-store? 3. Compacting the data (removing slack) etc etc. I think it's not easy at all !
- ddorian43 9y agoLmdb ?
- 72deluxe 9y agoC++ solution (the solution you didn't ask for, sorry) for saving to disk from memory (as some sort of periodic expensive sync) would be to write a stream operator for your objects and stream them to disk. Loading would involve reading them back; would only work on the same OS and architecture. Alternative approach: For Windows, you could just use disk but pass FILE_ATTRIBUTE_TEMPORARY to CreateFile to force it to be in memory if the file is small enough.