12 ms·
It depends on what is meant by in-memory database. The most useful kind IMO is the one which actually saves everything to disk, but is not designed like a tradi
by slaymaker1907 2y ago
It depends on what is meant by in-memory database. The most useful kind IMO is the one which actually saves everything to disk, but is not designed like a traditional RDMS in that it assumes everything, including indices, can be saved in memory. Therefore, you don't need a complicated buffer pool system and you don't need to touch disk at all after startup to service read queries. The most simple approach to such a database is just to MMAP a file.
This kind of workload is probably the most common in all of software development for the past couple of decades given how plentiful RAM is as well as most applications having some need for storing persistent data.
- hinkley 2y agoI always feel a little weird using memcached because it has never once crashed on us but when it goes down we have a bad time with circuit breakers. We only have problems with memcached when we create them ourselves. Disk backing store would soften that considerably.
- lanstin 2y agomemcached type API with RocksDB backing store is pretty good. Honestly, at this point hasn't every one written some in memory DB with various methods to persist and various consistency models as a result? At this point the magic is in client routing to the appropriate shard without having to redeploy to change the shard configuration and to allow multi-remote callers to access the data and still get access without disk load or mutex/locking around the data. I have a thing where the shards each have sub-shards in process and 1 go-routine per sub shard; communication to/from the remote callers is via channels to a per-request go-routine (or whatever it is gRPC does) and the main subs hard go-routine has no locking on itself. Just a big hash map and a DLL to implement an LRU so I have a hard cap on memory usage, and no allocations for lookups or mutations (just creations).
- anotherguy0 2y agoSince you mentioned MMAP: "Are You Sure You Want to Use MMAP in Your Database Management System?" https://www.cidrdb.org/cidr2022/papers/p13-crotty.pdf https://www.cidrdb.org/cidr2022/papers/p13-crotty.pdf
- riku_iki 2y agoDataset was 20x times larger than available RAM in that case, so it makes sense that OS cashing was useless and only induced overhead.. Another potential issue was that they compared their mmap code to fio O_DIRECT code, which kinda not clean experiment, fio could be just much more optimized itself..
- slaymaker1907 2y agoBesides this point, I mostly mentioned it as an example of probably the simplest in-memory yet disk-backed database imaginable. For example, if you want to implement any sort of transactions, you don't want to actually do anything like fsync until you start trying to commit the data.
- arandomusername 2y agoIt's also worth while reading the rebuttal from ravendb: https://ayende.com/blog/196161-C/re-are-you-sure-you-want-to-use-mmap-in-your-database-management-system https://ayende.com/blog/196161-C/re-are-you-sure-you-want-to... Also worth mentioning LMDB, probably the fastest embedded K/V db in terms of read perf uses mmap.