4 ms·
Not 100% what the article is about, just a short story. In one of my old jobs we had megabytes of infrequently accessed static key-value data. If we simply loa
by hakunin 3y ago
Not 100% what the article is about, just a short story.
In one of my old jobs we had megabytes of infrequently accessed static key-value data. If we simply loaded it into a const (i.e. a hash table), it would blow up the RAM so much, that we would need to upgrade to bigger VPSes. If we put it in database, it would make it annoying to keep these tables up to date, track changes in them.
I figured this was one of those in-between use cases, where the best solution is to have zero-RAM lookups from SSD. In my case, I wrote a little ruby library[1] that arranges data in equal cells in a file, and performs binary searches via `pread`. This was perfect for us, because we kept data in our repos, sacrificed no RAM at runtime, SSD lookups were fast enough, and we didn't have to support a more elaborate db.
[1]: https://github.com/maxim/wordmap https://github.com/maxim/wordmap
- CraigJPerry 3y agoI think i’m missing something here, could you not just mmap() the file in and let the kernel take care of memory pressure for you?
- sschnei8 3y agohttps://db.cs.cmu.edu/mmap-cidr2022/ https://db.cs.cmu.edu/mmap-cidr2022/
- hakunin 3y agoI'm no systems programmer, but I remember trying to research mmap approach. It's been 3 years, so I'm not sure what stopped me, but something didn't feel right. Perhaps it was the lack of control over how memory is used. I could clearly see how not use it, and didn't want any fluctuations. Edit: oh and I think I did come across some article like the one linked in the neighbor comment. It's starting to come back.
- lisper 3y agoSure -- until your system crashes and data is left in an inconsistent state because some of it was written to non-volatile storage and some of it wasn't and your OS had no concept of transactional consistency because mmap is a leaky abstraction. [UPDATE] I missed a crucial part of the problem setup, which is that the data is read-only. That obviously moots my objection.
- legulere 3y agoAnd Flatbuffers for data-layout and access.
- boywitharupee 3y agoBasically, your solution ended up residing on-disk, but the data was fragmented in such a way that efficient lookup was possible? Did this copy from kernel space to userspace, also?
- hakunin 3y ago>the data was fragmented in such a way that efficient lookup was possible Yeah, there's a build step that sorts and arranges data into "cells", making binary search possible. Not sure about kernel/user space. I'm just calling `pread` from ruby, so only a few bytes are loaded per lookup.
- rubiquity 3y ago> sacrificed no RAM at runtime Those lookups eventually made their way into the kernel's page cache.
- hakunin 3y agoAre you saying that the file would just entirely be loaded into a page cache eventually? I imagine, even if true, it still wouldn't result in OOM killing the server/worker daemons on account of this data?
- rubiquity 3y agoDepending on the file size and available memory, yes. In the event memory needs to be reclaimed Linux’s memory management system will free from the page cache first if processes need more anonymous memory.
- deleted 3y ago[deleted]
- coldtea 3y agoThe "log(n) * number of searches" page lookups, that the kernel could clean up at any time after a search, instead of all n items that it would have had to load and keep in memory? Yes, they did.
- teaearlgraycold 3y agoWhy not commit a sqlite db?
- hakunin 3y agoI considered that, but couldn't find a way to precisely control amount of RAM used when you read from it.
- jerrygenser 3y agoI have a similar use case and I have had success reading from sqlite files embedded in my repo