3 ms·
A caching layer for random reads where the dataset is 10x larger tham memory isn't hugely useful. If you get a hit, great, but you can't count on it. For memor
by prospero 12y ago
A caching layer for random reads where the dataset is 10x larger tham memory isn't hugely useful. If you get a hit, great, but you can't count on it.
For memory mapping to make sense you need to fetch a big chunk of data, whereas read() gets a page's worth. For the data size and read pattern described in the post, the latter is much more desirable.
- fiatmoney 12y agoIt's the opposite. As I said, read() and a read into a mmap'do array both hit the OS page cache first, and will bring in ~4K of data on miss; read() also has the overhead of a system call. For tiny read/writes the advice is to use mmap. This is different if you're doing "direct io" and bypassing the OS page cache because you have your own caching layer, but I don't think they do.
- prospero 12y agoWhose advice? Check out RocksDB's front page: http://rocksdb.org http://rocksdb.org. Empirically what you're saying isn't true in my experience, rather mmap should be used when there's decent coherency w.r.t. the available memory. Without knowing what you're basing your belief on, I really can't address it.
- fiatmoney 12y agoIt looks like they are making some claims about OS level bottlenecks specifically with the virtual memory subsystem. This is something I'd like to look into; all I can find is the particular quote but no explanation of where they think the bottleneck actually lies. The experience of, e.g., the SQLite folks seems to be different. https://www.sqlite.org/mmap.html https://www.sqlite.org/mmap.html