3 ms·
How did that compare to --mmap?
by electrum 16y ago
How did that compare to --mmap?
- angusgr 16y agoWow, I didn't realise grep had a separate --mmap option. Interesting. I don't think it would make much difference in the parent's case of many, small, fragmented files because if you're mmapping each file in turn and it's not cached, it still needs to be loaded from disk - it just happens in a page fault instead of the read() call. Possibly if you mmapped all of the files and then used madvise() or something to prefetch in front of where you are in the list of files. Maybe grep does that, I don't know? I guess the case where that technique would help is actually when you have a combination of (a) many files, (b) a computationally expensive pattern match (even just -i is a measurable hit) and (c) largeish files. Because on many small files and a simple match, the disk I/O is still going to be the major component - even if you prefetch you still can't get around needing to load all the file contents from disk.
- DarkShikari 16y agoPossibly if you mmapped all of the files and then used madvise() or something to prefetch in front of where you are. Maybe grep does that, I don't know? It doesn't do that, and that's the first method I used in my hack. fadvise() works just as well though, I think.
- angusgr 16y agoIt doesn't do that, and that's the first method I used in my hack. Fair enough. :) If you feel like sharing then I'm very curious as to what the additional methods were.
- DarkShikari 16y agoIt was a total hack, so I just opened files ahead of time, fadvised them, then closed them.
- dasht 16y ago"mmap" is not a general solution for grep. It's nice when it can be used but grep must be able to grep files much to large to reasonably mmap and various kinds of streams.
- nikcub 16y agoYes but you would think that a reasonable default solution would be to use both. ie. mmap by default and then if the file is larger than MAX (based on memory use, file size) then use read.