4 ms·
What does mmap do exactly? Why was the transition to using it a big improvement in llama.cpp?
by eachro 4y ago
What does mmap do exactly? Why was the transition to using it a big improvement in llama.cpp?
- programmarchy 4y agoMy understanding is that it maps a file directly to memory to reduce disk usage.
- debatem1 4y agoDoesn't change disk usage. The file is still on disk. The difference is between reading a file and memory mapping (mmap'ing) it. If you read a 1TB file into memory you use 1TB of disk and 1TB of physical memory. If you then access that data it's as fast as RAM because that's where it is. If you mmap a 1TB file you use 1TB of disk and 1TB of virtual memory. If you then access that data it may be mapped into virtual memory but not actually be in RAM. This triggers a page fault, at which point the correct page is loaded from disk to physical memory, and handed back to you. The key observation is that the amount of physical memory occupied by the mmap'd file is much smaller than the entirety of the file unless you access almost all of it. If your filesystem supports holes, this can also be useful for writes: it's possible to map files vastly larger than physical disk space, but so long as the actual number of places written to is quite small you won't run out. The combination is very useful for datastructures because you basically don't have to care about data being extremely sparse until that data is also getting quite large, which means you can use cheap/fast approaches to indexing, etc.
- Salgat 4y agoFor example, you create a 100MB file. You tell the operating system to map that file to memory, and it gives you a pointer to 100MB of memory. Whatever you read from that 100MB of memory is what is actually in the file. You can also write to that memory and commit it back to the file. The parts you read from the pointer are the only parts from the file that are loaded into memory. So if you memory map a 100GB file, the operating system won't actually load all 100GB into memory, only what is accessed (and this is all handled for you automatically). The operating system is free to load and cache the file into memory in whatever way it wants, so for large memory mapped files it'll often try to use all available memory to cache as much as possible. If another program needs memory, the operating system will simply lower the amount of memory available to the memory mapped file for caching. This is extremely useful for databases, since it greatly simplifies both how to persist the data along with how to load and cache the persisted data. This all comes with a big caveat however. The less memory you have, the more file accesses occur (similar to your pagefile when you're thrashing), which can dramatically slow down your memory operations. tldr; it lets you designate a file to use as a region of memory.
- Jasper_ 4y agoThe kernel page cache already exists, and is used on read() too. The distinction is that instead of read() copying from the kernel page cache, mmap will map that page cache page directly into the user's memory address space. It saves a copy on read, in exchange for a bit more overhead configuring the processes address space up front.
- chasd00 4y agoso basically an MRU cache for a file's contents? Is there enough information to know what bytes from the file are cache'd and which are not? It would interesting to see like a heatmap of a large model showing what portions of the neural network are being used the most. ..like a brain MRI more or less.
- AceJohnny2 4y agoQuoting another user, jcranmer [1]: > "The fundamental operation of mmap is to add new entries to the page table of a process, and the precise properties of those entries are heavily dependent on what the arguments to mmap are. > When you mmap a regular file, you're essentially adding an entry to the page table that shares the data with the kernel's filesystem cache." Now, understanding this requires some understanding of an OS's virtual memory function, and what "page tables" are. Those are what the OS use to track a process' memory, whose granularity is in "pages" (historically 4kB on Linux, though others use larger granularity such as 16kB). mmap() has a a lot of flags [2] that affect the properties of those mapped pages. It is the swiss-army knife of memory management on Linux. Some of those properties allow you to share memory with other processes (MAP_SHARED | MAP_ANONYMOUS), or just allocate memory (MAP_ANONYMOUS) or, by default, map an (open) file specified by the `fd` argument. (fun fact! On linux, when you malloc(), you don't actually get memory, just an IOU from the kernel. Only when you access that memory, ie accessing those memory pages, does the kernel actually make the effort of allocating you that memory.) [1] https://news.ycombinator.com/item?id=35412842 https://news.ycombinator.com/item?id=35412842 [2] https://linux.die.net/man/2/mmap https://linux.die.net/man/2/mmap
- xiphias2 4y agoOne thing I don't understand is that if I read a gigabyte from a file with the read call, why the kernel can't see that it's unused allocated memory, create copy on write pages, which could just be shared with the disk cache as long as the data is aligned (which it should be after a malloc + read call). It seems the intuitive way to implement read if there's a complex memory subsystem that does all these great things anyways.
- pkaye 4y agommap maps files into virtual memory. When a program read the mapped are of memory, that portion of the file is read.
- akiselev 4y agoThe transition to mmap offloaded memory management to the kernel, which can lazily load in parts of the file from disk as its memory mapped pages are accessed. The original version eagerly read the files into memory before running inference.
- deckard1 4y agothis is probably the best source to actually understand how memory works in Linux: https://manybutfinite.com/post/anatomy-of-a-program-in-memory/ https://manybutfinite.com/post/anatomy-of-a-program-in-memor... [1] You can't really understand mmap/sbrk without understanding virtual memory and process space layout. [1] Images are broken. Just open https://static.duartes.org/img/blogPosts/kernelUserMemorySplit.png https://static.duartes.org/img/blogPosts/kernelUserMemorySpl... and go to advanced and "Proceed to static.duartes.org" to workaround their https issues. The duartes.org host is owned by the blog author. Refresh the blog article and images should load now.