11 ms·
Why MMAP in llama.cpp hides true memory usage
- flurly 4y agoTL;DR > The way mmap handles memory makes it a bit tricky for the OS to report on a per-process level how that memory is being used. So it generally will just show it as "cache" use. That's why (AFAICT) there was some initial misunderstanding regarding the memory savings
- vadansky 4y agoI lost track since things move so quickly. Was there still memory savings just not as drastic? Or no memory savings, just a speed-up?
- chpatrick 4y agoIt's neither memory savings or a speed-up really. The advantage of mmap is that you can treat a file on a disk a block of memory, so pages from it can be loaded (or unloaded) as necessary instead of one big upfront load into RAM. The benefit is that you can work with data that's bigger than your physical RAM because the kernel can swap it back out to disk if needed. Another benefit could be that if only a small part of the data is needed to compute something then the OS will automatically only load those, but it's unclear to me whether this is the case with LLaMA.
- toxik 4y agoRegarding your last point: No, you need all of the weights all of the time. Edit: except embedding weights but those are not the problem.
- iforgotpassword 4y agoIt's still somewhat faster if you benchmark it. I assume the os is doing good enough prefetching in the mmap case to hide the loads from disk mostly. So it's not just hiding the initial load of 30gb from disk. Obviously if you're swapping because you don't have enough memory to hold the model in RAM, the mmap version is going to be much faster, since you don't need to swap anything out to disk but just discard the page and re-read from disk if you need it again later.
- antonvs 4y ago> So it's not just hiding the initial load of 30gb from disk. The issue is typically that that initial load involves some sort of transformation - parsing, instantiating structures, etc. If you can arrange it so that the data is stored in the format you actually need it in memory, then you can skip that entire transformation phase. I don’t know if that’s what’s been done with llama.cop though.
- simion314 4y agoIt is a speedup for me. When I run llama.cpp from CLI first tiem it takes a very long tiem to load the model in memory. If the program exits or I stop it with Ctrl+C and start it again it will start almost instant.
- chaboud 4y agoThat's down to caching. If you used your system to do something else for a while, you'd find those pages evicted and the performance back down to Earth. That's one of the things that makes mmap so useful, though. The system can take advantage of access patterns to dramatically improve performance.
- simion314 4y agoYes, makes sense. And is great. Though honestly not sure why it takes minutes to load a 23Gb model in RAM, I feel is not proportional with the smaller models.
- CyberDildonics 4y agoThey weren't asking about mmap, they were asking about the program itself.
- deleted 4y ago[deleted]
- rovr138 4y agoIt's basically paging to disk. Not necessarily memory savings, but the improvements here are that it will run on computers with less ram because it can page to disk. Not necessarily the number reported (since you do need to load chunks into ram), but still lower.
- chaboud 4y agoSort of, but without the duplication and initial wait to load. A traditional fat in-memory app would do this: file (DISK) to process active pages (RAM) to paged out virtual memory if saturated (elsewhere on DISK) Using mmap typically goes something like: file (DISK) to process active pages (RAM) to released and cached (RAM) to uncached if evicted (back to same place on DISK) OR back to process active pages from cache (RAM) For the cost of fixing up some process page tables, the physical memory pages necessary can be brought back from the cache rather than read from disk. It's an orders-of-magnitude performance savings.
- detrites 4y agoIf that's the case then part of this may be the different interpretations of "memory". One persons "paging the same memory requirement from disk to RAM" can be someone elses "requiring less memory/RAM".
- hnav 4y agomore like memory mis-reporting, since when you mmap a file in, IIRC that counts against page cache rather than memory usage (you can evict the page without causing write IO so the memory isn't "used")
- astrange 4y agoIt being safe to evict is why it's correct to not report it as "memory usage". It's part of the program's working set but measuring that is a completely different story.
- eigenvalue 4y agoIt did somewhat reduce the total memory used. Now you can load the 30B model while only using ~20gb of RAM, which is about the aggregate size of the 4bit quantized weight files for that model. The real win is that you can kill the main inference binary and try another prompt, and it will start doing inference basically immediately instead of spending 10-15 seconds loading up all the weights into RAM each time.
- thatcherc 4y agoGood thread! https://threadreaderapp.com/thread/1642726595436883969.html https://threadreaderapp.com/thread/1642726595436883969.html
- ImprobableTruth 4y ago>mmap is a really nifty feature of modern operating systems ... for some definition of "modern".
- antonvs 4y agoDoes this mean DEC’s TOPS-20 from 1969 is a modern OS?
- gumby 4y agoMmap was the only way to do disk I/O in Multics — we’re talking about a design a decade earlier than that. (Also you’re thinking of TOPS-10 — TOPS-20 and TWENEX were developed in the 70s. But your heart is in the right place!
- koito17 4y agoI wonder how the author would label Windows' MapViewOfFile :)
- deleted 4y ago[deleted]
- lionkor 4y agoyou dont get the devs who spend their day on twitter instead of writing code to read your thread if it doesnt mention "modern" at least once. /s Its quite obvious that the author "dumbs down" all the information for a very non-developer audience, so I assume that gives some leniency. If you tell people that we have had things like mmap() for decades, they may start asking too many questions about why every piece of software underperforms so horribly below what was possible decades ago. Bit of a rant, but I feel that we lose a little bit of potential every time a developer calls operator<< on a std::istream in a loop.
- chaxor 4y agoDevs of 2023: "why does this wheel have rubber on it? It needs to be more modern. Let's take away the rubber and add firecrackers in its place as a nice feature, so really pops and makes the experience more exciting"
- CyberDildonics 4y agoThis isn't why mmap in llama.cpp hides true memory usage, it is why mmap hides true memory usage.
- Thaxll 4y agoPeople are discovering mmap, it has been used for a very long time especially in databases.
- xiphias2 4y agoIt's true for databases, but this usage is closer to how shared libraries are loaded, just with extenal data. Also just to show how buggy mmap is, it's disabled in SQLite by default for example: https://www.sqlite.org/mmap.html https://www.sqlite.org/mmap.html Even now the windows version is buggy, but the great thing is that after it's fixed in llama.cpp, the open source community can just copy the solution. I don't know of many interactive utilities that were known for fast startup time _because_ of the use of mmap, but now we have one, which means many more are coming :)
- quotemstr 4y agoBuggy? In what way?
- masklinn 4y agoThe "But there are also disadvantages" of the linked pages provides some of the issues with mmap. This also matches burntsushi's experience with mmap in ripgrep: - depending on concurrent accesses mmap can just sigbus on you (e.g. if the mapped file is being truncated) - mmap simply does not work with virtual filesystems, and will blow up on large files on 32b systems (windows also has further limitations on mmaps) - depending on workload mmap may not be faster than regular reads and memory buffers, ripgrep will mmap when working on just a few files, but will use normal buffers for large file counts, because when you start reusing buffer you amortise allocation costs which you can't amortise when creating and destroying mappings - not only that but mmap/munmap are also globally blocking on the process, so in multithreaded processes it's very bad to map/unmap a lot, you can stall your own application - the semantics of mmap on crash are also somewhat risky, in that the OS will try very hard to sync, but that may not be desirable if the application crashes in the middle of a write
- eachro 4y agoWhat does mmap do exactly? Why was the transition to using it a big improvement in llama.cpp?
- programmarchy 4y agoMy understanding is that it maps a file directly to memory to reduce disk usage.
- debatem1 4y agoDoesn't change disk usage. The file is still on disk. The difference is between reading a file and memory mapping (mmap'ing) it. If you read a 1TB file into memory you use 1TB of disk and 1TB of physical memory. If you then access that data it's as fast as RAM because that's where it is. If you mmap a 1TB file you use 1TB of disk and 1TB of virtual memory. If you then access that data it may be mapped into virtual memory but not actually be in RAM. This triggers a page fault, at which point the correct page is loaded from disk to physical memory, and handed back to you. The key observation is that the amount of physical memory occupied by the mmap'd file is much smaller than the entirety of the file unless you access almost all of it. If your filesystem supports holes, this can also be useful for writes: it's possible to map files vastly larger than physical disk space, but so long as the actual number of places written to is quite small you won't run out. The combination is very useful for datastructures because you basically don't have to care about data being extremely sparse until that data is also getting quite large, which means you can use cheap/fast approaches to indexing, etc.
- Salgat 4y agoFor example, you create a 100MB file. You tell the operating system to map that file to memory, and it gives you a pointer to 100MB of memory. Whatever you read from that 100MB of memory is what is actually in the file. You can also write to that memory and commit it back to the file. The parts you read from the pointer are the only parts from the file that are loaded into memory. So if you memory map a 100GB file, the operating system won't actually load all 100GB into memory, only what is accessed (and this is all handled for you automatically). The operating system is free to load and cache the file into memory in whatever way it wants, so for large memory mapped files it'll often try to use all available memory to cache as much as possible. If another program needs memory, the operating system will simply lower the amount of memory available to the memory mapped file for caching. This is extremely useful for databases, since it greatly simplifies both how to persist the data along with how to load and cache the persisted data. This all comes with a big caveat however. The less memory you have, the more file accesses occur (similar to your pagefile when you're thrashing), which can dramatically slow down your memory operations. tldr; it lets you designate a file to use as a region of memory.
- Animats 4y agoWell, yeah. Short version: the claim that llama could use much less memory by memory-mapping the model data file was wrong. It's just that the Linux utilities for memory use don't count memory-mapped file area.
- pavon 4y agoI'm not convinced that there isn't more to the story than that. People (including Justine who implemented MMAPing the model and knows how to monitor memory use properly) are seeing that much of the file isn't being paged into memory, and are not quite sure why[1]. Is it a bug? Is the distribution of weights used highly dependent on the prompts? Are some weights somehow unused altogether? This twitter thread rules out some of those, but I don't think it closes the door. [1] https://news.ycombinator.com/item?id=35393615 https://news.ycombinator.com/item?id=35393615 edit: softened some of my claims after reading more updates from over the weekend.
- Animats 4y agoNow that's interesting. Entire memory pages of the model aren't being referenced?
- freilanzer 4y agoIf correct could this hint at sub models for specific tasks?
- sitkack 4y agoeBPF echo 3 > /proc/sys/vm/drop_caches
- wmf 4y agoOther people are saying that the whole file does get paged in and transformers access all their weights by design.
- peterfirefly 4y agoThe kernel will read more than just the page that faulted. Jart counts just the page faults and multiplies by the page size.
- mrbonner 4y agommap is modern? Lol. I used mmap in Java back in 2013 to offheap large matrices in our regression engine. Mmap probably exists long before that. Edit: and yes, the JVM only reported a few hundreds MBs used in the heap. In reality, the memory mapped file is several GBs in size.
- loeg 4y agoThe syscall was described in 4.2-4.3BSD in the 80s and shipped in 4.3BSD-Reno in 1990. The underlying concept of memory-mapped files dates back to at least the 70s. So all of this to say -- I agree, it's not especially new. https://en.wikipedia.org/wiki/Mmap https://en.wikipedia.org/wiki/Mmap
- peterfirefly 4y agohttps://en.wikipedia.org/wiki/Multics#Novel_ideas https://en.wikipedia.org/wiki/Multics#Novel_ideas
- Kiro 4y ago2013 is modern.
- queuebert 4y agoI have kids older than your usage of mmap.
- antonvs 4y ago> modern? Lol. 2013 How to tell if a dev is in their 20s or thereabouts
- btown 4y agoI mean, if we're being pedantic, I think it is possible to run llama.cpp with real memory savings this way... in a contrived situation where you were only generating a single token with a single forward pass, and you were limited in RAM. The new MMAP system would not require you to load the whole model from disk up front, but rather would allow you to load and immediately forget each layer of parameters as you do a forward pass through the GPT architecture. And if you had just tried to use normal swap paging to do this without MMAP, you'd essentially be doing 2x the loading work, since the page for the first layer likely got swapped out as you loaded the rest of the model into "memory." Of course, as soon as you want to generate a new token, you'd need to reload all the pages of the model parameters. Every token generated would take as long as it would if you were loading the model from scratch. But for a classification problem where you only need to generate one token, and you need or want to do this in a serverless environment where you'd need to load from disk anyways? I think you'd be able to do that workflow with significantly less RAM in the same amount of clock time, now.
- visarga 4y ago> Every token generated would take as long as it would if you were loading the model from scratch. Wouldn't make sense to keep most of the model in GPU (as much as it fits) and only load the remaining layers on each pass?
- alchemist1e9 4y agoIs it possible to do multi-GPU inference the way you describe and have say 4 GPUs but each doing different layers on a large model? If that is the case then perhaps a high end consumer system with max pcie 5 nvme bandwidth and multi gpus could do inference on large LLM models.
- rnnr 4y agoMemory mapped files also known as sections in VMS/NT have just two advantages: * Fewer context switches among user space / kernel syscalls * No need to copy _modified_ data into the swap. The behavior for read only data doesn't change That's it, nothing miraculous about it.
- astrange 4y agoNo, it also shares them between multiple runs of the same process, and it reuses the pages that were in your file cache anyway. > * Fewer context switches among user space / kernel syscalls This is not necessarily true. There's lots of cases where it's actually slower.
- jlokier 4y ago> Fewer context switches among user space / kernel syscalls Note that each page fault to read an mmaped page is also a context switch from user space into the kernel. The kernel entry/exit paths for syscalls and page faults have much in common. Mmaped pages can in principle use fault-ahead to reduce the number of page faults when sequential access is detected. An equivalent reduction is available for read syscalls by reading larger blocks at a time. In practice, mmap files are faster in some scenarios compared with read syscalls and slower in others. Same with O_DIRECT reads, those are faster in some scenarios and slower in others.
- riedel 4y agoBest part: >AI-generated summary: >"This thread explains how the mmap feature of modern operating systems can be used to reduce memory usage when running large deep learning models. It also explains how... Yet I wonder, if this was just another of those humans imitating a trustworthy AI...
- garganzol 4y agommap does not hide the real RAM usage, it just maps a part of virtual memory address space to an I/O-bound storage medium (e.g. file). That trick does not consume RAM in any significant amounts, so there is nothing to hide. However, such image mapping technique is not always suitable for data-intensive workloads due to page thrashing. If it works for llama.cpp then it can be considered a huge success.
- bragadiru_mafia 4y ago[flagged]
- mhh__ 4y agoSoft page faults considered genius
- plq 4y agommap() basically adds a new swap file that is only to be used for the contents of that specific file. This means it does page eviction on memory pressure and also results in deduplication as mmap'd memory of the same file can be used by any process that has access to it. If the data access is sequential in nature (or at least has some form of locality in function of time) it could end up using less memory indeed, even if the whole file ends up being read during the training process, because unused pages will be evicted by the OS to make room for needed ones. What is interesting to me here is that it's such a marvel to these ML guys. It's a fundamental OS facility and looks like using mmap() instead of using read() on malloc()'d memory should have been the strategy here from the start.
- mort96 4y agoSaying it's like "adding a new swap file that is only to be used for the contents of that specific file" isn't really correct. Swap works by letting the kernel write a memory page to the swap file/partition. After it has done that, the memory is marked as backed by a filesystem (i.e it's no longer anonymous), and the region of physical RAM can be repurposed for something else. With mmap'd files, the memory is always backed by a filesystem; there are no writes, and the physical memory can be discarded at any time.