5 ms·
you can also do 2M and 1G huge pages on x86, it gets kind of silly fast.
by lanigone 2y ago
you can also do 2M and 1G huge pages on x86, it gets kind of silly fast.
- ShroudedNight 2y ago1G huge pages had (have?) performance benefits on managed runtimes for certain scenarios (Both the JIT code cache and the GC space saw uplift on the SpecJ benchmarks if I recall correctly) If using relatively large quantities of memory 2M should enable much higher TLB hit rates assuming the CPU doesn't do something silly like only having 4 slots for pages larger than 4k ¬.¬
- ignoramous 2y agoWhat? Any pointers on how 1G speeds things up? I'd have taken a bigger page size to wreak havoc on process scheduling and filesystem.
- afr0ck 2y agoBecause of virtual address translation [1] speed up. When a memory access is made by a program, the CPU must first translate the virtual address to a physical address, by walking a hierarchical data structure called a page table [2]. Walking the page tables is slow, thus CPUs implement a small on-CPU cache of virtual-to-physical translations called a TLB [1]. The TLB has a limited number of entries for each page size. With 4 KiB pages, the contention on this cache is very high, especially if the workload has a very large workingset size, therefore causing frequent cache evictions and slow walk of the page tables. With 2 MiB or 1 GiB pages, there is less contention and more workingset size is covered by the TLB. For example, a TLB with 1024 entries can cover a maximum of 4 MiB of workingset memory. With 2 MiB pages, it can cover up to 2 GiB of workingset memory. Often, the CPU has different number of entries for each page size. However, it is known that larger page sizes have higher internal fragmentation and thus lead to memory wastage. It's a trade off. But generally speaking, for modern systems, the overhead of managing memory in 4 KiB is very high and we are at a point where switching to 16/64 KiB is almost always a win. 2 MiB is still a bit of a stretch, though, but transparent 2 MiB pages for heap memory is enabled by default on most major Linux distributions, aka THP [2] Source: my PhD is on memory management and address translation on large memory systems, having worked both on hardware architecture of address translation and TLBs as well as the Linux kernel. I'm happy to talk about this all day! [1] https://blogs.vmware.com/vsphere/2020/03/how-is-virtual-memory-translated-to-physical-memory.html https://blogs.vmware.com/vsphere/2020/03/how-is-virtual-memo... [2] https://docs.kernel.org/admin-guide/mm/transhuge.html https://docs.kernel.org/admin-guide/mm/transhuge.html
- ignoramous 2y agoThanks! > I'm happy to talk about this all day! With noobs, too? ;) > Often, the CPU has different number of entries for each page size. - Does it mean userspace is free to allocate up to a maximum of 1G? I took pages to have a fixed size. - Or, you mean CPUs reserve TLB sizes depending on the requested page size? > With 2 MiB or 1 GiB pages, there is less contention and more workingset size is covered by the TLB - Would memory allocators / GCs need to be changed to deal with blocks of 1G? Would you say, the current ones found in popular runtimes/implementations are adept at doing so? - Does it not adversely affect databases accustomed to smaller page sizes now finding themselves paging in 1G at once? > my PhD is on memory management and address translation on large memory systems If the dissertation is public, please do link it, if you're comfortable doing so.
- deleted 2y ago[deleted]
- afr0ck 2y ago> - Does it mean userspace is free to allocate up to a maximum of 1G? I took pages to have a fixed size. > - Or, you mean CPUs reserve TLB sizes depending on the requested page size? The TLB is a hardware cache with a limited number of entries that cannot dynamically change. Your CPU is shipped with a fixed number of entries dedicated for each page size. Translations of base 4 KiB pages could, for example, have 1024 entries. Translations of 2 MiB pages could have 512 entries and those of 1 GiB usually have a very limited number of only 8 or 16. Nowadays, most CPU vendors increased their 2 MiB TLBs to have the same number of entries dedicated for 4 KiB pages. If you're wondering why they have to be separate caches, it's because, for any page in memory, you can have both mappings at the same time from different processes or different parts of the same process, with possibly different protections. > - Would memory allocators / GCs need to be changed to deal with blocks of 1G? Would you say, the current ones found in popular runtimes/implementations are adept at doing so? > - Does it not adversely affect databases accustomed to smaller page sizes now finding themselves paging in 1G at once? Runtimes and databases have full control and Linux allows per-process policies via madvise) system call. If a program is not happy with huge pages, it can ask the kernel to be ignored, as it can choose to be cooperative. > If the dissertation is public, please do link it, if you're comfortable doing so. I'm still in the PhD process, so no cookies atm :D
- monocasa 2y agoIt's nice for type 1 hypervisors when carving up memory for guests. When page walks for guest virtual to host physical end up taking sixteen levels, a 1G page short circuits that in half to eight.
- afr0ck 2y agoThat's what most hypervisors (e.g. Qemu) do on Linux when THP are enabled and allowed for the process.
- SloopJon 2y agoSearch for huge pages in the documentation of a DBMS that implements its own caching in shared memory: Oracle [1], PostgreSQL [2], MySQL [3], etc. When you're caching hundreds of gigabytes, it makes a difference. Here's a benchmark comparing PostgreSQL performance with regular, large, and huge pages [4]. There was a really bad performance regression in Linux a couple of years ago that killed performance with large memory regions like this (can't find a useful link at the moment), and the short-term mitigation was to increase the huge page size from 2MB to 1GB. [1] https://blogs.oracle.com/exadata/post/huge-pages-in-the-context-of-exadata https://blogs.oracle.com/exadata/post/huge-pages-in-the-cont... [2] https://www.postgresql.org/docs/current/kernel-resources.html#LINUX-HUGE-PAGES https://www.postgresql.org/docs/current/kernel-resources.htm... [3] https://dev.mysql.com/doc/refman/8.4/en/large-page-support.html https://dev.mysql.com/doc/refman/8.4/en/large-page-support.h... [4] https://www.percona.com/blog/benchmark-postgresql-with-linux-hugepages/ https://www.percona.com/blog/benchmark-postgresql-with-linux...