11 ms·
Adding 16 kb page size to Android
- monocasa 2y agoI wonder how much help they had by asahi doing a lot of the kernel and ecosystem work anablibg 16k pages. RISC-V being fixed to 4k pages seems to be a bit of an oversight as well.
- ashkankiani 2y agoIt's pretty cool that I can read "anablibg" and know that means "enabling." The brain is pretty neat. I wonder if LLMs would get it too. They probably would.
- mrob 2y agoLLMs are at a great disadvantage here because they operate on tokens, not letters.
- platelminto 2y agoI remember reading somewhere that LLMs are actually fantastic at reading heavily mistyped sentences! Mistyped to a level where humans actually struggle. (I will update this comment if I find a source)
- thanatropism 2y agoTihs probably refers to comon mispelllings an typo's.
- HeatrayEnjoyer 2y agoIt's actually not. You can scramble every letter within words and it can mostly unscramble it. Keep the first letter and it recovers almost 100%.
- mrbuttons454 2y agoUntil I read your comment I didn't even notice...
- evilduck 2y agoQuestion I wrote: > I encountered the typo "anablibg" in the sentence "I wonder how much help they had by asahi doing a lot of the kernel and ecosystem work anablibg 16k pages." What did they actually mean? GPT-4o and Sonnet 3.5 understood it perfectly. This isn't really a problem for the large models. For local small models: * Gemma2 9b did not get it and thought it meant "analyzing". * Codestral (22b) did not it get it and thought it meant "allocating". * Phi3 Mini failed spectacularly. * Phi3 14b and Qwen2 did not get it and thought it was "annotating". * Mistral-nemo thought it was a portmanteau "anabling" as a combination of "an" and "enabling". Partial credit for being close and some creativity? * Llama3.1 got it perfectly.
- jandrese 2y agoSeems like there is a bit of a roll of the dice there. The ones that got it right may have just been lucky.
- HeatrayEnjoyer 2y agoRan it a few times in new sessions, 0 failures so far.
- slaymaker1907 2y agoI wonder how much of a test this is for the LLM vs whatever tokenizer/preprocessing they're doing.
- Retr0id 2y agofwiw I failed to figure it out as a human, I had to check the replies.
- Alifatisk 2y agoIs there any task Gemma is better at compared to others?
- evilduck 2y agoLocal LLM topics are a treadmill of “what’s best and what is preferred” changing basically weekly to monthly, it’s a rapidly evolving field, but right now I actually tend to gravitate to Gemma2 9b for coding assistance for Typescript work or general question and answer stuff. Its embedded knowledge and speed on the computers that I have (32GB M2 Max, 16GB M1 Air, 4080 gaming desktop) make for a good balance while also using the computer for other stuff, bigger models limit what else I can run simultaneously and are slower than my reading speed, smaller models have less utility and the speed increase is pointless if they’re dumb.
- im3w1l 2y agoI asked chatgpt and it did get it. Personally, when I read the comment my brain kinda skipped over the word since it contained the part "lib" I assumed it was some obscure library that I didn't care about. It doesn't fit grammatically but I didn't give it enough thought to notice.
- IshKebab 2y agoProbably wouldn't be too hard to add a 16 kB page size extension. But I think the Svnapot extension is their solution to this problem. If you're not familiar it lets you mark a set of pages as being part of a contiguously mapped 64 kB region. No idea how the performance characteristics vary. It relieves TLB pressure, but you still have to create 16 4kB page table entries.
- monocasa 2y agoSvnapot is a poor solution to the problem. On one hand it means that that each page table entry takes up half a cache line for the 16KB case, and two whole cache lines in the 64KB case. This really cuts down on the page walker hardware's ability to effectively prefetch TLB entries, leading to basically the same issues as this classic discussion about why tree based page tables are generally more effective than hash based page tables (shifted forward in time to today's gate counts). https://yarchive.net/comp/linux/page_tables.html https://yarchive.net/comp/linux/page_tables.html This is why ARM shifted from a Svnapot like solution to the "translation granule queryable and partially selectable at runtime" solution. Another issue is the fact that a big reason to switch to 16KB or even 64KB pages is to allow for more address range for VIPT caches. You want to allow high performance implementations to be able to look up the cache line while performing the TLB lookup in parallel, then compare the tag with the result of the TLB lookup. This means that practically only the untranslated bits of the address can be used by the set selection portion of the cache lookup. When you have 12 bits untranslated in a address, combined with 64 byte cachelines gives you 64 sets, multiply that by 8 ways and you get the 32KB L1 caches very common in systems with 4KB page sizes (sometimes with some heroic effort to throw a ton of transistors/power at the problem to make a 64KB cache by essentially duplicating large parts of the cache lookup hardware for that extra bit of address). What you really want is for the arch to be able to disallow 4KB pages like on apple silicon which is the main piece that allows their giant 128LB and 192KB L1 caches.
- aseipp 2y ago> What you really want is for the arch to be able to disallow 4KB pages like on apple silicon which is the main piece that allows their giant 128LB and 192KB L1 caches. Minor nit but they allow 4k pages. Linux doesn't support 16k and 4k pages at the same time; macOS does but is just very particular about 4k pages being used for scenarios like Rosetta processes or virtual machines e.g. Parallels uses it for Windows-on-ARM, I think. Windows will probably never support non-4k pages I'd guess. But otherwise, you're totally right. I wish RISC-V had gone with the configurable granule approach like ARM did. Major missed opportunity but maybe a fix will get ratified at some point...
- saagarjha 2y agoProbably very little, since the Android ecosystem is quite divorced from the Linux one.
- wren6991 2y agoRV64 has some reserved encoding space in satp.mode so there's an obvious path to expanding the number of page table formats at a later time. Just requires everyone to agree on the direction (common issue with RISC-V). For RV32 I think we are probably stuck with Sv32 4k pages forever.
- twoodfin 2y agoA little additional background: iOS has used 16KB pages since the 64-bit transition, and ARM Macs have inherited that design.
- arghwhat 2y agoA more relevant bit of background is that 4KB pages lead to quite a lot of overhead due to the sheer number of mappings needing to be configured and cached. Using larger pages reduce overhead, in particular TLB misses as fewer entries are needed to describe the same memory range. While x86 chips mainly supports 4K, 2M and 1G pages, ARM chips tend to support more practical 16K page sizes - a nice balance between performance and wasting memory due to lower allocation granularity. Nothing in particular to do with Apple and iOS.
- CalChris 2y agoArmv8-A also supports 4K pages: FEAT_TGran4K. So Apple did indeed make a choice to instead use 16K, FEAT_TGran16K. Microsoft uses 4K for AArch64 Windows.
- jsheard 2y agoMakes me wonder how much performance Windows is leaving on the table with its primitive support for large pages. It does support them, but it doesn't coalesce pages transparently like Linux does, and explicitly allocating them requires special permissions and is very likely to fail due to fragmentation if the system has been running for a while. In practice it's scarcely used outside of server software which immediately grabs a big chunk of large pages at boot and holds onto them forever.
- arghwhat 2y agoQuite a bit, but 2M is an annoying size and the transparent handling is suboptimal. Without userspace cooperating, the kernel might end up having to split the pages at random due to an unfortunate unaligned munmap/madvise from an application not realizing it was being served 2M pages. Having Intel/AMD add 16-128K page support, or making it common for userspace to explicitly ask for 2M pages for their heap arenas is likely better than the page merging logic. Less fragile. 1G pages are practically useless outside specialized server software as it is very difficult to find 1G contiguous memory to back it on a “normal” system that has been running for a while.
- a1o 2y ago> The very first 16 KB enabled Android system will be made available on select devices as a developer option. This is so you can use the developer option to test and fix > once an application is fixed to be page size agnostic, the same application binary can run on both 4 KB and 16 KB devices I am curious about this. When could an app NOT be agnostic to this? Like what an app must be doing to cause this to be noticeable?
- sweeter 2y agoWine doesn't work on 16 KB page size among other things.
- mananaysiempre 2y agoThis seems especially peculiar given Windows has a 64K mapping granularity.
- tredre3 2y agoWindows uses 4KB pages.
- nullindividual 2y ago4K, 2M ("large page"), or 1G ("huge page") on x86-64. A single allocation request can consist of multiple page sizes. From Windows Internal 7th Edt Part 1: On Windows 10 version 1607 x64 and Server 2016 systems, large pages may also be mapped with huge pages, which are 1 GB in size. This is done automatically if the allocation size requested is larger than 1 GB, but it does not have to be a multiple of 1 GB. For example, an allocation of 1040 MB would result in using one huge page (1024 MB) plus 8 “normal” large pages (16 MB divided by 2 MB).
- mananaysiempre 2y agoRight (on x86-32 and -64, because you can’t have 64KB pages there, though larger page sizes do exist and get used). You still cannot (e.g.) MapViewOfFile() on an address not divisible by 64KB, because Alpha[1]. As far as I understand, Windows is mostly why the docs for the Blink emulator[2] (a companion project of Cosmopolitan libc) tell you any programs under it need to use sysconf(_SC_PAGESIZE) [aka getpagesize() aka getauxval(AT_PAGESZ)] instead of assuming 4KB. [1] https://devblogs.microsoft.com/oldnewthing/20031008-00/?p=42223 https://devblogs.microsoft.com/oldnewthing/20031008-00/?p=42... [2] https://github.com/jart/blink/blob/master/README.md#compiling-and-running-programs-under-blink https://github.com/jart/blink/blob/master/README.md#compilin...
- lostmsu 2y agoNot entirely related (except the block size), but I am considering making and standardizing a system-wide content-based cache with default block size 16KB. The idea is that you'd have a system-wide (or not) service that can do two or three things: - read 16KB block by its SHA256 (also return length that can be <16KB), if cached - write a block to cache - maybe pin a block (e.g. make it non-evictable) I would be like a block-level file content dedup + eviction to keep the size limited. Should reduce storage used by various things due to dedup functionality, but may require internet for corresponding apps to work properly. With a peer-to-peer sharing system on top of it may significantly reduce storage requirements. The only disadvantage is the same as with shared website caches prior to cache isolation introduction: apps can poke what you have in your cache and deduce some information about you from it.
- monocasa 2y agoI'd probably pick a size greater than 16KB for that. Windows doesn't expose translations less than 64KB in their version of mmap, and internally their file cache works in increments of 256KB. And these were numbers they picked back in the 90s.
- treyd 2y agoI would go for higher than 16K. I believe BitTorrent's default minimum chunk size is 64K, for example. It really depends on the use case in question though, if you're doing random writes then larger chunk sizes quickly waste a ton of bandwidth, especially if you're doing recursive rewrites of a tree structure. Would a variable chunk size be acceptable for whatever it is you're building?
- lostmsu 2y agoI could feasibly do 2x partitioning. E.g. have caches for 16KB, 32KB, etc provided there's some mechanism to automatically combine/split pieces
- devit 2y agoSeems pretty dubious to do this without adding support for having both 4KB and 16KB processes at once to the Linux kernel, since it means all old binaries break and emulators which emulate normal systems with 4KB pages (Wine, console emulators, etc.) might dramatically lose performance if they need to emulate the MMU. Hopefully they don't actually ship a 16KB default before supporting 4KB pages as well in the same kernel. Also it would probably be reasonable, along with making the Linux kernel change, to design CPUs where you can configure a 16KB pagetable entry to map at 4KB granularity and pagefault after the first 4KB or 8KB (requires 3 extra bits per PTE or 2 if coalesced with the invalid bit), so that memory can be saved by allocating 4KB/8KB pages when 16KB would have wasted padding.
- username81 2y agoShouldn't there be some kind of setting to change the page size per program? AFAIK AMD64 CPUs can do this.
- saagarjha 2y agoYes, ARM CPUs can do it too.
- fouronnes3 2y agoCould they upstream that or would that require a fork?
- mgaunard 2y agowhy does it break userland? if you need to know the page size, you should query sysconf SC_PAGESIZE.
- eyalitki 2y agoRHEL tried that in that past with 64KB on AARCH64, it led to MANY bugs all across the software stack, and they eventually reverted it - https://news.ycombinator.com/item?id=27513209 https://news.ycombinator.com/item?id=27513209. I'm impressed by the effort on Google's side, yet I'll be surprised if this effort will pay off.
- nektro 2y agoapple's m-series chips use a 16kb page size by default so the state of things has improved significantly with software wanting to support asahi and other related endeavors
- kcb 2y agoNvidia is pushing 64KB pages on their Grace-Hopper system.
- rincebrain 2y agoI didn't realize they had reverted it, I used to run RHEL builds on Pi systems to test for 64k page bugs because it's not like there's a POWER SBC I could buy for this.
- daghamm 2y agoCan someone explain those numbers to me? 5-10% performance boost sounds huge. Wouldn't we have much larger TLBd if page walk was really this expensive? On the other hand 9% increase in memory usage also sounds huge. How did this affect memory usage that much?
- scottlamb 2y ago> 5-10% performance boost sounds huge. Wouldn't we have much larger TLBd if page walk was really this expensive? It's pretty typical for large programs to spend 15+% of their "CPU time" waiting for the TLB. [1] So larger pages really help, including changing the base 4 KiB -> 16 KiB (4x reduction in TLB pressure) and using 2 MiB huge pages (512x reduction where it works out). I've also wondered why the TLB isn't larger. > On the other hand 9% increase in memory usage also sounds huge. How did this affect memory usage that much? This is the granularity at which physical memory is assigned, and there are a lot of reasons most of a page might be wasted: * The heap allocator will typically cram many things together in a page, but it might say only use a given page for allocations in a certain size range, so not all allocations will snuggle in next to each other. * Program stacks each use at least one distinct page of physical RAM because they're placed in distinct virtual address ranges with guard pages between. So if you have 1,024 threads, they used at least 4 MiB of RAM with 4 KiB pages, 16 MiB of RAM with 16 KiB pages. * Anything from the filesystem that is cached in RAM ends up in the page cache, and true to the name, it has page granularity. So caching a 1-byte file would take 4 KiB before, 16 KiB after. [1] If you have an Intel CPU, toplev is particularly nice for pointing this kind of thing out. https://github.com/andikleen/pmu-tools https://github.com/andikleen/pmu-tools
- 95014_refugee 2y ago> I've also wondered why the TLB isn't larger. Fast CAMs are (relatively) expensive, is the excuse I always hear.
- ein0p 2y agoNo mention of Apple on the page. Apple has been using 16K pages for years now.
- zahlman 2y agoWhy would an "Android news" blog mention what competitors are doing?
- ein0p 2y agoBecause the whole thing sounds like they’re doing something new, but they’re just catching up to something Apple has done back when they switched to aarch64.
- deadlydose 2y agoA page right out of Apple's playbook then.
- pflanze 2y agoI would expect that this increases the gap between new and old phones / makes old phones unusable more quickly: new phones will typically have enough RAM and can live with the 9% less efficient memory use, and will see the 5-10% speedup. Old phones are typically bottlenecked at RAM, now 9% earlier, and reloading pages from disk (or swapping if enabled) will have a much higher overhead than 5-10%.
- dboreham 2y agoTime to grab some THP popcorn...
- taeric 2y agoI see they have measured improvements in the performance of some things. In particular, the camera app starts faster. Small percentage, but still real. Curious if there are any other changes you could do based on some of those learnings? The camera app, in particular, seems like a good one to optimize to start instantly. Especially so with the the shortcut "double power key" that many phones/people have setup. Specifically, I would expect you should be able to do something like the lisp norm of "dump image?" Startup should then largely be loading the image, not executing much if any initialization code? (Honestly, I mostly assume this already happens?)
- saagarjha 2y agoA big part of the challenge for launching the camera app is getting the hardware ready and quickly freeing up RAM for image processing.
- taeric 2y agoMakes sense. That it is so much faster on repeat wakeups does seem to hint that it could be computing a something. I'm assuming you are saying that most of what is getting computed is related to paging in/out memory? That would track on how it could be better with larger pages.
- quotemstr 2y agoGood. It's about time. 4KB pages come down to us from 32-bit time immemorial. We didn't bump the page size when we doubled the sizes of pointers and longs for the 64-bit transition. 4KB has been way too small for ages, and I'm glad we're biting the minor compatibility bullet and adopting a page size more suited to modern computing.
- jeffbee 2y agoNow do 512B LBAs on NVMe devices.
- lxgr 2y agoNow I wonder: Does increased page size have any negative impacts on I/O performance or flash lifetime, e.g. for writebacks of dirty pages of memory-mapped files where only a small part was changed? Or is the write granularity of modern managed flash devices (such as eMMCs as used in Android smartphones) much larger than either 4 or 16 kB anyway?
- tadfisher 2y agoFlash controllers expose blocks of 512B or 4096KB, but the actual NAND chips operate in terms of "erase blocks" which range from 1MB to 8MB (or really anything); in these blocks, an individual bit can be flipped from "0" to "1" once, and flipping any bit back to "0" requires erasing the entire block and flipping the desired bits back to "1" [0]. All of this is hidden from the host by the NAND controller, and SSDs employ many strategies (including DRAM caching, heterogeneous NAND dies, wear-leveling and garbage-collection algorithms) to avoid wearing the storage NAND. Effectively you must treat flash storage devices as block devices of their advertised block size because you have no idea where your data ends up physically on the device, so any host-side algorithm is fairly worthless. [0]: https://spdk.io/doc/ssd_internals.html https://spdk.io/doc/ssd_internals.html
- lxgr 2y agoWrites on NAND happen at the block, not the page level, though. I believe the ratio between the two is usually something like 1:8 or so. Even blocks might still be larger than 4KB, but if they’re not, presumably a NAND controller could allow such smaller writes to avoid write amplification? The mapping between physical and logical block address is complex anyway because of wear leveling and bad block management, so I don’t think there’s a need for write granularity to be the erase block/page or even write block size.
- to11mtm 2y ago> Even blocks might still be larger than 4KB, but if they’re not, presumably a NAND controller could allow such smaller writes to avoid write amplification? Look at what SandForce was doing a decade+ ago. They had hardware compression to lower write amp and some sort of 'battery backup' to ensure operations completed. Various bits of this sort of tech is in most decent drives now. > The mapping between physical and logical block address is complex anyway because of wear leveling and bad block management, so I don’t think there’s a need for write granularity to be the erase block/page or even write block size. The controller needs to know what blocks can get a clean write vs what needs an erase; that's part of the trim/gc process they do in background. Assuming you have sufficient space, it works kinda like this: - Writes are done to 'free-free' area, i.e. parts of the flash it can treat like SLC for faster access and less wear. If you have less than 25%-ish of drive free this becomes a problem. Controller is tracking all of this state. - When it's got nothing better to do for a bit, controller will work to determine which old blocks to 'rewrite' with data from the SLC-treated flash into 'longer lived' but whatever-Level-cell storage. I'm guessing (hoping?) there's a lot of fanciness going on there, i.e. frequently touched files take longer to get a full rewrite. TBH sounds like a fun thing to research more
- CalChris 2y agoiOS has had 16K pages since forever. OSX switched to 16K pages in 2020 with the M1. Windows is stuck on 4K pages, even for AArch64. Linux has various page sizes. Asahi is 16K.
- nullindividual 2y agoWindows has 4K, 2M, and 1G page sizes on x86-64.
- CalChris 2y agoNormal, large and huge. But default normal pages (which is what Android is changing) are 4K. FWIW, Itanium and Alpha had 8K default pages. https://devblogs.microsoft.com/oldnewthing/20210510-00/?p=105200 https://devblogs.microsoft.com/oldnewthing/20210510-00/?p=10... I wonder why Microsoft stayed with 4K for AArch64.
- Kwpolska 2y agoMicrosoft wanted to make x86 compatibility as painless as possible. They adopted an ABI in which registers can be generally mapped 1:1 between the two architectures.
- nullindividual 2y agoI was confused as to why you were posting incorrect information when this thread already contained the correct information.
- baby_souffle 2y ago> and 1G page sizes on x86-64. I wonder who requested the 1G page size be implemented and what they use it for...
- Kwpolska 2y agoAnother thread says virtual machines.
- iam-TJ 2y agoIn Debian kernel we've very recently enabled building an ARM64 kernel flavour with 16KiB page size and we've discussed adding a 64KiB flavour at some point as is the case for PowerPC64 already. This will likely reveal bugs that need fixing in some of the 70,000+ packages in the Debian archive. That ARM64 16KiB page size is interesting in respect of the Apple M1 where Asahi [0] identified that the DART IOMMU has a minimum page size of 16KiB so using that page size as a minimum for everything is going to be more efficient. [0] https://asahilinux.org/2021/10/progress-report-september-2021/ https://asahilinux.org/2021/10/progress-report-september-202...
- antalya070707 2y ago[dead]
- antalya070707 2y ago[dead]