6 ms·
Slap, slap, slap. I do not think you did a complete or accurate analysis. Randomization is NOT free. Randomization either requires that the values are pre-comp
by szc 6y ago
Slap, slap, slap. I do not think you did a complete or accurate analysis.
Randomization is NOT free. Randomization either requires that the values are pre-computed before install (which means delivering and pre-computing N different versions) or it is computed it on device.
If the randomization is computed on-device how to you validate that the binary or a library has not been "substituted" - persistent malware, APT?
The "compute on device" was a feature of very old macOS versions - it was annoying and took quite a lot of resources.
"TOTAL ASLR" depends on a CPU arch if it fully endorses over all addresses position independent code and data (Q: homework for ARM, x86_64...). If the ABI allows violations of this you cannot glide / slide all code and data addresses without significant runtime costs. This will likely result in a compromise.
- saagarjha 6y agoI believe my analysis is both accurate and complete; actually, I would dispute many of the things you mention. "Baking in" randomization into a binary before installing it is rare; on Apple's platforms ASLR is done at runtime. I am well aware that ASLR does not come for free; it requires position-independent code to function and support from the kernel to do randomization (for image base addresses and anonymous mmaps) and a dynamic linker aware of how to apply relocations. On iOS PIE code has been required for many years, and the various OS subsystems are not only aware of ASLR but ensure validity of code signatures regardless of load address. I suspect the expensive process that you are referring to is shared cache prebinding, which was never a thing on iOS AFAIK and is no longer used on macOS either. To be clear, I am not complaining about a lack of ASLR where it would be prohibitive, such as mapping the shared cache at a different address for every process (which, unless done carefully, would kill the benefits of it being in shared memory as the pages would all be dirtied). I am talking more about various instances where Apple has generally used very poor slides for reasons that aren't all that great, leading to the randomization being easy to break.
- szc 6y agoSorry no. This is not about the main executable - this is a misdirection. The comments you make about the "main binary" are 100% correct. But, they do not address my comment at all. My comments were entirely about the shared cache region and how it can be moved. I again ask you to do the homework - please calculate how it could be done better given the address spaces involved. Address spaces going from 32bit to 64bit have gotten better, but this does need to be kept into consideration given the size of the object involved and the API / CPU instructions available (please, I ask you to consider all of the addressing modes available to the ABI for all of the currently supported platforms) [added] It is likely the individual objects that compose the share cache region are compiled independently. There are lots of individual objects! Resolving the dependencies of the composite shared object are likely expensive. Historically I observe that the contextual data available to a static or dynamic linker has been very constrained, which makes relinking / reallocation objects a challenge.
- joosters 6y agoStop setting homework questions in your posts, please. If you have something relevant to say, why not say it yourself instead of roleplaying a teacher delivering work to your students.
- szc 6y agoYou are right, it is rude, sorry. Will do better.
- saagarjha 6y agoPlease assume good faith–I am familiar with how the shared cache works and my comments apply equally to either the main executable or the shared cache. My point is that the shared cache is not very well randomized at all–for its size, the region it has to slide around in is fairly small. From the original blog post detailing the ASLR break (https://googleprojectzero.blogspot.com/2020/01/remote-iphone-exploitation-part-2.html https://googleprojectzero.blogspot.com/2020/01/remote-iphone...) the number of probes required to find the mapping is quite small, purely because it can only be located within a single 4 GB region for whatever reason. On a system with terabytes of virtual address space, this is a strange choice. Oh, and because you mentioned it: yes, the shared cache has many objects in it. The process of making it is fairly involved, especially since many optimizations go into it (string deduplication, perfect hashing for Objective-C runtime metadata, shortening intra-cache procedure calls, …) But this is all done once when it is built by Apple's B&I, so it's not a problem on-device.
- szc 6y agoSorry, I got animated and committed the faux-pas of familiarity. I forgot that this medium is not the same as a lively discussion between people who "know" each other - I see lots of your posts and the thoughts and depths of your knowledge. There has to be some reason.
- saagarjha 6y agoI really have no issue with assuming familiarity, the part I was concerned about was you calling my comments a "misdirection" when it was not my intention to be misleading. And, I am sure there is a reason, I just have not seen anything that indicates anything good enough that it is worth drastically reducing the quality of the ASLR they provide.
- Hackbraten 6y ago> mapping the shared cache at a different address for every process (which, unless done carefully, would kill the benefits of it being in shared memory as the pages would all be dirtied) Assume foo.dylib and bar.dylib are system libraries, both live in the shared cache, and foo links to bar. Both are loaded and mapped to a running user-space process. If foo links to bar, then there must be a symbol table somewhere in physical memory with an entry that points to bar’s TEXT. That symbol table is part of the shared cache, right? Doesn’t it already follow that bar’s TEXT needs to be at the same virtual address in every process?
- saagarjha 6y ago> That symbol table is part of the shared cache, right? Doesn’t it already follow that bar’s TEXT needs to be at the same virtual address in every process? Yes, and this is how the shared cache works. If you wish to map the shared cache elsewhere there would have to be another copy of it in memory, which is why this would be a massive pessimization if done without designing for it. Perhaps you might have some idea as to what would need to change to make this not be as bad.
- Hackbraten 6y agoGot it, thanks!
- floatboth 6y ago> "Baking in" randomization into a binary before installing it is rare OpenBSD relinks the kernel with new randomization on every boot ("KARL"). Fun stuff, but probably a nightmare to make it work with signature checking :D