3 ms·
Yes. They are all reasonable ideas, and we actually have experience with most (all?) of them: * Distributing IR applications: We do a form of this by supportin
by lgg 5y ago
Yes. They are all reasonable ideas, and we actually have experience with most (all?) of them:
* Distributing IR applications: We do a form of this by supporting bitcode for App Store submissions on iOS, tvOS, and watchOS. For watchOS it was a great success in the sense that it allowed us to transparently migrate all the binaries from armv7k (a 32 bit ABI on a 32 bit instruction set) to arm64_32 (a 32 bit on a 64 bit instruction set), but we had to carefully design both ABIs in parallel in order to allow that to be efficient. It also introduces serious burdens on developer workflows like crash reporting. Those probably are not significant issues for people deploying binaries to servers, but it can be pretty difficult for developers trying to aggregate crash statistics from apps deployed to consumer devices. It also causes security issues with code provenance since you have to accept locally signed code, have the transforms performed by a centrally trusted entity who signs it, or you have to limit to your optimizations to things you can verify through a provable chain back to the IR.
* Split stacks: We don't do this per se, but we do something semantically equivalent. When we designed the ABI from arm64e we used PAC (pointer authentication codes) to sign the return addresses on the stack, which means that while you can smash the stack and overwrite the return pointer, the authentication code won't match any more and you will crash.
* Pre-mapped libraries: We have done this since then Mac OS X Developer Releases back in the later 90s, though the mechanism has changed a number of times over the years. Originally we manually pre-assigned the address ranges of every dylib in the system (and in fact we carefully placed large gaps between the __TEXT and __DATA of the library and have the libraries in overlapping ranges so that we could place all the system __TEXT in single adjacent region and all the __DATA in a second adjacent region that way we could exploit the batch address translation registers on PPC processors to avoid polluting the TLBs). When the static linker built a dylib we look in a file with the mappings and build the dylib with segment base addresses that we looked up from the that manually maintained list, and when the OS booted we would pre-map all the libraries into their slots so every process just had them mapped.
That was a huge pain in the neck to maintain, since it meant whenever a library grew too large it was would break the optimizations and someone would need update the file and rebuild the OS. It also only dealt with rebasing, to deal with binding we locally ran update_prebinding which users hated because it was slow, sysadmins hated because it means we routinely rewrote every binary on the system. That was also before anyone had deployed ASLR or codesigning.
Nowadays we use the dyld shared cache to achieve similar ends, but it is a far more flexible mechanism. We essentially merge almost all the system dynamic libraries at OS build time into one mega-dylib that we can pre-bind together and sign. We also use VM tricks to allow the system to rebase it on page in rather than the dynamic linker doing any work, and we pre-map it into every address space by default.
We even perform a number of additional optimizations when we build it, such as analyzing the dependencies of every executable in the base system and pre-calculating a lot of the work dyld and even the dynamic runtimes like ObjC would have to do in order to avoid doing them during app launch.
So in short, I think your ideas there have more merit than you perhaps suspect, and from experience I can say if/when you implement them you implement them you might find your views on the trade offs of dynamic vs static linking change.
- ghoward 5y agoThank you. And after reading your response, I might give up on at least one of those ideas.