4 ms·
Every time I read one of these essays extolling the virtues of static linking it makes me pretty sad. Engineering is about trade offs, and sometimes it does mak
by lgg 5y ago
Every time I read one of these essays extolling the virtues of static linking it makes me pretty sad. Engineering is about trade offs, and sometimes it does make sense to statically link, but the fact that those trade offs are so lopsided on Linux has very little to do with dynamic linking vs static linking as generic concepts, and more to do with the fact that the design of ld.so is straight out of the 1990s and almost nothing has be done to either exploit the benefits dynamic linking brings, nor to mitigate the issues it causes.
On Darwin derived systems (macOS, iOS, tvOS, watchOS) we have invested heavily in dynamic linking over the last two decades. That includes features we use to improve binary compatibility (things like two level namespaces (aka Direct Binding in ELF) and umbrella frameworks), middleware distribution (through techniques like bundling resources with their dylibs into frameworks), and mitigate the security issues (through technologies like __DATA_CONST).
Meanwhile, we also substantially reduced the cost of dynamic linking through things like the dyld shared cache (which makes dynamic linking of most system frameworks almost free, and in practice it often reduces their startup cost to below the startup cost of statically linking them once you include the cost of rebasing for PIE), and mitigate much of the rest of the cost of dynamic linking through things like pre-calculated launch closures. It does not hurt that my team owns both the static and dynamic linkers. We have a tight design loop so that when we come up with ideas for how to change libraries to make the dynamic linker load them faster we have the ability to rapidly deploy those changes through the ecosystem.
As I explained in here[1] dynamic linking on macOS is way faster than on Linux because we do it completely differently, and because of that it is used way more pervasively on our systems (a typical command line binary loads 80-150 dylibs, a GUI app around 300-800 depending on which of our OS you are talking about). IOW, by mitigating the costs we drove up the adoption which amplifies the benefits.
And it is not like we have squeezed all the performance out of the dynamic linker that we can, not by a long shot. We have more ideas than we have time to pursue, as well as additional improvements to static linking. If you have any interest in either static or dynamic linking we're looking for people to work on both[2].
[1]: https://www.realworldtech.com/forum/?threadid=197081&curpostid=197486 https://www.realworldtech.com/forum/?threadid=197081&curpost...
[2]: https://jobs.apple.com/en-us/details/200235669/systems-engineer-dyld https://jobs.apple.com/en-us/details/200235669/systems-engin...
(edit: fixed link formatting)
- ghoward 5y agoAuthor here. I'm curious: did you read to the end of the post where I lay out ideas about how to get the benefits of both in one?
- lgg 5y agoYes. They are all reasonable ideas, and we actually have experience with most (all?) of them: * Distributing IR applications: We do a form of this by supporting bitcode for App Store submissions on iOS, tvOS, and watchOS. For watchOS it was a great success in the sense that it allowed us to transparently migrate all the binaries from armv7k (a 32 bit ABI on a 32 bit instruction set) to arm64_32 (a 32 bit on a 64 bit instruction set), but we had to carefully design both ABIs in parallel in order to allow that to be efficient. It also introduces serious burdens on developer workflows like crash reporting. Those probably are not significant issues for people deploying binaries to servers, but it can be pretty difficult for developers trying to aggregate crash statistics from apps deployed to consumer devices. It also causes security issues with code provenance since you have to accept locally signed code, have the transforms performed by a centrally trusted entity who signs it, or you have to limit to your optimizations to things you can verify through a provable chain back to the IR. * Split stacks: We don't do this per se, but we do something semantically equivalent. When we designed the ABI from arm64e we used PAC (pointer authentication codes) to sign the return addresses on the stack, which means that while you can smash the stack and overwrite the return pointer, the authentication code won't match any more and you will crash. * Pre-mapped libraries: We have done this since then Mac OS X Developer Releases back in the later 90s, though the mechanism has changed a number of times over the years. Originally we manually pre-assigned the address ranges of every dylib in the system (and in fact we carefully placed large gaps between the __TEXT and __DATA of the library and have the libraries in overlapping ranges so that we could place all the system __TEXT in single adjacent region and all the __DATA in a second adjacent region that way we could exploit the batch address translation registers on PPC processors to avoid polluting the TLBs). When the static linker built a dylib we look in a file with the mappings and build the dylib with segment base addresses that we looked up from the that manually maintained list, and when the OS booted we would pre-map all the libraries into their slots so every process just had them mapped. That was a huge pain in the neck to maintain, since it meant whenever a library grew too large it was would break the optimizations and someone would need update the file and rebuild the OS. It also only dealt with rebasing, to deal with binding we locally ran update_prebinding which users hated because it was slow, sysadmins hated because it means we routinely rewrote every binary on the system. That was also before anyone had deployed ASLR or codesigning. Nowadays we use the dyld shared cache to achieve similar ends, but it is a far more flexible mechanism. We essentially merge almost all the system dynamic libraries at OS build time into one mega-dylib that we can pre-bind together and sign. We also use VM tricks to allow the system to rebase it on page in rather than the dynamic linker doing any work, and we pre-map it into every address space by default. We even perform a number of additional optimizations when we build it, such as analyzing the dependencies of every executable in the base system and pre-calculating a lot of the work dyld and even the dynamic runtimes like ObjC would have to do in order to avoid doing them during app launch. So in short, I think your ideas there have more merit than you perhaps suspect, and from experience I can say if/when you implement them you implement them you might find your views on the trade offs of dynamic vs static linking change.