7 ms·
Speeding up ELF relocations for store-based systems
- kreetx 2y agoThis article confuses static linking and deterministic builds. I.e Nix, a "store-based system" (author's term), still very much dynamically links. Static linking means copying actual program code from libraries into the executable itself, such that external .so files don't need to be loaded.
- geddawm 2y agoI believe you're referring to: > Store-based systems, however, are static in nature, with all dependencies being resolved at build time. I think the author is saying that the shared libraries (.so) are available at build time on store-based systems and never change. Thus, the dynamic linker can speed up symbol resolution by doing the symbol resolution at build time and sticking the result in output binary. This is distinct from static linking which sticks the entire library (.a) into the output binary.
- lgg 2y agoWindows and macOS both use a form of two level name-spacing, which does the same sort of direct binding to a target library for each symbol. Retrofitting that into a binary format is pretty simple, but retrofitting it into an ecosystem that depends on the existing flat namespace look up semantics is not. I think it is pretty clever that the author noticed the static nature of the nix store allows them to statically evaluate the symbol resolutions and get the launch time benefits of two level namespaces. I do wonder if it might make more sense to rewrite the binaries to use Direct Binding[1]. That is an existing encoding of library targets for symbols in ELF that has been used by Solaris for a number of years. 1: https://en.wikipedia.org/wiki/Direct_binding https://en.wikipedia.org/wiki/Direct_binding
- JonChesterfield 2y agoThat is much better than the Linux model! Not only is there less crawling around looking for symbols, you're no longer in trouble when two libraries export the same symbol. Especially given libraries are found by name, and symbols by name, where "type information" or "is that actually the library I wanted" are afterthoughts.
- cryptonector 2y ago> you're no longer in trouble when two libraries export the same symbol. Whether you use direct binding or symbol versioning, either way you don't have a problem with multiple libraries exporting the same symbol. By the way, this is the fundamental problem with static linking for C: it's still stuck with 1970s semantics and you can't get the same symbol conflict resolution semantics as with ELF because the static linker-editors do not record dependencies in static link archives. The key insight is that when you link-edit your libraries and programs you should provide only the direct dependencies, and the linker-editor should then record in its output which one of those provided which symbol. Compare to static linking where only the final edit gets the dependency information and that dependency tree has to get flattened (because it has to fit on a command-line, which is linear in nature).
- o11c 2y agoThe advantage of the Linux model is that you can refactor which library actually contains a function, which is done quite often.
- cryptonector 2y agoFilters and auxiliary filters also do that, but Linux doesn't support them well, which is really sad because they make for a very neat system.
- glandium 2y agoI think you can get an effect similar to direct binding with symbol versioning.
- cryptonector 2y agoAnother option is the Solaris/Illumos "direct binding" scheme where each object stores for each external symbol the SONAME of the object meant to provide it. It's a lot like pre-linking, but a) less intrusive, b) less fast (because it's less intrusive). EDIT: Ay, https://news.ycombinator.com/item?id=40268546 https://news.ycombinator.com/item?id=40268546 mentions this.
- cryptonector 2y agoI don't like the term "store-based" for what Nix does. Nix uses a partial transitive universe hash to compute deploy-time locations for built artifacts at build configuration time so that those can be hard-coded into built artifacts -- "store-based" hardly captures this. I don't know what to call this, but "store-based" is insufficiently suggestive.
- Hello71 2y agoThis is interesting, but I'm not sure it actually makes much difference for realistic workloads. ffmpeg probably has the most link-time dynamic linking of any Linux package, and ffmpeg -loglevel quiet takes only ~40 ms on Alpine and ~60 ms on Debian. Other programs using ffmpeg tend to either statically link it (Chromium) or link it at run-time (Firefox, most video editors), neither of which would be improved by this optimization. Would it be nice to shave 60 ms off of every ffmpeg/mpv invocation? In isolation, sure, but considering the maintenance burden and potential inconsistencies I don't think it's worth it. Nix is supposed to ensure that the dependencies are always the same, but currently if something breaks somehow, the wrong version will be loaded or an error will be emitted, whereas with this optimization, it will crash or silently invoke the wrong functions which seems extremely difficult to debug.
- setheron 2y agomy thinking: 40ms might matter if you consider the number of times a process might restart, the number of instances you run and on how many hosts. That could be considerable wasted cycles in aggregate. (i.e. cloud provider) Also it's not only a function of how many shared libraries but how many symbols they each individually have as well -- also their symbol length as well.
- bfrog 2y agoThis is really neat! I’ll have to take a closer look and see if some ideas could be reused in zephyr’s new elf loading facilities
- setheron 2y agoCool! Reach out to me. I'd love to experiment with this + additional ideas that further OS when more control is known such as in NixOS. I have another idea that's quite crazy but I have done some rough benchmarking to prove its efficacy.
- Hortinstein 2y agoHahaha funny seeing you @fzakaria here, knew I recognized that name! We worked on a hypemachine scraper a long time ago 12 or 13 years ago, glad to see you are still around writing great software and interesting articles!
- setheron 2y agoWow! That was a nice reply to see. Made my week. The knowledge of that scraper went into a Chrome extension that sees some good downloads to this day (I guess people this use hypemachine... I have kids now so my music listening time is on a pause) Hope you are well as well.
- phire 2y ago> This optimization is only possible for store-based systems since the set of shared libraries is fixed and immutable. In a traditional Linux distribution, each shared library could be updated at any time, which would invalidate the cached symbol resolutions and their offsets. Yes and no. The optimisation of having a cache could be implemented on more traditional linux systems. You would just need to check the modification time of all the shared libraries hasn't changed before using the cache. Alternatively, the job could be given to the package manager, make it invalidate and regenerate the cache for any binaries that that use and shared libraries that have been updated. What a store-based system does is make it so much simpler to implement such optimisations, because you simply don't need to worry about various invalidation based edge cases.
- setheron 2y ago(author) Of course with utmost care a lot of things are possible.... ;) unfortunately on a traditional system nothing is stopping anything (user?) from mucking with the system outside the knowledge of the package manager.