6 ms·
Inside the Linker (2012)
- userbinator 8y agoReplace "atom" with "section" throughout this article, and you'll have the same thing as what other linkers can do and have done for a very long time, without the additional obfuscatory language: https://elinux.org/Function_sections https://elinux.org/Function_sections https://docs.microsoft.com/en-us/cpp/build/reference/opt-optimizations https://docs.microsoft.com/en-us/cpp/build/reference/opt-opt... http://www.drdobbs.com/cpp/the-most-underused-compiler-switches-in/240166599?pgno=6 http://www.drdobbs.com/cpp/the-most-underused-compiler-switc...
- deleted 8y ago[deleted]
- CalChris 8y agoYou'll need some more explanation for your claim. Second sentence: It is not "section" based like traditional linkers which mostly just interlace sections from multiple object files into the output file.
- comex 8y agoAs input to a linker, Mach-O and ELF files look pretty similar: they both split their contents into named "sections", and then there's a symbol table, which is basically a list of (name, address) pairs. In C code, each symbol generally represents (the start of) a separate function or variable. All references between them are explicitly marked as relocations; thus, the linker should be free to reorder them, and any function/variable that isn't explicitly referenced is unused and can be removed (unless you're linking a shared library and it's meant to be publicly exported from that). On the other hand, in assembly code, symbols are not necessarily independent. You can have things like a function that "falls through" into into another function. For example, here's a hypothetical assembly file that implements both bzero() and memset(), and has the former fall through into the latter: // void bzero(void *s, size_t n) // (zeroes memory) _bzero: // Set up arguments for memset mov r2, r1 // 3rd argument to memset = 2nd argument to bzero mov r1, #0 // 2nd argument to memset is 0 // (1st argument to memset = 1st argument to bzero, no move needed) // Fall through into memset // void *memset(void *b, int c, size_t len); _memset: ...memset implementation... As for sections: traditionally, all code gets put into the same section (named ".text" for ELF, "__text" for Mach-O); all data gets put into the same section (".data" / "__data"); etc. The Darwin (Mach-O) linker is optimized for the properties of C code, and implicitly treats the data between each symbol and the symbol following it as a separate unit ("atom" or "subsection"). Or, more specifically, it does this if the SUBSECTIONS_VIA_SYMBOLS flag is set in the Mach-O header, which is always the case for object files compiled from C (as of 2005 or so). Thus, it always has the ability to remove unused functions/variables, though it doesn't actually bother to do so unless you pass -dead_strip. ELF linkers are more traditional and treat each section as an indivisible unit; symbols aren't taken into consideration at all. So, by default, it's not possible to strip unused functions and variables, nor to reorder multiple symbols that came from a single object file. However, you can pass "-ffunction-sections -fdata-sections" to GCC (the compiler, not the linker) to make it put every single function and variable, respectively, in its own section in the .o file. For example, a function named "foo" would appear in a section called ".text.foo". Then the linker will coalesce all the ".text.*" sections back into a single ".text" output, and similarly for other types of sections. But first it can strip unused sections (if you pass --gc-sections) – which is equivalent to stripping unused functions/variables, since each section contains only a single function/variable. These are basically two different ways to accomplish the same thing, which probably explains what userbinator said. Both approaches feel kind of hacky to me. On the ELF side, object files with a bazillion sections are annoying to look at (if you examine them with readelf or other tools), and putting everything in its own section is not really how sections were originally intended to work. On the Mach-O side, well, the symbol table wasn't originally meant to to be used to split up the input data; in particular, unlike with ELF, Mach-O symbols don't have a size field (which is why the atom implicitly lasts until the next symbol). And it feels wrong that object files compiled from assembly have to be treated differently from everything else (they don't have SUBSECTIONS_VIA_SYMBOLS, unless you explicitly ask for it in the assembly file). Personally I prefer the Mach-O approach just because it requires fewer flags to enable stripping of unused functions/data. Heck, I don't understand why -dead_strip isn't enabled by default. But if you were to design a new object file format from scratch, it could probably handle this much more elegantly than either ELF or Mach-O.
- bear_child 8y agoCan you say more about the last paragraph? I have often wondered if you could avoid identifier mangling with a better designed object file format
- comex 8y agoHmm… well, when it comes to identifier mangling, object formats aren't really the problem. Both ELF and Mach-O use nul termination for symbol (and section) names, so they can't contain 00 bytes, but there's nothing in the binary format preventing them from containing any other bytes. So you could make a symbol named foo::bar(int,int) …and most likely, everything that deals with binaries would have no problem with it. A bigger obstacle might be the assembler, whose input is text. Assembly files usually write symbol names without any escaping or quoting, so non-alphanumeric characters could be misinterpreted. But in fact, it seems that both GNU as and LLVM's assembler (currently used on macOS) allow optionally surrounding symbol names in quotes, allowing those characters to be used: "foo::bar(int,int)": jmp "foo::bar(int,int)" Also, it seems that Clang will use this syntax where necessary when generating assembly files. This compiles: int bar(int a, int b) asm("foo::bar(int,int)"); int bar(int a, int b) { return a + b; } …but GCC apparently doesn't use it; I just tried it on the latest version of GCC (8.1.0) and it produces assembly that uses the name unquoted, which then makes the assembler spit out errors. However, I've left one thing out. I think C++ symbol mangling mainly originated as a hack to support existing assemblers, but it also achieves a basic form of compression. For example, keywords are represented by a single character, and there's a "substitution" syntax for reusing the same token sequence more than once. C++ symbol names already tend to be crazy long, and having a less succinct mangling would make them even longer, which would make binaries larger and might make dynamic linking slower – though to be honest, I have no idea how much (if at all) this would be noticeable in the relative scheme of things. Also, there has to be a single canonical mangling of any given declaration, so even if a platform decided to use C++ syntax directly in symbol names, it would probably omit spaces, unnecessary parentheses, etc., and the result might be harder to read than what you get after demangling. So a demangler might still be desirable. Still, it would certainly be more readable than the current mangling! But that's all assuming that the overall compilation scheme would still look like today, with a 'dumb' linker that only knows about symbols and assembly code, not types or anything about C++ semantics. You could go a step further and design an object format with native support for C++, even things like templates. Imagine being able to define a template in one .cpp file, link it into a library, and then instantiate that template from another executable! That would be enormously cool. In fact, the C++ spec used to define an 'export template' syntax that was supposed to do this, but essentially no compilers implemented it, and it was removed in C++11. (C++ modules are also kind of a form of this, but they're meant to be compiler-specific, private build artifacts rather than something defined at the system level.) I can think of three distinct drawbacks, though: 1. C++ template semantics are very tightly bound to its syntax; there's little you can say about a template definition without knowing what it's instantiated with. Indeed, if you're going to encode template definitions in object files, the format would probably be nothing more complicated than pre-tokenized source code. More modern languages do this somewhat better – in fact, Swift actually plans to have a stable ABI for generics. 2. Similarly, C++ template semantics are very C++-specific; other languages would probably require separate support in the format rather than being able to reuse the C++ functionality. In comparison, existing 'dumb' object formats are basically language-agnostic. 3. The biggest problem: If you allowed templates to be exported from dynamic libraries, the dynamic linker's functionality would have to be transformed from a series of quick name lookups and fixups that can be done every time a binary is launched, to a full-fledged C++ compiler with an expensive code generation step. Even if you cached the output it would still be slow on first launch, especially on low-powered platforms like mobile devices… And yet despite all those drawbacks, I still dream of a system that has… at least some form of this. (I've thought a bit about possible designs: perhaps it could be designed as a component of the package manager rather than of the linker directly.) Why? Well, Debian just started packaging Rust code, and look at how that's going. Each library package ("crate") gets a libfoo-dev that just contains a copy of the package source code, with no libfoo binary package; each executable is statically linked, and a new package version will be released whenever any of its dependencies update. Which is going to mean a lot of redundant upgrades. That's Rust, not C++, but to the extent C++ libraries avoid this problem, it's usually by eschewing templates altogether for anything that's meant to have a stable ABI. If the API does include templates (at least ones that clients instantiate with their own parameters), then clients have to be rebuilt whenever the library changes, same as Rust. I find this quite annoying, since I think the future should be full of ergonomic, strongly-typed APIs taking full advantage of the features of modern languages… yet I don't want to burden sysadmins with pointless upgrades. :) And it just feels wrong that linkers have basically never progressed beyond C.
- gok 8y agoNeeds a (2012), this was written with the Xcode 4.4 release.
- dang 8y agoAdded. Thanks!
- comex 8y agoAnd these days ld64 is on the way toward being obsoleted by LLVM’s lld, although the latter still uses the same basic model.
- AndyKelley 8y agoI wish. In reality the Mach-O LLD code is unmaintained and can not compile a simple hello world program without asserting. https://bugs.llvm.org/show_bug.cgi?id=32254 https://bugs.llvm.org/show_bug.cgi?id=32254
- EdJiang 8y agoRelated: at the most recent WWDC, Apple gave a great presentation about the Xcode build process: https://developer.apple.com/videos/play/wwdc2018/415/ https://developer.apple.com/videos/play/wwdc2018/415/
- User23 8y agoThis is really beautiful HTML
- ljcn 8y agoIndeed, retro. Not quite valid though unfortunately. Apart from the annoying character encoding, doctype, etc. problems, there's a </p> without a <p>, and a <ul> missing a preceding <li>.
- ma2rten 8y agoI work on C++ code at a large tech company. My workflow is such that I compile and run tests often after I made incremental changes. I feel like I often wait for the linker for ages. I am wondering if linking could be sped up. I feel like a lot of the steps mentioned here could be cached for instance.
- int0x80 8y agoIf it does LTO it gets worse. Maybe you can disable it for dev, if its enabled?
- mehrdadn 8y agoCan/do you have incremental linking enabled?
- pjmlp 8y agoDepends what you use as C++ compiler. Visual C++ supports incremental compilation and linking. https://blogs.msdn.microsoft.com/vcblog/2018/03/14/build-time-improvement-recommendation-turn-off-map-use-pdbs/ https://blogs.msdn.microsoft.com/vcblog/2018/03/14/build-tim... https://blogs.msdn.microsoft.com/vcblog/2014/11/12/speeding-up-the-incremental-developer-build-scenario/ https://blogs.msdn.microsoft.com/vcblog/2014/11/12/speeding-... https://blogs.msdn.microsoft.com/vcblog/2018/01/04/visual-studio-2017-throughput-improvements-and-advice/ https://blogs.msdn.microsoft.com/vcblog/2018/01/04/visual-st... https://blogs.msdn.microsoft.com/vcblog/2016/10/05/faster-c-build-cycle-in-vs-15-with-debugfastlink/ https://blogs.msdn.microsoft.com/vcblog/2016/10/05/faster-c-...
- johncolanduoni 8y agoI've seen (sometimes drastic) performance improvements when switching to LLVM's new linker lld[1]. It works quite well on both Linux and Windows in my experience. [1]: https://lld.llvm.org/ https://lld.llvm.org/
- jstimpfle 8y agoI, too, have been wondering that. On a high level, a linker does these things: - Read in objects file and parse linking information (needed symbols etc) - Do some in-memory calculations - resolving symbols and some other stuff. Simple graph theory. Is there any need for more than O(n log n) runtime?. - Write out object files I don't think anything should take so long. No caching needed. I read the series on linkers by gold's author, where significant speedups were claimed. And also, that object-oriented C++ was the right approach to software architecture. I strongly disagree with that second point and didn't understand too much from the article series to be honest. So I am not surprised at all that there are today even much faster linkers.
- AceJohnny2 8y agoThis is the default linker on macOS, LD64. Its "atom"-based concept carried over into LLVM's experimental ATOM-based ld: https://lld.llvm.org/AtomLLD.html https://lld.llvm.org/AtomLLD.html However as AndyKelley notes elsewhere in these comments, that variant of LLD hasn't seen progress in years. That said, that AtomLLD appears to be a side-project of LLD, LLVM's LD project, which appears well-supported on Linux (ELF) and experimental on Windows (PE/COFF). I wonder what Apple's plans are.