11 ms·
The .a file is a relic: Why static archives were a bad idea all along
- tux3 1y agoI actually wrote a tool a to fix exactly this asymmetry between dynamic libraries (a single object file) and static libraries (actually a bag of loose objects) I never really advertised it, but what it does is take all the objects inside your static library, and tells the linker to make a static library that contains a single merged object. https://github.com/tux3/armerge https://github.com/tux3/armerge The huge advantage is that with a single object, everything works just like it would for a dynamic library. You can keep a set of public symbols and hide your private symbols, so you don't have pollution issues. Objects that aren't needed by any public symbol (recursively) are discarded properly, so unlike --whole-archive you still get the size benefits of static linking. And all your users don't need to handle anything new or to know about a new format, at the end of the day you still just ship a regular .a static library. It just happens to contain a single object. I think the article's suggestion of a new ET_STAT is a good idea, actually. But in the meantime the closest to that is probably to use ET_REL, a single relocatable object in a traditional ar archive.
- stabbles 1y agoIt sounds interesting, but I think it's better if a linker could resolve dependencies of static libraries like it's done with shared libraries. Then you can update individual files without having to worry about outdated symbols in these merged files.
- tux3 1y agoIf you mean updating some dependency without recompiling the final binary, that's not possible with static linking. However the ELF format does support complex symbol resolution, even for static objects. You can have weak and optional symbols, ELF interposition to override a symbol, and so forth. But I feel like for most libraries it's best to keep it simple, unless you really need the complexity.
- amluto 1y agoIs there any actual functional difference between the author’s proposed ET_STAT and an appropriately prepared ET_RET file? For that matter, I’ve occasionally wondered if there’s any real reason you can’t statically link an ET_DYN (.so) file other than lack of linker support.
- tux3 1y agoI think everything that you would want to do with an ET_STAT file is possible today, but it is a little off the beaten path, and the toolchain command line options today aren't as simple as for dynamic libraries (e.g. figuring out how to hide symbols in a relocatable object is completely different on the GNU toolchain, LLVM on Linux, or Apple-LLVM which also supports relocatable objects, but has a whole different object file format). I would also be very happy to have one less use of the legacy ar archive format. A little known fact is that this format is actually not standard at all, there's several variants floating around that are sometimes incompatible (Debian ar, BSD ar, GNU ar, ...)
- harryvederci 1y agoMinor suggestion: the article refers to a RHEL 6 developer guide section about static linking. Maybe a more recent article can be used (if their viewpoint hasn't changed).
- dzaima 1y agoHow possible would it be to have a utility that merges multiple .o files (or equivalently a .a file) into one .o file, via changing all hidden symbols to local ones (i.e. alike C's "static")? Would solve the private symbols leaking out, and give a single object file that's guaranteed to link as a whole. Or would that break too many assumptions made by other things?
- Joker_vD 1y agoLike, a linker, with "objcopy --strip-symbols" run as the post-step? I believe you can do this even today.
- dzaima 1y ago--localize-hidden seems to be more what I was thinking of. So this works: ld --relocatable --whole-archive crappy-regular-static-archive.a -o merged.o objcopy --localize-hidden merged.o merged.o This should (?) then solve most issues in the article, except that including the same library twice still results in an error.
- reactordev 1y agoI did this with my dependencies for my game engine. Built them all as libs and used linker to merge them all together. Makes building my codebase as easy as -llibutils
- benreesman 1y agoI routinely tear apart badly laid-out .a files and re-ar them into something useful. It's a few lines of bash.
- tux3 1y agoThis works, but scripting with the ar tool is annoying because it doesn't handle all the edge cases of the .a format. For instance if two libraries have a source file foo.c with the same name, you can end up with two foo.o, and when you extract they override each other. So you might think to rename them, but actually this nonsense can happen with two foo.o objects in the same archive. The errors you get when running into these are not fun to debug.
- amiga386 1y ago> Yet, what if the logger’s ctor function is implemented in a different object file? This is a contrived example akin to "what if I only know the name of the function at runtime and have to dlsym()"? Have a macro that "enables use of" the logger that the API user must place in global scope, so it can write "extern ctor_name;". Or have library specific additions for LDFLAGS to add --undefined=ctor_name There are workarounds for this niche case, and it doesn't add up to ".a files were a bad idea", that's just clickbait. You'll appreciate static linkage more on the day after your program survives a dynamic linker exploit > Every non-static function in the SDK is suddenly a possible cause of naming conflict Has this person never written a C library before? Step 1: make all globals/functions static unless they're for export. Step 2: give all exported symbols and public header definitions a prefix, like "mylibname_", because linkage has a global namespace. C++ namespaces are just a formalisation of this
- Joker_vD 1y ago> This is a contrived example akin to "what if I only know the name of the function at runtime and have to dlsym()"? Well, you just do what the standard Linux loader does: iterate through the .so's in your library path, loading them one by one and doing dlsym() until it succeeds :) Okay, the dynamic loader actually only tries the .so's whose names are explicitly mentioned as DT_NEEDED in the .dynamic section but it still is an interesting design choice that the functions being imported are not actually bound to the libraries; you just have a list of shared objects, and a list of functions that those shared objects, in totality, should provide you with.
- lokar 1y agoAlso, don’t use automatic module init, make the user call an init function at startup. And prefix everything in your library with a unique string.
- layer8 1y agoWhat if you use two libraries A and B that both happen to use library C under the hood? Is the application expected to initialize all dependencies in the right order at the top level? Or is library initialization supposed to be idempotent? This all works as long as libraries are “flat”, but doesn’t scale very well once libraries are built on top of each other and want to hide implementation details.
- benreesman 1y agoIt is unclear to me what the author's point is. Its seems to center on the example of DPDK being difficult to link (and it is a bear, I've done it recently). But its full of strawmen and falsehoods, the most notable being the claims about the deficienies of pkg-config. pkg-config works great, it is just very rarely produced correctly by CMake. I have tooling and a growing set of libraries that I'll probably open source at some point for producing correct pkg-config from packages that only do lazy CMake. It's glorious. Want abseil? -labsl. Static libraries have lots of game-changing advantages, but performance, security, and portability are the biggest ones. People with the will and/or resources (FAANGs, HFT) would laugh in your face if you proposed DLL hell as standard operating procedure. That shit is for the plebs. It's like symbol stripping: do you think maintainers trip an assert and see a wall of inscrutable hex? They do not. Vendors like things good for vendors. They market these things as being good for users.
- throwawayffffas 1y agoCouldn't agree more with you the whole reason docker exists is to avoid having to deal with dynamic libraries we package the whole userland and ship it just to avoid dealing with different dynamic link libraries across systems.
- benreesman 1y agoRight, the popularity of Docker is proof of what users want. The implementation of Docker is proof of how much money you're expected to pay Bezos to run anything in 2025.
- ethin 1y agoThe only exception to this general rule (which, to be clear, I agree with) is when your code for whatever links to LGPL licensed code. A project I'm a major contributor of does this (we have no choice but to use these libraries, due to the requirements we have, though we do it via implib.so (well, okay, the plan is to do that)), and so dynamic linking/DLL hell is the only path we are able to take. If we link statically to the libraries, the LGPL pretty much becomes the GPL.
- stabbles 1y agoMuch of the dynamic section of shared libraries could just be translated to a metadata file as part of a static library. It's not breaking: the linker skips files in archives that are not object files. binutils implemented this with `libdep`, it's just that it's done poorly. You can put a few flags like `-L /foo -lbar` in a file `__.LIBDEP` as part of your static library, and the linker will use this to resolve dependencies of static archives when linking (i.e. extend the link line). This is much like DT_RPATH and DT_NEEDED in shared libraries. It's just that it feels a bit half-baked. With dynamic linking, symbols are resolved and dependencies recorded as you create the shared object. That's not the case when creating static libraries. But even if tooling for static libraries with the equivalent of DT_RPATH and DT_NEEDED was improved, there are still the limitations of static archives mentioned in the article, in particular related to symbol visibility.
- TuxSH 1y ago> This design decision at the source level, means that in our linked binary we might not have the logic for the 3DES building block, but we would still have unused decryption functions for AES256. Do people really not know about `-ffunction-sections -fdata-sections` & `-Wl,--gc-sections` (doesn't require LTO)? Why is it used so little when doing statically-linked builds? > Let’s say someone in our library designed the following logging module: (...) Relying on static initialization order, and on runtime static initialization at all, is never a good idea IMHO
- deleted 1y ago[deleted]
- astrobe_ 1y agoYes, these are really esoteric options, and IIRC GCC's docs say they can be counter-productive.
- jeffbee 1y ago-ffunction-sections has 750k hits on github. It is among the default flags for opt mode builds in Bazel. There are probably people who consider them defaults, in practice.
- astrobe_ 1y agoWell, C and C++ together have around 7M repos, so about 10%. Actually not entirely esoteric, but Github is only a fraction of the world's codebase and users of these repos probably never looked in the makefile, so I'd say 10% of C/C++ developers knowing about this is a very optimistic estimate.
- wtallis 1y agoLooking at GitHub is probably significantly undersampling the kinds of C projects that would be doing static linking, many of which pre-date GitHub.
- 1y ago
- EE84M3i 1y agoSomething I've never quite understood is why can't you statically link against an so file? What specific information was lost during the linking phase to create the shared object that presents that machine code from being placed into a PIE executable?
- sherincall 1y agowcc can do that for you: https://github.com/endrazine/wcc https://github.com/endrazine/wcc
- LtWorf 1y agoYou can, but why?
- EE84M3i 1y agoAt a fundamental level I don't understand why we have two separate file types for static and dynamic libraries. It seems primarily for historical reasons? The author proposes introducing a new kind of file that solves some of the problems with .a filed - but we already have a perfectly good compiled library format for shared libraries! So why can't we make gcc sufficiently smart to allow linking against those statically and drop this distinction?
- alexvitkov 1y agoBecause with the current compilation model shared libraries (.so/.dll) are the output of the linker, but static libraries are input for the linker. It is historical baggage, but as it currently stands they're fairly different beasts.
- high_na_euv 1y ago.so .o .a .pc holy shit, what a mess Why things that are solved in other programming ecosystems are impossible in c cpp world, like sane building system
- sparkie 1y agoBecause those other ecosystems assume that someone has already done the work on the base system and libraries that they don't have to worry about them, and can focus purely on their own little islands.
- pjmlp 1y agoMore like, in other ecosystems, especially the compiled languages that weren't born as part of UNIX like C and C++, the whole infrastructure also takes building and linking as part of the whole language. Note that ISO C and ISO C++ ignore the existence of compilers, linkers and build tools, as per legalese there is some magic way how the code gets turned into machine code, the standards don't even consider the existence of filesystems on header files and translation units locations, they are talked about in the abstract, and can in all standard compliant way be stored in a SQL database.
- adev_ 1y ago> Why things that are solved in other programming ecosystems are impossible in c cpp world, like sane building system This is such an ignorant comment. Most other natively compiled languages have exactly the same concept behind: Object files, Shared Libraries, collection of object and some kind of configuration description of the compilation pipeline. Even high level languages like Rust has that (to some extend). The fact it is buried and hidden under 10 layers of abstraction and fancy tooling for your language does not mean it does not exist. Most languages currently do rely on the LLVM infrastructure (C++) for the linker and their object model anyway. The fact you (probably) never had to manipulate it directly just mean your higher level superficial work never brought you deep enough where it starts to be a problem.
- high_na_euv 1y ago
- rixed 1y agoDo people who write this kind of pieces with such peremptory titles really believe that they finally came about to understand everything better after decades of ignorance? Chesterton’s Fence yada yada?
- cap11235 1y agoWell, it is on medium.com, so probably yes?
- immibis 1y agoLinking works the way it does because of inertia. There's a pretty broad space of possible ways to link (both static and dynamic). Someone once wrote one of them, and it was good enough, so it stuck, and spread, and now everything assumes it. It's far from the only possible way, but you'd have to write new tooling to do it a different way, and that's rarely worth it. The way they selected is pretty reasonable, but not optimal for some use cases, and it may even be pathological for some. At least Linux provides the necessary hooks to change it: your PT_INTERP doesn't have to be /lib64/ld-linux-x86-64.so.2 There are many different ways linking can work (both static and dynamic); the currently selected ways are pretty reasonable points in that space, but not the only ones, and can even be pathological in some scenarios. Actually, a whole lot of things in computing work this way, from binary numbers, to 8-bit bytes, to filesystems, to file handles as a concept, to IP addresses and ports, to raster displays. There were many solutions to a problem, and then one was implemented, and it worked pretty well and it spread, even though other solutions were also possible, and now we can build on top of that one instead of worrying about which one to choose underneath. If you wanted to make a computer from scratch you'd have to decide whether binary is better than decimal, bi-quinary or balanced ternary... or just copy the safe, widespread option. (Contrary to popular belief, very early computers used a variety of number bases other than binary)
- queenkjuul 1y agoWell you'd first have to invent the universe of course
- pjmlp 1y ago
- dale_glass 1y agoOh, static linking can be lots of "fun". I ran into this interesting issue once. 1. We have libshared. It's got logging and other general stuff. libshared has static "Foo foo;" somewhere. 2. We link libshared into libfoo and libbar. 3. libfoo and libbar then go into application. If you do this statically, what happens is that the Foo constructor gets invoked twice, once from libfoo and once from libbar. And also gets destroyed twice.
- rramadass 1y agoBut this is expected behaviour. The Linker cannot know about your intent but is "dumb" in that it only follows some simple rules. Both libfoo and libbar have their own copy of the .o from libshared containing the "Foo foo" instance. Thus the .init/.fini sections in libfoo and libbar make calls to the ctor/dtor of their own "Foo foo" instances resulting in the observed two calls in the app. The way people generally solve this problem is by using a helper class in the library header file which does reference counting for proper initialization/destruction of a single global instance. For an example see std::ios_base::Init in the standard C++ library - https://en.cppreference.com/w/cpp/io/ios_base/Init https://en.cppreference.com/w/cpp/io/ios_base/Init To understand the basics of how linking (both static and dynamic) works see; 1) Hongjiu Lu's ELF: From the Programmer's Perspective - https://ftp.math.utah.edu/u/ma/hohn/linux/misc/elf/elf.html https://ftp.math.utah.edu/u/ma/hohn/linux/misc/elf/elf.html 2) Ian Lance Taylor's 20-part linker essay on his blog; ToC here - https://lwn.net/Articles/276782/ https://lwn.net/Articles/276782/
- layer8 1y ago> Something like a “Static Bundle Object” (.sbo) file, that will be closer to a Shared Object (.so) file, than to the existing Static Archive (.a) file. Is there something missing from .so files that wouldn’t allow them to be used as a basis for static linking? Ideally, you’d only distribute one version of the library that third parties can decide to either link statically or dynamically.
- dwattttt 1y agoShared libraries are linked together in a lossy step. I don't believe it's theoretically impossible; as an unsatisfying proof of concept, you could 'statically' link the .so by archiving it in the final binary, unpacking it at runtime, and dynamically linking it. The static linker would be prevented from seeing multiple copies of code too.
- kazinator 1y ago.a archives can speed up linking of very large software. This is because of assumptions as to the dependencies and the way the traditional Unix-style linker deals with .a files (by default). When a bunch of .o files are presented to the linker, it has to consider references in every direction. The last .o file could have references to the first one, and the reverse could be true. This is not so for .a files. Every successive .a archive presented on the linker command line in left-to-right order is assumed to satisfy references only in material to the left of it. There cannot be circular dependencies among .a files and they have to be presented in topologically sorted order. If libfoo.a depends on libbar.a then libfoo.a must be first, then libbar.a. (The GNU Linker has options to override this: you can demarcate a sequence of archives as a group in which mutual references are considered.) This property of archives (or of the way they are treated by linking) is useful enough that at some point when the Linux kernel reached a certain size and complexity, its build was broken into archive files. This reduced the memory and time needed for linking it. Before that, Linux was linked as a list of .o files, same as most programs.
- parpfish 1y agorelic isn't the right word. relics are really old things that are revered and honored. i think they just want archaic which are old things that are likely obsolete
- Biganon 1y agovestige
- cryptonector 1y agoIt's not that .a files and static linking are a relic, but that static linking never evolved like dynamic linking did. Static linking is stuck with 1978 semantics, while dynamic linking has grown features that prevent the mess that static linking made. There are legit reasons for wanting static linking in 2025, so we really ought to evolve static linking like we did dynamic linking. Namely we should: - make -l and -rpath options in .a generation do something: record that metadata in the .a - make link-edits use that meta- data recorded in .a files in the previous item I.e., start recording dependency metadata in .a files and / so we can stop flattening dependency trees onto the final link-edit. This will allow static linking to have the same symbol conflict resolution behaviors as dynamic linking.
- flohofwoe 1y agoLibrary files are not the problem, deploying an SDK as precompiled binary blobs is ;) (I bet that .a/.lib files were originally never really meant for software distribution, but only as intermediate file format between a compiler and linker, both running as part of the same build process)
- eyalitki 1y agoYeah, but when the product is an SDK, and customers develop on top of it (using their own toolchains) there isn't a lot left for me to play with.
- triknomeister 1y agoSDK could ship the source, lol, stop kneecapping your consumers.
- jhallenworld 1y agoOn the private symbol issue... there is probably a solution to this already. You can partially link a bunch of object files into a single object file (see ld -r). After this is done, 'strip' the file except for those symbols marked with non-hidden visibility- I've not tried to do this, maybe 'strip -x' does the right thing? Not sure.
- eyalitki 1y ago1. "Advanced" compilation environments (meson) probably limit this ability to some extent. 2. Package managers (rpmbuild for instance) mandate build with debug symbols and they do the strip on their own so to create the debug packages. This limits our control of these steps.
- kazinator 1y ago> Yet, what if the logger’s ctor function is implemented in a different object file? Well, tough luck. No one requested this file, and the linker will never know it needs to link it to our static program. The result? crash at runtime. If you have spontaneously called initialization functions as part of an initialization system, then you need to ensure that the symbols are referenced somehow. For instance, a linker script which puts them into a table that is in its own section. Some start-up code walks through the table and calls the functions. This problem has been solved; take a look at how U-boot and similar projects do it. This is not an archive problem because the linker will remove unused .o files even if you give it nothing but a list of .o files on the command line, no archives at all.
- uecker 1y agoIsn't this what partial linking is for, combining object files into a larger one?
- SanjayMehta 1y agoUnix originated on the PDP-11, a machine with very limited memory and disk space. At that time, this was not only the right solution, it was probably the only solution. Calling it “a bad idea all along” is undeserved.