12 ms·
The C++ Build Process Explained
- amelius 8y agoAside, C++ would have been so much nicer without the need for header files ...
- garaetjjte 8y agoThere is modules TS: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/n4720.pdf http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/n472...
- grandmczeb 8y agoThe draft for the merged version expected to be in C++20 is here: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p1103r1.pdf http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p110...
- flukus 8y agoIf java is any guide developers would go out of their way to add interfaces everywhere anyway, often they're essentially header files for the OO gods.
- kjeetgill 8y agoI haven't done any C++ in a while, can you swap out the implementation backing a header? I suppose at the linker layer maybe but even that is just once globally for the program. Right? The OOP God's demand polymorphism! If memory serves me, the C++ version is a class with all empty virtual functions.
- flukus 8y agoYeah, it can only be swapped out at link time (without crazy magic). That satisfies two very common cases though, when there is only one implementation (and the interface is a sacrifice to the OOP gods) and when the second implementation is for unit testing only. Those 2 plus other crazy rules (like only one interface defined per file) mean that many projects will end up with just as many interface files as a c++ project would header files.
- tjoff 8y agoIn principle I agree. An IDE should be able to automatically generate something like a header file regardless of language. But still, no IDE or tool I've seen does this better. A header file gives you a very nice overview, and it really isn't any hassle to speak of to keep them up to date. It takes getting used to (as with everything when starting out with a new language), but before you know it you might even start to miss them in other languages.
- int_19h 8y agoSeparation of module interface and module implementation is useful, but headers are a horrible way to do that. For an example of how it can be done right, look at Borland Pascal dialects (with "interface" and "implementation" section in each unit, where "interface" can be extracted if desired), or Ada, or ML.
- tech_tuna 8y ago>it really isn't any hassle to speak of to keep them up to date Dear lord, all those broken builds I've seen over the years would love to talk to you about that.
- setpatchaddress 8y agoThat has not been my experience in 35 years of dealing with headers in C and C++ in large-scale software projects. Ruby, Swift, and even Python (or Pascal if we want to go way back) are significantly easier to deal with for anything sophisticated.
- kazinator 8y ago> An IDE should be able to automatically generate something like a header file regardless of language. https://www.hwaci.com/sw/mkhdr/ https://www.hwaci.com/sw/mkhdr/
- pcwalton 8y agoHeader files absolutely are a hassle to keep up to date. Besides, they're a terrible experience. You have to pay attention to include order! You have to write forward declarations! You have to write include guards, in 2018! They're also bad as an overview. They may have been good in the 1980s, but nowadays a proper documentation generator gives you better formatting, search features, and cross-references. Inline documentation is particularly obnoxious in header files: I like seeing function-level documentation alongside interface and implementation, but if I put the documentation in both the .cpp and the .h file it's duplicated in two places and easily gets out of date.
- muricula 8y agoHeaders allow you to ship a binary without the full source code. If you want to build a linux kernel module you don't need all of the linux sources, just the headers.
- fermienrico 8y agoIn principle you're right, but the notion that header file exposes just the "interface" is completely false. Class definition, private variables and functions, etc. are all exposed in header file. Header files are not a way to only expose the interface. You give up a lot more in C++. I've never had to deal with pesky header files until I started developing C and it immediately struct me as a royal pain in the ass. Even after couple of years of developing C/C++, I find the whole concept of header files archaic. Include preprocessor directive literally copy pastes stuff with no intelligence what-so-ever. The user is now burdened to ensure #includes are guarded to what I call a patchy half-baked hacked up solution - #IFDEF/#DEFINE/#ENDIF and #pragma in C++. It should be handled automatically by the compiler/preprocessor or IDE and I believe it is now being addressed in the C++17/20 spec with the advent of "modules". This thing should have been written up way back in 1989.
- jlarocco 8y ago> In principle you're right, but the notion that header file exposes just the "interface" is completely false. Sorry, but it is not "completely false". Doing it properly requires a carefully designed interface to hide internal data structures, and splitting out the end user headers from the internal headers, but it works. > I've never had to deal with pesky header files until I started developing C and it immediately struct me as a royal pain in the ass. It's like saying, "I've never had to deal with pesky .py files until I started developing Python."
- fermienrico 8y agoHow do you compare .py files with .h files? I am confused as to how you are drawing this analogy.
- jacobush 8y agoI use https://github.com/mjspncr/lzz3 https://github.com/mjspncr/lzz3 - it generates header files and source files. If feels quite weird and nice. :)
- drkl 8y agoI dont care
- drkl 8y agoGood information
- joker3 8y agoI've been looking for something like this to send to some C++ newbies. This is almost what I need but not exactly. Is there something similar that explains how libraries work?
- aidenn0 8y agoLibraries are just a bundle of .o files (e.g. run "ar t /usr/lib/x86_64-linux-gnu/libpython2.7.a" on an ubuntu with devtools installed and you'll see all the .o files that are in libpython2.7). The special thing about them though is that the linker will not include any .o files for which no symbols are referenced. You can think of the typical linker algorithm as follows: If I am missing a symbol X scan through all not used objects in each library until you find it, then include that .o file. Now check if any symbols are missing again and repeat until done.
- colanderman 8y ago> Now check if any symbols are missing again and repeat until done. That properly describes linking using --start-group/--end-group. Without those flags, the process is closer to "look in the first .a file for definitions of all currently undefined symbols; then look in the next .a file for definitions of remaining undefined symbols; etc.". The difference becomes apparent if you link chains of libraries in the wrong order, or if you have cyclic dependencies between libraries; normally they will not be resolved unless you use the grouping flags. (But really, you should avoid making such cycles in the first place!)
- 8y ago
- deleted 8y ago[deleted]
- aidenn0 8y agoIf anybody is curious about how templates work, there are many ways, but the two most historically popular are: Prelinker: 1. Compile each file, noting which templates are needed in a section in the object file 2. Have a special program called the "prelinker" run before the linker that reads each .o file, and then somehow instantiates each template (which usually, but not always requires reparsing the C++ file) Weak Symbols: 1. When compiling instantiate every needed template, but mark it somehow in the object file as being weak, so that the linker only pulls in one instantiation for each definition. The prelinker used to be more popular, as if you e.g. instantiate the same template in every single file, your compiler does tons more work with the weak symbols approach, but now weak symbols are popular both because they are much simpler to implement and the fact that compilation is usually parallel, while linker is typically not means that walk clock times may even be faster.
- int_19h 8y agoWho actually used the prelinking implementation, aside from those very few compilers that valiantly tried to support "export template"?
- Quekid5 8y agoNobody, AFAIK. Even with an explicit "export template", it's basically impossible because of of the interaction of all the features of the C and C++ and CPP parts of the language. (Precompiled headers are/were are thing, but they're very brittle. Of course, one assumes you already know this, I'm trying to provide additional exposition.) People started to use templates for metaprogramming and from that point on the scope for "reuse" of templates isn't really there. (Reusing parsing might be plausible, but it's really difficult because parsing is extremely context-sensitive because of SFINAE, #defines, etc.) Some might comment that "modules" is "export template" all over again, but this time there are actually 2-3 implementations of 2-3 of the proposals and everyone is confident that the remaining minor problems can be resolved satisfactorily... and they're all exchanging experiences to help each other!
- beached_whale 8y ago
- tiagoma 8y agoWatch this: https://www.youtube.com/watch?v=dOfucXtyEsU https://www.youtube.com/watch?v=dOfucXtyEsU
- wahern 8y ago> You could put a function's definition in every source file that needs it but that's a terrible idea since the definition has to be the same everywhere if you want anything to work. Instead of having the same definition everywhere, we put the definition in a common file and include it where it is necessary. This common file is what we known as a header. In C and C++ parlance "definition" should be "declaration" and "implementation" should be "definition".[1] The terminology is important if you don't want to get confused when learning more about C and C++. This is compounded by the fact that some languages describe these roles in the author's original terms. (Perhaps the author's terminology reflects his own confusion in this regard?) [1] This is indisputable given the surrounding context, but I didn't want to paste 3-4 whole paragraphs.
- deleted 8y ago[deleted]
- kccqzy 8y agoYes, indeed. In extremely old code, I sometimes see people preferring to manually write declarations of functions (even libc functions) in every source file instead of including a header. To add to the confusion, in certain cases, it is possible to put a function's definition in a header file (for example if it's a function template, or in an anonymous namespace, or the static keyword is used to indicate internal linkage). So it is possibly to write this function definition manually in every translation unit. Otherwise, the ODR rule requires functions to be defined exactly once. > One and only one definition of every non-inline function or variable that is odr-used is required to appear in the entire program (including any standard and user-defined libraries). The compiler is not required to diagnose this violation, but the behavior of the program that violates it is undefined.
- quietbritishjim 8y agoInterestingly a form the "one definition rule" even applies to some of those functions that can appear in multiple translation units, specifically inline functions and template functions (not static functions or those in an unnamed namespace). In those case it says that they must be identical in all the translation units that they're defined in, so it's more like a "unique definition rule" for them. This sounds like it would be easy – just put the definition a header file. But even if the text of a function is identical in different translation units, it can still be ODR-different between them if the symbols that they look up are different due to other header files included before them declaring different thing or doing things like "using namespace std". Argument-dependent lookup is especially dangerous here. As with other ODR violations, this causes undefined behaviour and compilers/linkers aren't required to issue a diagnostic (and they usually don't!). I believe C++20 modules will solve this problem.
- roel_v 8y ago"Basically, the compiler has a state which can be modified by these directives. Since every .c file is treated independently, every .c file that is being compiled has its own state. The headers that are included modify that file's state. The pre-processor works at a string level and replaces the tags in the source file by the result of basic functions based on the state of the compiler." I almost sort of get what the author means here but then I don't really. I mean, there is no 'state' for the compiler that is modified by precompiler directives, so this is probably an analogy or simplification he's making here, but I don't really understand how he gets to the mental image of 'compiler state'. Why not just say it like it is: the preprocessor generates long-assed .i (or whatever) 'files' in memory before the actual compiler compiles them, the content of which can be different between compilation units, because preprocessor preconditions might vary between compilation units?
- mbel 8y agoIt's neither an analogy nor simplification. When he speaks about "compiler state" he means "preprocessor symbol table state". When preprocessor processes a file its state is mutated -- symbols get defined, redefined or undefined. What you propose as a replacement (an in-memory file) does not provide any insight into why the same file preprocessed twice may end up looking different or why order of included files matter.
- roel_v 8y agoWell in that case, I guess it's a definition thing. When I teach C++, I find it much more useful to make a clear separation between 'preprocessor' and 'compiler', and not make the preprocessor part of the compiler and then make the... uh... 'actual compiler' also part of the compiler. When you take the preprocessed state of a compilation unit, by having the preprocessor write it out to disk, and show someone what the effects are of passing one or the other -D flag, or change the order of includes - that directly and concretely shows what is going on. And then this preprocessed file is passed on to the actual compiler. There is a clear separation between stages, easy to understand, and useful to boot when the time comes you have to debug an issue related to it and you want to look at the preprocessed file to see what's going.
- marco_craveiro 8y agoBit of a side-question, but somewhat related. Is anyone working on "whole program compilation"? I don't mean whole program optimisation, I mean an attempt to read all files for a given target in memory at the same time and then generate all translation units in one go (all in memory? and maybe linking them in memory too?). Clearly, there would be caveats (strange header inclusion techniques relying on macros to modify text of include files would break, gigantic use of memory and so forth), but for those willing to take the risk, presumably this should result in faster builds right? In fact, ideally you'd even generate all binaries for a project in one go but that may be taking it a step too far :-) At any rate, I searched google scholar for any experience reports on this and found nothing. It must have a more technical name I guess...
- Khoth 8y agoThe terms you're looking for are "single compilation unit" or "unity build". It's used sometimes, I think mostly to help the compiler optimise better. Build times for a full rebuild may be faster, but may not, since traditional builds can use many CPU cores. However, it stops incremental builds from working - if you modify one source file, you have to recompile everything.
- eequah9L 8y agoOr a compile server / incremental compilation. Tom Tromey worked on supporting this in GCC years ago, and blogged about the roadblocks that he met on the way. I don't remember the details, but eventually the project was abandoned. It might still be interesting to read through this stuff--throw "tromey gcc compile server" at a search engine and see what comes up.
- deleted 8y ago[deleted]
- archi42 8y agoWe do "kind of that": Our source is split across various cpp files to keep it well organized. Our build scripts then generates a single cpp file with ~30 #includes. Compilation of the 11Mb object file takes about 30s and requires at most 1GB memory (the lib is not that huge, 20kloc), but I think when we started on it, the build took a few minutes for the lib alone. So we save some dev time, but building all unit test executables still takes an additional 1m30s. So that's only a minor improvement. But I think the real gain is a much better optimization (the architecture of the lib is great to maintain and bugs are at least critical or might even have a lethal impact; there is a lot of potential for inlining/LTO).
- leni536 8y agoI just watched Matt Godbolt's recent talk about the linking process[1]. It's a pretty good talk. [1] https://www.youtube.com/watch?v=dOfucXtyEsU https://www.youtube.com/watch?v=dOfucXtyEsU
- Const-me 8y agoI don’t think that article’s accurate. At least not anymore. Modern C++ compilers do less while compiling, and much more while linking. This allows them to inline more stuff and apply some other optimizations. VC++ calls that thing “Link-time Code Generation”: https://docs.microsoft.com/en-us/cpp/build/reference/ltcg-link-time-code-generation?view=vs-2017 https://docs.microsoft.com/en-us/cpp/build/reference/ltcg-li... LLVM calls it “Link Time Optimization”, pretty similar: http://llvm.org/docs/LinkTimeOptimization.html http://llvm.org/docs/LinkTimeOptimization.html
- gpderetta 8y agoLTO is still an opt in thing though. I suspect that most pronects still don't use it.
- IloveHN84 8y agoI hope someday the build and linking process could be standardized, but I don't believe it will happen, because many members of the committee come from Microsoft, Google and other tech giant who want to sell their compiler (or give it for free, but still). There are too much Interests and the standardization would kill many of them
- crumbshot 8y agoThe section describing how a function call is made appears to be slightly incorrect. The return value of the `add` function, in most ABI definitions, would be stored in a register. After that, the `main` function may then copy that value to its own space it has reserved on the stack. This is at odds with the description in the article, which seems to describe `add` passing its return value to `main` via the stack. (This is assuming no optimizations - all this would most likely be inlined anyway, with no function call.)