9 ms·
Make's underlying design is great (it builds a DAG of dependencies, which allows for parallel walking of the graph), but there's a number of practical problems
by ejholmes 9y ago
Make's underlying design is great (it builds a DAG of dependencies, which allows for parallel walking of the graph), but there's a number of practical problems that make it a royal pain to use as a generic build system:
1. Using make in a CI system doesn't really work, because of the way it handles conditional building based on mtime. Sometimes you just don't want the condition to be based on mtime, but rather a deterministic hash, or something else entirely.
2. Make is _really_ hard to use to try to compose a large build system from small re-usable steps. If you try to break it up into multiple Makefiles, you lose all of the benefits of a single connected graph. Read the article about why recursive make is harmful: http://aegis.sourceforge.net/auug97.pdf http://aegis.sourceforge.net/auug97.pdf
3. Let's be honest, nobody really wants to learn Makefile syntax.
As a shameless plug, I built a tool similar to Make and redo, but just allows you to describe everything as a set of executables. It still builds a DAG of the dependencies, and allows you to compose massive build systems from smaller components: https://github.com/ejholmes/walk https://github.com/ejholmes/walk. You can use this to build anything your heart desires, as long as you can describe it as a graph of dependencies.
- Groxx 9y agoVery much agreed :| It's great when it's kept fairly simple (for both automating random stuff and building things! and it's installed everywhere! and has some cross-platform tools!), but it quickly turns into a nest of vicious footguns unless you're extraordinarily careful.
- ejholmes 9y agoAlso, I'm always surprised how little mention there is of graphs when we start talking about build systems. If you're making a build system (for literally anything) the DAG is your best friend. This is how all major build system tools work (make, terraform, systemd (yes, it's a build system when you think about it)) and it's how we're able to parallelize execution so easily. It's important to be conscious of the fact that this is what your doing when your making a build system; connecting a graph. Highly recommend reading https://ocw.mit.edu/courses/electrical-engineering-and-computer-science/6-042j-mathematics-for-computer-science-spring-2015/readings/MIT6_042JS15_Session17.pdf https://ocw.mit.edu/courses/electrical-engineering-and-compu... for some theory on parallel execution for graphs, if your interested in things like this.
- deleted 9y ago[deleted]
- majewsky 9y ago> Using make in a CI system doesn't really work, because of the way it handles conditional building based on mtime. Any CI I've seen starts from a fresh repo checkout and rebuilds everything every time, so it's not an issue with practice. OTOH, probably all my projects were small enough for it to not hurt when the CI builds everything from scratch every time. I might look at this from another angle if the things I worked on were mega-LOC C++ programs, and not kilo-LOC Go programs.
- regularfry 9y agoCircleCI (at least) caches build dependencies for speed, but it does it based on specific directories, not timestamps. As a result, it's not as fast as it could be because unless you munge your build to fit, it doesn't cache intermediate build products of your own code.
- adrianN 9y agoMost languages don't build as quickly as go. Having to wait tens of minutes for a clean build is not unusual.
- majewsky 9y agoThat brings back not-so-fond memories of trying to compile Servo on my notebook. :) By the way, I just noticed that compilation times are another argument for microservices.
- ejholmes 9y agoThat works, until your build system is sufficiently large or time consuming or not C/C++ like make was originally built for. For example, at my company we have a build system for building all of our Amazon Machine Images (AMIs). It doesn't make sense to re-build images unless they're dependencies have changed, but mtime just doesn't make any sense for a build system like this. Trying to coerce make into doing what we wanted was like pulling teeth.
- wruza 9y ago
- shangxiao 9y ago> Let's be honest, nobody really wants to learn Makefile syntax. I haven't found that to be the case. Once I show people how simple it is they realise make is somewhat approachable, similar to OP's article.
- wruza 9y agoMain downside, as with many tools much better than make: >walk is currently not considered stable, and there may be breaking changes before the 1.0 release. :) It also needs go, which is pretty non-lightweight dependency on windows, I guess. (And personally, I don't find these dir-hierarchy connection and bash's "case $target in a) ... b) ..." any friendlier at all, though make has some quirks with variable interpolation.) I'm not to argue to use make for very complex situations, but usual src->intermediate->executable and generate this from that of any size and count is a perfect task for make. Makefiles are unsuitable not for big projects, but for complicated build systems, where it's not enough to just do them steps in correct order. If your build system is complicated, it should be worth it at least. Otherwise use make. Fix: typos/grammar
- ejholmes 9y agowalk is built with Go, which also means it's a single statically linked binary with no dependencies if you download the pre-compiled binary. If you don't like bash syntax, you can use Python, or Ruby, or Perl, etc. Any executable can serve as a Walkfile. It's a completely fair point that make is installed on pretty much every machine by default, which is why we won't see it going away anytime soon (nor should it, it's still good).
- vog 9y ago> As a shameless plug, I built a tool similar to Make and redo Also be sure to have a look at tup, which operates vastly more efficient by simply walking the DAG in the opposite direction: http://gittup.org/tup/ http://gittup.org/tup/ That is, instead of looking what you want to build and checking all timestamps of dependency files, it can use e.g. inotify, then know exactly which files changed, and rebuilds everything that depends on those files. Moreover, it performs the modification check only once, at the beginning. During the build, it doesn't need to re-check everything, because it already knows which files were recreated.
- adgasf 9y agoGood idea, but it misses some clever tricks. What happens when a dependency does change, but not in a way that actually matters? Buck has a nice feature where Java libraries are only fully recompiled when the API of their dependencies change, since everything is dynamically linked. https://www.youtube.com/watch?v=uvNI_E0ZgZU https://www.youtube.com/watch?v=uvNI_E0ZgZU
- vog 9y ago> What happens when a dependency does change, but not in a way that actually matters I believe that tup does have some features in that direction, but I may be mistaken.
- Jtsummers 9y agotup can be made to ignore unimportant changes. I had to look up the syntax, hopefully I got it correct. I have foo.c, bar.c and a rule like: : foreach foo.c bar.c |> ^o^ gcc -c %f -o %o |> %B.o {objs} : {objs} |> gcc %f -o %o |> baz The ^o^ part tells it to not trigger the next rules if there's been no change to the output. So if you just change your source files to use an updated license, reformatted it for clarity, etc., then nothing else will happen. I used this with some of my literate code projects where I had tup running org-tangle for me (via an emacs script). If I'd only updated the documentation, and the code hadn't changed, nothing else would build. If I'd only changed unimportant parts of the code that generated the same object files, no new binary or library would be built.
- vog 9y ago> 1. Using make in a CI system doesn't really work, because of the way it handles conditional building based on mtime This is my #1 gripe with Make, and many other build systems as well. There are so many flaws in the timestamp approach. Most are easily fixed with cryptographic hashes. I like the OCaml build system OPAM for that matter, it internally just stores checksum. I believe it also uses timestamps to speed things up, but only for equality comparison (not for older/newer comparisons which may easily lead to wrong result).
- monktastic1 9y agoAny reason the hashes need to be cryptographic? Is there a security consideration I'm missing?
- Groxx 9y agoDepending on how deep down the rabbit hole you want to go, you could argue that e.g. using md5 could allow "an attacker" to submit e.g. an innocuous image that conflicts with another source file, causing it to be excluded from the build, causing a security hole to be opened. But that's kinda silly. I might argue in favor of (fast) cryptographic hash algorithms in general since they're fairly well understood / implemented / hardware accelerated / tend to have extremely "balanced" random output thus less likely to accidentally conflict... but that's about all I can think of.
- vog 9y ago> 2. Make is _really_ hard to use to try to compose a large build system ... recursive make is harmful Note that "recursive make is harmful" does not argue against multiple Makefiles! There's nothing wrong with using multiple Makefiles per-se, as long as they include rather than call each other. In other words, the article just says that Makefiles should use sub-Makefiles via "include" rather than executing those through separate (recursive) call to Make. However, I agree that composition of Makefiles is still a pain, given that the included Makefile must be aware that it is executed from another (parent/grandparent) directory.
- sorbits 9y ago> As a shameless plug, I built a tool similar to Make and redo […] https://github.com/ejholmes/walk https://github.com/ejholmes/walk Your README’s example show `parse.h` as output from `Walkfile deps parse.o`, I think that is a mistake. As for your build system (and comment), I have some questions: 1. How do you achieve using a deterministic hash as condition (and aren’t all hash functions deterministic)? 2. Why would you not be able to use mtime as a dependency? The only case I have run into is when the build depend on data from remote machines, but in that case, I think the proper solution is to have an initial “configure” step where you retrieve the data your build depend on and/or a build rule to update this data. 3. Does your build system execute the Walkfile for every node in the graph on each walk? Because that sounds like a quite noticeable overhead for larger projects. 4. Am I right in thinking that the primary advantage with your system, over make, is that a shell command is executed to obtain a target’s dependencies?
- ejholmes 9y ago1. Yeah, hash functions are deterministic, but your input needs to be determinsitic across machines too. For example, on a unix system, you may want to conditionally build if any files have changed. To do that, you could generate a deterministic hash of the dependencies with something like `find . | sort | xargs sha1sum | sha1sum | cut -d ' ' -f 1`. Including mtime in that would break across machines. 2. Mainly because doesn't translate across machines; it only applies to your machine. If someone checks out the repo on their machine, mtime is different. As soon as you move a build system to CI, you need something better than mtime, like content adressable hashes, if you intend to cash targets. 3. It executes the Walkfile twice for each target in the graph; once to determine the targets deps, and once to execute the target. This definitely hasn't been a bottle neck for anything I've used walk(1) with so far. 4. Correct! But even more so, replace "shell script" with "executable". The Walkfile can be written in any language you want, as long as it has the executable bit set.
- sorbits 9y agoThanks for the clarifications. As for #1, what I do not understand is how a `Walkfile` allows me to use a content hash change to trigger a rebuild. Your documentation says that a file list should be returned for `deps`, so how does the `Walkfile` communicate that e.g. `main.o` should be updated if `sha1sum main.c` is different than on last invocation?
- Nursie 9y agoMakefile syntax is pretty simple, once you're used to it. Multiple Makefiles works fairly simply with the include directive. The thing I've found painful before, with C projects, is getting Make to recognise that it needs to rebuild when a header has changed. There are various ways around this (makedepend etc) but they've all been quite painful to set up and not quite perfect. That said, with modern machines that have NVMe storage and massively fast processors, a complete rebuild is seldom a big time cost
- deorder 9y agoOr you can let the compiler generate the header-dependency Makefiles for you: http://make.mad-scientist.net/papers/advanced-auto-dependency-generation/#combine http://make.mad-scientist.net/papers/advanced-auto-dependenc... You already tried that? I do not find it painful to set up.
- Nursie 9y agoMaybe not using exactly that method, but I am pretty sure I have tried using gcc for it. Will have a proper read of that later. The syntax is a tad hairy though,and I'd want to adapt it - I tend to avoid compiling individual C files to objects these days, due to WHOPR optimisation.
- chubot 9y agoIf you don't care about incremental builds and want just a full rebuild, I would write a shell script instead of a Makefile. That said, I think incremental builds are important for most use cases.
- Nursie 9y agoI'd argue that the make syntax and built-in features are a huge boon over starting from plain-old-shell regardless.
- chubot 9y agoWhat are some examples of that? Shell can do basically everything make can. The syntax of both is ugly but I'll grant that make is more uniform. And btw there is no way to use make without shell, but you can use shell without make.
- ynezz 9y ago> 2. Make is _really_ hard to use to try to compose a large build system Are Android and LEDE/OpenWrt big enough? > 3. Let's be honest, nobody really wants to learn Makefile syntax. That's probably very subjective, I find JavaScript, C++, PHP or Perl much worse :-) Anyway, for the past years I'm cheating a lot and using CMake to get complex Makefiles almost for free.
- ejholmes 9y ago> That's probably very subjective, I find JavaScript, C++, PHP or Perl much worse :-) Oh definitely. I actually like Makefile's, but in my experience, more teams have a deep familiarity with some programming language, than with Makefile syntax. I haven't met very many people who have read the GNU make manual, and know all the idiosyncracies around Makefile syntax.
- 0xdeadbeefbabe 9y agoThis lovely cough line brought to you by a lede/openwrt Makefile: target/linux/ar71xx/image/Makefile: CMDLINE = $$(if $$(BOARDNAME),board=$$(BOARDNAME)) $$(if $$(MTDPARTS),mtdparts=$$(MTDPARTS)) $$(if $$(CONSOLE),console=$$(CONSOLE))
- Nursie 9y agoThat's just a bit of cmd line building. Three clauses to build a string - If BOARDNAME is defined, add board=$(BOARDNAME) to the string If MTDPARTS is defined, add mtdparts=$(MTDPARTS) to the string If CONSOLE is defined, add console=$(CONSOLE) to the string Pretty simple, if a little like the ternary operator in C. You must have seen more complex clauses than that in shell scripts and all sorts of places.
- mto 9y ago+1 for cmake. It works so well, in my latest C++ project the cmake file is more or less just a listing of source files and it does it job on windows, Mac and a few different Linux distros without any platform specific stuff. So well that I replaced the huge makefiles of the libraries I use with really small cmake files that just work and don't require modifications every 3 months because this specific distro causes problems or whatever. It feels much more like in those IDEs where you just drag all the source files in and you're done. I even prefer editing visual Studio project files by hand to many makefiles out there..
- ynezz 9y ago> 2. Make is _really_ hard to use to try to compose a large build system Are Android and LEDE/OpenWrt big enough? > 3. Let's be honest, nobody really wants to learn Makefile syntax. That's probably very subjective, I find JavaScript, C++, PHP or Perl much worse :-) Anyway, for the past years I'm cheating a lot and using CMake to get complex Makefiles almost for free.
- GlennS 9y agoThe number of times I've forgotten that unzip restores timestamps and so make has decided to rebuild everything... `unzip -D` is your friend.
- ericfrederich 9y agoccache is your friend ;-)
- GlennS 9y agoHonestly, I've never actually written a makefile to compile a C or C++ program. In fact, I haven't written any C of my own since university. My makefiles are usually just a way to record a data pipeline. Get these files, shove them through these scripts here and those programs there. Launch a web server to show the output.
- gkya 9y ago> Let's be honest, nobody really wants to learn Makefile syntax. Make's syntax is quite simple. It's a bunch of variables one has to memorise, and man page is a command away. And any decent editor would know to insert literal tabs. > Make is _really_ hard to use to try to compose a large build system from small re-usable steps. I have no experience myself, but Linux's build system is a bunch of homebrew makefiles, all the BSDs and their ports trees build with bmake. These are enough positive examples for me to think that Make is good for big systems.
- majewsky 9y ago> Make's syntax is quite simple. I'll just leave this here: https://github.com/sapcc/limes/blob/62e07b430e2019a6c189144350bae34c574b2f55/Makefile#L24-L27 https://github.com/sapcc/limes/blob/62e07b430e2019a6c1891443... (then used in line 42)
- dmitriid 9y agoI could leave the entirety of https://github.com/ninenines/erlang.mk https://github.com/ninenines/erlang.mk https://github.com/ninenines/erlang.mk/blob/master/core/compat.mk#L17 https://github.com/ninenines/erlang.mk/blob/master/core/comp... https://github.com/ninenines/erlang.mk/blob/master/core/index.mk#L20 https://github.com/ninenines/erlang.mk/blob/master/core/inde... etc.
- gkya 9y agoThat's a variable assignment. What's complex with that?
- majewsky 9y agoThe fact that I even need to use variables because the language does not have proper string literals.
- kazinator 9y ago> If you try to break it up into multiple Makefiles, you lose all of the benefits of a single connected graph Only if each Makefile is treated as a separate rule set processed with a separate make invocation. > Read the article about why recursive make is harmful That (now) venerable, old paper in fact shows how to break up into multiple makefiles (called "module.mk" in its examples) which are included into one graph. (It's possible to actually have this file be called Makefile. Not only that, but it's possible to have it so you can actually type "make" in any directory of the project where there is a Makefile, and it will correctly load and execute the rules relative to the root of the tree.)
- lomnakkus 9y agoI honestly have no idea why there's so much fandom towards Make in this thread, but for me there are a few absolutely devastating problems with Makefiles: a) Mtimes-as-change-detection is fundamentally broken given the reality of networked file systems and... physics. (Minor problem, but extremely annoying to work around in practice.) b) Nobody can actually really understand all the interdependencies between all the code in their system(s!), and yet Makefiles expect you to specify all of that explicitly?!? Yes, you could technically specify that -- and you'll want to -- but you won't, because you don't know and don't have the time. c) Make support for builds that change the structure of the build during the build is abysmal. E.g. after "processing foo.xml we now have more files than we had before! What are you going to do?". Well, in Make it's some sort of custom solution with ".d" files and "gcc -M" or whatever. This is utterly broken in that it pushes all the complexity onto the user. So, yes, an elegant model, but it doesn't actually solve the problem. If you want to see a better solution see the "Shake" paper.