8 ms·
Mtime comparison considered harmful
- esotericn 8y agoWhy not just use a complete checksum of the file? The average project is what, a few MB? Less? Most of which is going to be cached after the first compile anyway. Even on an enormous codebase, as long as you have an SSD or a nontrivial amount of RAM I can't see this being an issue. You don't care if a file is newer - you care if it's different!
- ams6110 8y agoThis is discussed near the end of the article. It works best when the file system itself stores a checksum in its metadata so it does not have to be calculated for each file for each build. It's not appropriate when your build may include dependencies based on other side effects besides file content. For example, sometimes you depend on the timestamp of an empty file, or the success or failure of another step based on e.g. a log message, to trigger other actions
- esotericn 8y agoIs it common to build over NFS? I can't immediately see a use case - collaborative editing or something? Even in that case, wouldn't it be easier to build on the box? The other case seems valid in a sort of 'if your build intentionally makes use of mtime, you'll need to look at mtime'. It seems like an odd thing to do in the first place - I guess Makefile as deployment rather than for building? A 470K LoC project I have here with >1000 files takes 0.04 seconds to do a full sha256sum traversal on my box from the cache. That's single-threaded. If I drop caches, it takes approximately 1 second (from spinning rust, not SSD).
- erik_seaberg 8y agoNFS is getting so rare that some systems aren't even organized to accommodate "mount -o ro /usr" anymore.
- esotericn 8y agoI use NFS a decent amount, just not for anything like this, because even on a link with e.g. 5ms latency you end up with issues all over the place. It seems like solving a problem that could be fixed more easily by just rsyncing or cloning the codebase. Storage is cheap.
- ams6110 8y agoNFS is widely used in HPC to mount user home directories on compute nodes.
- __float 8y agoAt work, we frequently build in a Linux VM (using VirtualBox through Vagrant) from a macOS host, and the default shared folder does not support symlinks. We use NFS as a workaround.
- livueta 8y agoI do a lot of nfs builds at work as part of development on a proprietary OS. Some things will only nicely build on-OS, which for dev purposes is generally running in a non-local VM. The relevant git repos are massive enough (I'll frequently be building a small part of the tree, but still have to clone the whole thing), and the VMs disposable enough, that dealing with slower builds via an nfs mount from my workstation and/or homedir server is faster than repeatedly cloning. You can probably argue that this is a consequence of bad tooling rather than any strength of nfs builds, but it is an example of a non-trivial number of developers frequently building over nfs.
- koala_man 8y agoFor the median project, you could just rebuild everything from scratch every time and avoid the problem. It's the huge projects with millions of files and tens of gigabytes of source and assets that need these optimizations the most, and that's also where checksumming is the most painful. It's not as unrealistic or monstrous as it sounds. It happens in monorepos when you include all of a project's thousand dependencies (down to things like openssl and libpng).
- floatboth 8y agohm! i wonder if ZFS exposes the hash to the outside world..
- mwkaufma 8y agoCheck out the source sizes for Firefox or Unreal Engine. A project that's so small is a project that isn't putting big demands on the build system anyway, so that exception proves the rule.
- qznc 8y agoFrom this article I learned that build systems don't have the fundamental choice: mtime or checksum. Instead a better solution is mtime plus a bunch of other things. The article explains the faults of mtime and checksum clearly. This insight makes me want to try redo. One thing I dislike about redo is that it probably does not work well on Windows. Has anybody ever tried? Redo also makes me wonder: Is a build directory just a habit from using Make or is it a flaw of redo to not support that well? With "build directory" i mean the concept where the build process generates files in an extra directory which can simply be deleted and nothing does pollute the source directory.
- ams6110 8y agoI think all of the authors observations are valid, but in real life I've never encountered any of them. I don't have builds triggered automatically from I notify or the like, and I'm not so fast with my fingers that it takes me less than a second between saving a file and kicking off a make in another terminal window. And anytime I do an initial check out from any version control, "make clean" is always the first step.
- thristian 8y ago> One thing I dislike about redo is that it probably does not work well on Windows. Has anybody ever tried? Yes, although only for a toy project. I'm sure it'd work fine with Cygwin or WSL, it might work with MSYS, but I've definitely had it working with busybox-w32[1]. > Is a build directory just a habit from using Make or is it a flaw of redo to not support that well? You can use redo in a Make-like fashion by putting a single `default.do` in the root of your project that decides what to do by examining the filename it's been asked to build. That does give up some of the benefits of redo, however (since a single file builds everything, when you edit that file redo wants to rebuild everything). Having a separate build directory that can be easily wiped is a good idea, but I'm a lot less worried about it (or things like 'make clean') now that I have 'git clean -dxf'. [1]: https://frippery.org/busybox/ https://frippery.org/busybox/
- JdeBP 8y agoDaniel J. Bernstein, who originally designed redo, also came up with the slashpackage idea. In that system, the build directory starts out as a writable copy of the (potentially read-only) source directory, complete with all of the .do files in the case of systems built with redo. * http://jdebp.eu./FGA/slashpackage.html http://jdebp.eu./FGA/slashpackage.html
- contras1970 8y agohttp://jdebp.eu/FGA/introduction-to-redo.html http://jdebp.eu/FGA/introduction-to-redo.html hoping jdebp will chime in..
- qznc 8y agoI have never before seen this idea to put CXXFLAGS into a file and treat it as a file dependency. That would also work with Make. Clever idea. Nevertheless, a build system which does this implicitly is better, imho.
- apenwarr 8y agoPerhaps unsurprisingly, djb seems to have pioneered this with his Makefiles, which generally produce, then depend on and run, a 'compile' script that contains the flags.
- peff 8y ago> the .git/index file, which uses mmap, is synced incorrectly by file sync tools relying on mtime This part implies that the index file is written via mmap, but that's not true. It is fully rewritten to a new tempfile/lockfile, and then atomically renamed into place. Git does not ever mmap with anything but PROT_READ, because not all supported platforms can do writes (in particular, the compat fallback just pread()s into a heap buffer).
- gonzo 8y ago> Purists make me sad. People who don’t understand Unix make me sad.
- chungy 8y agoThe "popular misconceptions" section seems to have a couple of the author's own misconceptions. * On precision, he notes "almost no filesystems provide that kind of precision" (nanoseconds), but I would honestly say the exact opposite statement. ext4, xfs, btrfs, ZFS are some of the very common file systems that support this. He cites that his ext4 system only has 10ms granularity, which is most certainly not the default, but likely a result of upgrading from ext2/ext3 to ext4. As an aside, NTFS only has a granularity of 100ns. * It is unclear what he means by "If your system clock jumps from one time to another...". If this is talking about NTP, it's probably accurate. My first reading was "daylight saving" or time zone changes, in which case, everyone uses UTC internally and such changes don't affect the actual mtimes. (You might get strange cases where a file listing regards a file modified at 01:45 to be older than a file modified at 01:20, but if you display in UTC, you can see it's just DST nonsense)
- mahkoh 8y agoHe cites that his ext4 system only has 10ms granularity, which is most certainly not the default, but likely a result of upgrading from ext2/ext3 to ext4. What is the granularity of your file system? Mine appears to be 3.33ms.
- chungy 8y agoNanosecond granularity on all the ones I use, which includes ZFS, ext4, and tmpfs.
- jlokier 8y agoCheck the timestamps which are actually set on the files. I was surprised and disappointed to find Linux sets mtime to the nearest clock tick (250Hz on my laptop) on filesystems whose documentation says they provide nanosecond timestamps. It's not obvious because the numbers actually stored still have 9 random looking digits. But the chosen mtime values actually go up only on clock ticks. If you're running those filesystems on Linux, try it yourself: (n=1; while [[ $n -le 10000 ]]; do > test$n; n=$((n+1)); done) ls -ltr --full-time You should see the timestamp nanoseconds increment every few files, in batches. If they were truly nanosecond accurate mtimes, they would be different for every file. That's why some of my programs on Linux now set the mtime explicitly with a call to clock_gettime() followed by futimens(), after writing the file. To make sure the timestamps do change each times files are replaced, in case it's more than once inside a 250Hz tick.
- oever 8y agoThe article suggests writing explicit rules to check for changes in the toolchain. These dependencies can be recorded automatically with LD_PRELOAD. LD_PRELOAD can redefine functions such as fopen and let you record what files are read when running e.g. cc. This makes it feasible to record the entire relevant state of a system, files and environment variables, at build time. An argument in favor of checksums is the use of build caches. Switching between branches on large codebases triggers a lot of rebuilds. With a build cache, that can be avoided. SCons is a build system that uses such a cache.
- jake_the_third 8y agoDepending on LD_PRELOAD is extremely fragile and finicky. Not only can a process sidestep libc entirely by calling the `open`(2) syscall, but there are often many ways of combining function calls to achieve the same outcome. This method will also fail completely on systems that have new, previously unknown functions that are not monitored by the LD_PRELOAD solution. Worst of all, a LD_PRELOAD solution would not cover operations that are done on the behalf of the target program by external programs via IPC (think system daemons and dbus), at least not without intercepting and interpreting all io that target does. In short, it doesn't scale.
- zzo38computer 8y agoThose looks like some good ideas, because currently I do use just mtime base (for programs with multiple files; many of my programs are only one file and so don't need to deal with stuff like that).
- eecc 8y agoUgh, these "* considered harmful" blog posts... the only message I read is that someone wants to believe they're as smart as Dijkstra. /s
- saagarjha 8y ago> Random side note: on MacOS, the kernel does know all the filenames of a hardlink, because hardlinks are secretly implemented as fancy symlink-like data structures. You normally don't see any symptoms of this except that hardlinks are suspiciously slow on MacOS. But in exchange for the slowness, the kernel actually can look up all filenames of a hardlink if it wants. I think this has something to do with Aliases and finding .app files even if they move around, or something. I don’t think this is true anymore with APFS.
- evmar 8y agoIn Ninja I sorta stumbled through some of the same issues described here. I eventually realized that the interesting question is "does this output file reflect the state of all the inputs" and not anything in particular about mtimes, and that"inputs" includes not only the contents of the input files, but also the executables and command lines used to produce the output. If you squint, mtime/inode etc. behave like a weak content signature of the input. And once you have that perspective, you say "if mtime != mtime I had last time, rebuild", without caring about their relative values, and that sidesteps a lot of clock skew related issues. It does "the wrong thing" if someone intentionally pushes timestamps to a point in the past (e.g. when switching branches to an older branch) as an attempt to game such a system, but playing games with mtime is not the right approach for such a thing, totally hermetic builds are. One nice trick is that you can even capture all the "inputs" with a single checksum that combines all the files/command lines/etc., and that easily transitions between truly looking at file content or just file metadata. The one downside is that when the build system decides to rebuild something, it's hard to tell the user why -- you end up just saying "some input changed somewhere".
- chubot 8y agoIs there a description of Ninja's algorithm anywhere? I looked at the manual [1] and didn't quite see it. Does Ninja use a database like sqlite? It seems like it has to if it does something better than Make's use of mtimes. (e.g. the command line, which Make doesn't consider.) I looked at redo (linked in the article) and it uses sqlite to store the extra metadata. [1] https://ninja-build.org/manual.html https://ninja-build.org/manual.html
- evmar 8y agoNo, sorry. And I also mixed what Ninja actually does with some random observations in that comment. Ninja does use some database-like things, but they are just in a simple text/binary format. It's actually been long enough that I have forgotten the details. https://ninja-build.org/manual.html#ref_log https://ninja-build.org/manual.html#ref_log (contains a hash of the commands used) / https://ninja-build.org/manual.html#_deps https://ninja-build.org/manual.html#_deps (database-like thing with some mtimes, see https://github.com/ninja-build/ninja/blob/master/src/deps_log.h#L29 https://github.com/ninja-build/ninja/blob/master/src/deps_lo... )
- gumby 8y agoThe checksum issue could be addressed by having the compiler generate the checksums in a "sidecar" file and have the build system depend on those. Obviously this faces the "not every toolchain will support this" but you could have a switch to use checksums and continue to use the older approach by default. You would have to call cksum instead of touch in your Makefile. With a modules system (like C++20 is trying to embrace) you could in theory generate a vector of checksums (with subranges in the file) for each file and only recompile portions of the file that needed it. It's rather an accident of history that we use the granularity of a file at all.