5 ms·
Lots of progress for Debian's reproducible builds
- christop 12y agoThe reproducible builds talk at 31C3 also does a nice job of explaining some of the many possible attack vectors that make reproducible builds desirable, and many of the subtleties involved in making it work: http://media.ccc.de/browse/congress/2014/31c3_-_6240_-_en_-_saal_g_-_201412271400_-_reproducible_builds_-_mike_perry_-_seth_schoen_-_hans_steiner.html http://media.ccc.de/browse/congress/2014/31c3_-_6240_-_en_-_...
- walterbell 12y agoBaserock (http://wiki.baserock.org http://wiki.baserock.org) may have a repeatable build of OpenEmbedded for automotive systems.
- desdiv 12y agoWhat's the relationship between Baserock and OpenEmbedded? Glancing through OpenEmbedded's wiki page, it seems to a recipe/build system for embedded linux. Baserock seems to be in the same league.
- walterbell 12y agoIt's a distro/downstream stable version of OE, used in automotive, http://www.genivi.org/ http://www.genivi.org/. There are several OE "distros", e.g. Angstrom, http://www.angstrom-distribution.org/ http://www.angstrom-distribution.org/ . There's also Yocto, the overall build system, https://www.yoctoproject.org/ https://www.yoctoproject.org/
- contingencies 12y agoI drank with the Yocto people in 2011. They absolutely live and breathe release engineering. On subjects like repeatable builds, I trust their opinion. At present, they don't appear to mention it at all on their docs or wiki.
- csirac2 12y agoIt can be surprisingly difficult. Funnily enough moving from svn git in one project I know of probably did a lot of the necessary work to achieve this, by having to remove reliance on $SVN tags and pre/post-"build commits" which used to be a part of the release process. It's an interesting use-case for Docker as well: you can ship the build environment (or its Dockerfile describing it) for people to run builds under the same env as the official released build.
- stsp 12y agoCan you elaborate on what these pre/post-"build commits" did and why they were needed? If these commits were used to adjust version numbers in source files, the following trick should eliminate them: Check out the release branch into a working copy, adjust the version numbers in the working copy, and then copy the working copy to the tag's URL (as in: cd working-copy; svn copy . ^/tags/1.0) If this was a maven project: The maven release plugin is doing this wrong and performs 3 commits to an SVN repository per release...
- csirac2 12y agoIt was a bespoke build system but this aspect of it was similarly "wrong". IIRC pre-build commit, yes, was version-number related. It also gathered issues fixed in the release being built and updated change logs/release notes/upgrade info documentation automatically. IIRC the post-build commit helped confirm in the commit history that a particular build for release x.y.z was successful and the version number can now be incremented (occasionally it took multiple attempts for the release manager to build successfully). Of course, one could go through the tags but most of the developers liked seeing release management stuff in trunk/release branches. I didn't mean to criticize SVN as being inherently incapable of reproducible builds (if anything that's harder to achieve with git, especially with hacks like I've listed above), but the act of cleaning up the SVN repos and preparing for migration to git where a lot of our old SVN (and RCS!) habits would be problematic, also seem like the same kind of housecleaning you'd need to prepare for reproducible builds.
- DanielDent 12y agoI had the same hope for Docker, that it would ease reproducible builds. It quickly became clear to me how much work still needs to be done for that to be realistic. I maintain a packaging of Meteor using Docker designed to increase reproducibility (https://registry.hub.docker.com/u/danieldent/meteor/ https://registry.hub.docker.com/u/danieldent/meteor/). It does some checksums on the Meteor code, but the build process itself introduces entropy, and Dockerfile doesn't really have the primitives to make it easy to prevent that. Many many build processes have hidden from-the-network dependencies which are yet another huge source of problems (both in terms of having high-availability for the build process and in terms of understanding exactly what is getting built). All of the upstream Docker images (including the bottom base images which most people never look at carefully) would need to be built in a reproducible way. And the Docker code would need to actually verify checksums more carefully than it currently does. Having projects like Debian doing work like this will make it a lot easier for everything else to become reproducible. Which in addition to the security benefits is also pretty useful when bug hunting - having a system that makes it clear which (if any) of your dependencies changed narrows the list of things which need to be checked when something breaks. Ubuntu Core, Snappy, and Nix are also on my list of things to watch in this area.
- chubot 12y agoI actually did some work making debootstrap reproducible. So even if the 100 or so .deb builds it depends on are reproducible, then the chroot image resulting from debootstrap will not be reproducible byte-for-byte, due to the debootstrap shell script itself and the tools it calls. Offhand, I remember that /etc/{passwd,group} are copied from the host machine by design. There is also a random seed file, to save entropy across reboots. And there is some nondeterminism in the dynamic linker cache AFAIK. And timestamps in logs. If anyone is interested in this let me know.
- csirac2 12y agoI'm interested! I actually have had to work around the /etc/{passwd,group} shenanigans for other reasons, interested to see what you did there.
- chubot 12y agoUnfortunately I haven't published it yet, but I can describe what I did. I have like 3000 lines of shell script to make containers, and maybe 800-1000 lines are related to debootstrap. I have a cron job running daily on multiple machines, doing a deterministic debootstrap of Debian Wheezy on i386 and amd64. It basically wraps debootstrap, strips down the image a bit, and stamps out the nondeterminism. Every day it has gives a checksum for i386 and one for amd64, which holds across multiple host machines (one is Debian Wheezy, the other is Ubuntu Trusty). So it is free from host influence. The /etc/apt/sources.list just has the "wheezy" repo now (i.e. not wheezy-updates). So Debian 7.7 had one pair of checksums, and then on Jan 10 2015, I noticed Debian 7.8 was released. They changed that day, and have been stable/reproducible every day since. Part of this is also mirroring the Release/Packages metadata daily and storing version history in Git. One nice thing I found out about Debian through doing this is that the Release file completely describes the input, since it's hashes all the way down (a "Merkle tree", basically like Git itself.) My scripts also make it so you can store versioned metadata in one tree, while keeping data immutable in "pool" (all this really requires is symlinks and file:// URLs for the repo). What project were you working with in this area? I'm basically doing this to make reproducible builds of containers. I'm kind of surprised that Docker completely punts on this problem.
- sanxiyn 12y agoDebian does amazing amounts of system-wide initiatives. Off the top of my head, there are multiarch https://wiki.debian.org/Multiarch https://wiki.debian.org/Multiarch, clang rebuild http://clang.debian.net/ http://clang.debian.net/, and automated code analysis https://qa.debian.org/daca/ https://qa.debian.org/daca/.
- agumonkey 12y agoLittle bit of related trivia : Lunar (J.Bobbio) worked on hOp, a GHC based Haskell micro kernel so you can write drivers in it. See https://github.com/dls/house https://github.com/dls/house. A knowledgeable fellow.
- Alupis 12y agoCan anyone comment on why all builds are not currently "reproducible"? I mean, if a package is compiled on the same system, with the same compiler, with the same build script -- should it not produce the same output?
- agwa 12y agoThe biggest offender is timestamps - compiled binaries, documentation, archives, etc. often contain the time at which the file was built. Other problems include non-deterministic filesystem order, randomized hash algorithms, and even the fact that Markdown processors mangle email addresses randomly. A highly vexing problem that I'm currently trying to solve is that libxslt implements the XSLT generate-id() function by taking the memory address of the XML node struct, which makes documentation generated with XSLT non-reproducible (memory addresses are non-deterministic because of address space randomization and randomized hash tables). (I'm the author of strip-nondeterminism, a tool in our custom toolchain that attempts to normalize files after they're built.)
- contingencies 12y agostrip-nondeterminism "is a Perl module for stripping bits of non-deterministic information, such as timestamps and file system order, from files such as gzipped files, ZIP archives, and Jar files. It can be used as a post-processing step to make a build reproducible, when the build process itself cannot be made deterministic. It is used as part of the Reproducible Builds project." Browse source @ https://anonscm.debian.org/cgit/reproducible/strip-nondeterminism.git/tree/ https://anonscm.debian.org/cgit/reproducible/strip-nondeterm... PS. non-deterministic filesystem order .. doesn't this go away if you use tmpfs?
- jzwinck 12y agoMaybe you could fix the libxslt problem by using a pool allocator for the nodes. The pool would contain only the nodes, and you could use the offset in the pool as their ID, rather than the full virtual address. Just some food for thought.
- 12y ago
- jml7c5 12y agoWill this provide a guaranteed method for reproducible builds, or will it still be technically possible to create build scripts that produce different results (e.g., by pulling from /dev/random, or grabbing timing information from various sources, or by writing a multithreaded program whose threads all write to a single file)?
- SwellJoe 12y agoHow could anyone or anything prevent you from building something nonreproducible? This is about making build processes that are intentionally reproducible... You seem to be asking if one could continue to do things as they currently are done, which given that the entire system is open source, of course you can. This can't possibly force all users to only make repeatable builds. This seems like such an odd question that I think I must be misunderstanding you.
- deleted 12y ago[deleted]
- leonhandreke 12y agoThis is a link shared by an LWN subscriber - usually, articles only become available for free 7 days after publication. If you read this article, please think about supporting LWN financially.
- aplanas 12y agoI love this kind of projects, and I think that for Debian is one of the best things that can happens. Also openSUSE have reproducible builds/packages since ages via OBS (http://build.opensuse.org http://build.opensuse.org) and now Factory/Tumbleweed have reproducible packages + automatic CI (using openQA: https://openqa.opensuse.org https://openqa.opensuse.org) Quite an achievement for a rolling distribution.