8 ms·
Reproducible builds for Debian: a big step forward
- wmanley 5y agoOn the subject of reproducible debian-based environments I wrote apt2ostree[1]. It applies the cargo/npm lockfile idea to debian rootfs images. From a list of packages we perform dependency resolution and generate a "lockfile" that contains the complete list of all packages, their versions and their SHAs. You can commit this lockfile to git. You can then install Debian or Ubuntu into a chroot just based on this lockfile and end up with a functionally reproducible result. It won't be completely byte identical as your SSH keys, machine-id, etc. will be different between installations, but you'll always end up with the same packages and package versions installed for a given lockfile. This has saved us on a few occasions where an apt upgrade had broken the workflow of some of our customers. We could see exactly which package versions changed in git history and roll-back the problematic package before working on fixing it properly. This is vastly better than the traditional `RUN apt-get install -y blah blah` you see in `Dockerfile`s. You know exactly what was installed before an update and exactly what is installed after and you can rebuild old versions. IMO it's also more convenient than debootstrap as you don't need to worry about gpg keys, etc. when building the image. Dependency resolution and gpg key stuff is done at lockfile generation time, so the installation process can be much simpler. In theory it could be made such that only dpkg is required to do the install, rather than the whole of apt, but that's by-the-by. apt2ostree itself is probably not interesting to most people as it depends on ostree and ninja but I think the lockfile concept as applied to debian repos could be of much broader interest. [1]: https://github.com/stb-tester/apt2ostree#lockfiles https://github.com/stb-tester/apt2ostree#lockfiles [2]: https://ostreedev.github.io/ostree/ https://ostreedev.github.io/ostree/
- ztcfegzgf 5y agoan article arguing that reproducible builds are a lot of effort, and the benefits are not that big: https://blog.cmpxchg8b.com/2020/07/you-dont-need-reproducible-builds.html https://blog.cmpxchg8b.com/2020/07/you-dont-need-reproducibl...
- Foxboron 5y agoThis blogpost can be summarized as with an XKCD essentially. It positions itself with the following assertions: > Q. If a user has chosen to trust a platform where all binaries must be codesigned by the vendor, but doesn’t trust the vendor, then reproducible builds allow them to verify the vendor isn’t malicious. > I think this is a fantasy threat model. If the user does discover the vendor was malicious, what are they supposed to do? The malicious vendor can simply refuse to provide them with signed security updates instead, so this threat model doesn’t work. Which only works in the context of proprietary vendors and not in the context of FOSS distributions. Nothing can be denied as everything is freely distributed. You want to have the ability to verify the work done by packagers and build servers. Next up is the essentially the claim that "reproducible builds can't solve bugdoors. Thus it's insufficient to solve any problems". But this is essentially just an XKCD argument; https://xkcd.com/2368/ https://xkcd.com/2368/ Reproducible Builds is a nice property of any build system for multiple reasons. It's also part of the supply-chain security story and not the entire story alone. As for how much effort it is? It's a lot. But considering the core community of reproducible builds people is below 50 people, and we are still able to come close to an 88% reproducible builds in real world distributions should point out how achievable this goal is. https://reproducible.archlinux.org/ https://reproducible.archlinux.org/
- dane-pgp 5y agoAs another data point for achievability: "bookworm [the next Debian stable release] on amd64 is 95.6% reproducible right now! " https://isdebianreproducibleyet.com/ https://isdebianreproducibleyet.com/
- Foxboron 5y agoSadly, they are not "real" numbers. This is taken from the integration suite which Debian have been running for years. This represents checking out the code, and building twice. This is not distributed packages from Debian. This inflates the number a little bit. Holger explains this in a thread a few years back. https://lists.debian.org/debian-devel/2019/03/msg00017.html https://lists.debian.org/debian-devel/2019/03/msg00017.html
- e12e 5y agoNot to detract from the article, but given: > It took 3-4 months to get 4.2 TB of data Maybe send an email and ask someone to mail a couple of hard drives? (offering to reimburse hw, labour and shipping obviously).
- signa11 5y agocan someone please elucidate the _benefits_ of reproducible builds ? perhaps i am missing something trivial? thank you kindly!
- passionforfruit 5y agohttps://reproducible-builds.org/docs/buy-in/ https://reproducible-builds.org/docs/buy-in/
- rocqua 5y agoIt ensures two important things. The first is that it makes sure any maintainers who try to backdoor a package need to document that fact somewhere in the build process. Hence it becomes possible to get more confidence from an audit of a package. This first goal is done very simply. You rebuild the package, and check that the resulting binary is bit-for-bit identical. With reproducible builds, the same build process should lead to the same binary. So if someone tampered with the source code or the build process, that can be detected if you have reproducible builds. The second goal is to just have a consistent system which makes debugging easier. A nice example of the importance of reproducible builds is the app signal, which I mention because of your user-name. We trust the source code of signal, but how are we supposed to be sure that the binaries offered by apt, the google play store, the app store, or whatever source you have for signal, are actually made from the source code? With reproducible builds, one person can rebuild the binary and check to see that the output is the same. This is much better than telling everyone to build signal from source, because now the lazy people who trust the binaries get to benefit from the skeptical people who build from source.
- brnt 5y agoI would turn that around: if your builds are irreproducable, how would you troubleshoot build errors?
- goodpoint 5y agoPeople already listed: - security against backdoors - validating build environment - reproducing bugs But there's more: - legal liabilities: many bad build systems pull trees of dependencies from the Internet during the build. How can you prove that no license breaching occurred in any of the dependencies? Debian explicitly verifies licensing in each package. - prevent pulling dependencies from deleted (or hijacked!) repositories on the Internet - reproduce performance improvements: non reproducible builds can often lead to non reproducible performance. Even things like the length of the local hostname can leak into a binary and affect memory alignment.
- meirelles 5y agoWe all agree about how much important debian is for all tech community. Why is it so difficult to scale their snapshot service? Serving static files at scale is a solved problem. Am I missing something? Can't a cloud provider help them out?
- lrem 5y agoIsn't Debian opinionated about the freedom of the stack it stands on? Would the community be happy to build a dependency on a vendor?
- goodpoint 5y agoDebian is opinionated about software freedom - thankfully. But it's OK to accept donations in hardware or money as long as there are no strings attached.
- BiteCode_dev 5y agoIt's very, very, expensive. I don't know the details for debian, but pypi.org, hosting python packages, costs $800k/month: https://twitter.com/dstufft/status/1236331765846990848 https://twitter.com/dstufft/status/1236331765846990848 I imaging debian is also super expensive, and so scaling that must not be easy. Every decision could be thousands of dollars.
- est31 5y agoIs that with at least some attempt at building a CDN? Generally cloud providers don't charge for traffic between hosts in the same availability zone. One could think about putting a slave into each major availability zone of various cloud providers. The main service would then only be used to create HTTP redirects to the specific slave, or if none exists, or the package isn't replicated yet at the slave, just answer directly. Even if such a system isn't built, with that kind of money on the table you could get a team of FANG scale developers to build it for you.
- rightbyte 5y agoThis lazy constant pulling of dependencies by CI systems and containers is not very substainable. pypi should set up limits and make people use some cache proxy.