5 ms·
Is anyone actually implementing the concept of checking hashes with trusted builders? This is all wasted effort if that isn't needed. I've seen it pointed out
by advisedwang 2y ago
Is anyone actually implementing the concept of checking hashes with trusted builders? This is all wasted effort if that isn't needed.
I've seen it pointed out (by mjg59, perhaps?) that if you have a trusted builder, why don't you just use their build? That seems to be the actual model in practice.
Reproducibility seems only to be useful if you have a pool of mostly trustworthy builders and somehow want to build a consensus out of that. Which I suppose is useful for a distributed community but does seem like a stretch for the amount of work going in to reproducible builds.
- c0balt 2y ago> is anyone actually implementing [..] Not for NixOS as far as I can tell. You only have this for source derivations where a hash is (usually in a PR) submitted and must be reproducable in CI. This specific example however has the problem that linkrot can be hard to detect unless you regularly check upstream sources.
- mschwaig 2y agoYou also couldn't feasibly do that for derivations that actually build packages, instead of fixed output derivations only, because if you the update the package set to include a newer version of the compiler, which would often produce a different output, in addition to having to rebuild everything, you would have to update all of the affected hashes. What you should be able to do in the future with a system like nix plus a few changes is use nix as a common underlying mechanism for precisely describing build steps, and then use whatever policy you like to determine who you trust. One policy can be about having an attestation for every build step, another one can be about two different builders being in agreement about the output of a specific build step. That way you can construct a policy that expresses reproducibility, and reproducibility strengthens any other verification mechanism you have, because it makes it so that you can aggregate evidence from different sources. and then have different build hosts
- c0balt 2y ago> You also couldn't feasibly do that for derivations that actually build packages, [..] you would have to update all of the affected hashes. You can actually, changes to stdenv are possible and "just" a lot of work. You will regularly see them between releases or on unstable and they cause mass rebuilds. This doesn't just affect a compiler but also all stdenv tooling as these changes tend to cause rebuilds across nixpkgs. This would be verifiable but it obviously multiples the amount of compute spent. Hint: If you look at PRs for nixpkgs you will notice labels indicating the required amounts of rebuilds, e. G., rebuild-darwin:1-10. See for example https://github.com/NixOS/nixpkgs/pull/377186 https://github.com/NixOS/nixpkgs/pull/377186 with the rebuild-darwin:5001+ label.
- mschwaig 2y agoI know about mass rebuilds, but in the parent comment you were talking about fixed output derivations, and committing the hashes for a mass rebuild to version control is technically possible, but not a reasonable workflow, because it makes all changes that are mass rebuilds conflict. What works better is keep track of those hashes as part of the signatures, which is already happening. There's a lot of interesting things that can be done with that kind of information, I'm one of the people working on that kind of stuff. Basically I have a paper out about how verifiable and reproducible can come together like that in Nix: https://dl.acm.org/doi/10.1145/3689944.3696169 https://dl.acm.org/doi/10.1145/3689944.3696169
- arccy 2y agoThe superior distro Arch Linux does it: https://reproducible.archlinux.org/ https://reproducible.archlinux.org/ maintainers build the packages, other people check: https://wiki.archlinux.org/title/Rebuilderd#Package_rebuilders https://wiki.archlinux.org/title/Rebuilderd#Package_rebuilde...
- __MatrixMan__ 2y ago> if you have a trusted builder, why don't you just use their build Pardon my tinfoil hat, but doing this would make them a high-value target. If I like them enough to trust their builds, I probably also like them enough to avoid focusing the attentions of the bad guys on them. Better would be to have a lot of trusted builders all comparing hashes... like, every NixOS user you know (and also the ones they know) so that there's nobody in particular to target.
- Timber-6539 2y agoThat's no different from how NixOS does it. You are still comparing hashes from the first build done by the distribution. A more pure approach would be to use the source code files (simple sha256sum will suffice) as the first independent variable in the chain of trust.
- __MatrixMan__ 2y agoI'm not sure what you mean. It's your machine that calculates the hashes when it encounters the code. If you bulld the directed graph made by the symlinks in the nix store, and walk it backwards, a sha256 of the source files is what you'll find, both in the form of a nix store path and possibly in a derivation that relies on a remote resource but provides a hash of that resource so we can know it's unchanged when downloaded later. The missing piece is that they're not gossipped between users. So if I find some code in a dark alley somewhere and it has a nix flake to make building it easy, I've got no way to take the hashes and determine who else has experience with the same code and can help me decide if it's trustworthy.
- mschwaig 2y agoThere are also some other gaps left to close to implement this vision, mentioned in this post an my reply to it: https://news.ycombinator.com/item?id=43030046 https://news.ycombinator.com/item?id=43030046
- __MatrixMan__ 2y ago
- whazor 2y agoThere is also an additional benefit to reproducible builds, where getting the same output every time could help avoiding certain regressions. For instance, if GitHub actions performs extensive testing on a particular executable. Then you want to be able to get the exact same executable in the future, not one that is slightly different.
- mschwaig 2y agoYes. Reproducibility also makes it possible to aggregate information about the links in dependency trees and distribute trust on that basis. That stuff is useful to humans, but it is also really useful for cold hard automated logical reasoning about dependency trees.
- sublimefire 2y ago> is useful for a distributed community but does seem like a stretch for the amount of work going in to reproducible builds Good point but even in the case of a larger monolithic systems you want to be sure it is possible to forensically analyze your source, to audit it. Once you can trust that one hash relates to this specific thing you can sign it, etc. This can then be "sold" with some added value of trust down the stream. Tracking of hashes also becomes easier once they are reproducible because they mean much more than just a "version".