3 ms·
They use the word "Reproducible Builds" for linking the VCS commit to PyPI's repacked source code upload, that's why "reproduce the source code" sounds a little
by kpcyrd 2mo ago
They use the word "Reproducible Builds" for linking the VCS commit to PyPI's repacked source code upload, that's why "reproduce the source code" sounds a little confusing.
For projects with native bindings you can still fairly trivially solve this with SBOMs, and some Linux distributions are doing this for many years now. The compiler is a dependency that needs documenting, and they explicitly write they need means to re-create the documented environment from that SBOM.
- akoboldfrying 2mo ago> For projects with native bindings you can still fairly trivially solve this with SBOMs For bitwise identical results there are many small details that need to be taken care of, even in the pure-Python case: Ensuring that files appear in the wheel (zip) file in the same order (glob() doesn't guarantee this and you can definitely see different orders from run to run), timestamps are set to some fixed standard time. Any kind of build-time code generation or pulling in of external info like git commit hashes needs to be handled. Then once you include compiled build dependencies, there might be nondeterminism in the order that a multithreaded compiler writes object code, hidden timestamps, hidden absolute paths, and things like runtime detection of CPU version leading to different instructions being emitted on different machines.
- kpcyrd 2mo agoI'm well aware of those issues (having worked on this for many years), but I still believe the majority of python packages won't be affected by this. :) Most of these you would fix once on the relevant [build-system] and be done with it, no need to fix each individual python library.
- crabbone 1mo ago> fairly trivially solve this with SBOMs Please try building with conda-build to understand the actual problems this entails. Your solution is a fantasy because of how library linkage works. Dynamically loaded libraries need to be able to find each other at runtime. In order to do that, they need to store the location (the filesystem path) to the library they want to load. How these links are resolved is beyond the scope of Python (on Linux, this is managed by ld and ldconf). A lot of Python package maintainers don't understand this problem and don't understand how to make a portable library. Often times builds contain either an absolute path to another library, or a relative path... into nowhere, or a collection of paths... into questionable places. And there's no standard that tells the library authors how to do this the right way. As long as there's no standard, anyone can claim to have done it right and expect their users to accommodate them (which is what currently happens in Python). About your other ideas: making a compiler a dependency is... an awful idea. Unless you can require the same compiler is used for every dependency in your project, at the minimum, you will have an uncontrolled (how are you going to make sure that the right compiler is used to build a package?) zoo of compilers and their versions in your project. In the worst case, compilers will embed their signatures in the generated binary code preventing other binaries from loading such code, if they are generated with a different compiler version (kind of like what happens if you try to load Linux drivers that weren't compiled with the same compiler that compiled the kernel). * * * Like I said: these people have next to no practical experience with the problem they are trying to solve. Why can't they just go some place else and apply themselves elsewhere is beyond my comprehension.