4 ms·
Would you please publish an analysis that supports your first statement about bad design, that would be helpful and interesting! Also please consider writing s
by BodyCulture 2y ago
Would you please publish an analysis that supports your first statement about bad design, that would be helpful and interesting!
Also please consider writing some articles about typical packaging problems in Python that you come across your work.
Please also consider setting up some way of sending you money to support this interesting work!
- crabbone 2y agoIn my day job in a company I no longer work for I wrote a program that took apart a list of Wheels (that's what Python packages are called) and combined them into a single Wheel. The reason to do this was to expedite deployment (at that time pip installing from a list of wheels saved locally would take about a minute per wheel...). Later, I also wanted to write a Python packaging tool, but somewhere down the line I've lost interest (the project, while not functional, can be found here: https://gitlab.com/python-packaging-tools/cog https://gitlab.com/python-packaging-tools/cog ). This is just to set the context for where I'm coming form. So, in order to work on the projects listed above I had to study the spec. I won't dwell on all the problems I've discovered, here are just some highlights: the first time my eyebrows began to rise was when I read that Wheel despite being a binary package discourages programmers from putting binary artifacts in it... Yes, you read this correctly. Now, to elaborate on the matter: when Python interpreter loads Python source files installed in platlib (and possibly some other locations, I haven't researched this subject in-depth) it byte-compiles the sources for future "expedited" loading. At first, the byte-compiled Python code used to be stored alongside the sources, but later it migrated into \_\_pycache\_\_ directory. The unfortunate decision to byte-compile when loading rather than when installing is what back in the days created a lot of frustration for the dummy Python users who would try to install Python packages on their Linux system when running as sudo, but later unable to run their programs because Python interpreter would break trying to put byte-compiled files in a directory not owned by the user running the interpreter. And this is how --user option of pip was born. A history of kludges and bad fixes for a self-inflicted problem. So... the advise to not put the byte-compiled Python code in the Wheels is what tripped the poor unwitting Python programmers: had they put the byte-compiled files in there, the problem would've gone away, and they could happily install and run their programs in a simple and straight-forward way. Today, the number of kludges around this problem grew by a lot, and simply undoing this advise will not work, but this isn't the point. Anyways, the motivation for this advise? -- premature optimization. The authors of Wheel format decided to "save space" for programmers publishing their packages. In their mind, if someone published a binary packages, but with... sources in it instead of binaries, that package could be applicable to more than one combination of OS/architecture/Python version. A huge improvement, considering most Python packages are hundreds Kilobytes big! Not to mention that anyone who packages native libraries with Python packages has to packages them for all those combinations anyways. And that's like half of all the useful Python packages. This is what should've been done instead: copy from Java JARs. Have packages with byte-compiled code, have them used for deployment, and have source packages for those who want editor intellisence etc. This story is just a drop in a bucket of all the bad decisions made when designing the Wheel format, but for the lack of space and time I will not go further into details. Similarly, I will only touch on some problems with installation, just to give an example. Python allows specifying arbitrary URL as a package dependency. This makes auditing or even ensuring stable builds a huge issue. I.e. you might think that all the packages you are installing are coming form the PyPI index, unless you configured pip to use something else... but it's possible that some dependency will specify to load from a URL that was convenient for the author submitting the package at that time. Another problem is what happens when a package for the desired combination of OS/architecture/Python version doesn't exist in the index known to pip: in this case, instead of failing, pip will download the source archive and will try to build the package locally. This means that the users will get the version of the package the authors of the package are guaranteed to never even have run... And, unfortunately, quite often this process "succeeds" in the sense that some package is produced and installed. Infrequently, but still often enough for it to be a problem, such packages will have bugs related to API version mismatch. Some such problems may be sometimes swept under a rug. I've personally encountered a bug that resulted from this situation where some NumPy arrays were assumed to have 16 bit integers but in fact had 32 bit integers. Those arrays represented channels in ECG data (readings from electrodes attached to patient's scalp). The research was made and the paper was published before this hilarity came to light. (You may say that ECG is a borderline scam anyways, but still...) Now, to the last part: the module loading. The whole reason why Python came up with the kludge of virtual environment is due to how modules are loaded. Python source doesn't have a way of specifying package version when requesting to import a module. Therefore, if more than one version of a package is found on sys.path, there's no telling what will be loaded. What should've been done instead of virtual environments: Python modules should've only been loaded from platlib (not from the project source directory as it's often done during development). When loading modules, the package info directory should've been examined, and the dependencies from the META file parsed. These dependencies then would be kept in memory and refined every time new module loading request is made to narrow down the selection of versions that could be imported. This would allow Python to install and use multiple versions of the same package in the same Python installation without conflicts. This might not be a huge deal for developers working on their (single) projects, but it would be huge for developers packaging their projects for system use. Today many Python-based projects available on Linux are packaged each with their entire virtual environment and often even the Python interpreter and a bunch of accompanying libraries. I.e. in order to install a project that has single digits Megabytes of useful code, often hundreds of Megabytes of duplicated code are pulled into the system. As for the money part: dealing with Python problems pays the bills! :) I even sometimes get my name on scientific publications because that's often the kind of projects I have to deal with. So, I have both fame and compensation in good order (at least for now). But thank you for suggestion!