8 ms·
The magic of dependency resolution
- OscarCunningham 3y agoIdeally, one would also prefer higher versions of each package. How much longer would it take to find a solution with the guarantee that there's no other solution with all packages at least as good?
- droelf 3y agoUnfortunately, it is sometimes not even possible to say which solution should be preferred because different branches lead to different sets of packages having higher versions. There is not really an "ultimate" heuristic to determine which of the two sets is better. However, users are always encouraged to add more constraints to shape the solution, and packagers encouraged to take good care of the metadata. At prefix we are quite interested in figuring out how we can help packagers to determine e.g. compatibility ranges more correctly by doing static analysis and such things.
- OscarCunningham 3y agoRight, but we could still ask for Pareto optimality: no way to make one package better, while keeping the rest at least as good.
- crabbone 3y agoIdeally, there should be a choice that allowed to look for the minimal and the maximal versions that satisfy given requirement. Developers usually want the newest, but maintainers usually want the oldest.
- JohnFen 3y ago> Ideally, one would also prefer higher versions of each package Ideally, yes. In practice? I know that I rather strongly prefer not to use the latest version of any piece of software.
- wofo 3y agoAuthor here. I glossed over this in the article, but the solver is actually privileging higher versions as you suggest. In the "set" step, when we arbitrarily set a variable to true, we have a heuristic that picks the highest possible version of the package.
- javajosh 3y agoProps to the author for thanking their tester for breaking their stuff until it worked! Good testers are worth their weight in gold.
- JohnFen 3y agoI'll second this. Good testers are at least as valuable and essential as good devs.
- crabbone 3y agoOh, this is about Conda?.. I think, this is going to be the third go at the problem? The first two were hands-down awful leading to creation of Mamba. But speeding things up is only part of the solution here. Conda package ecosystem has a bunch of issues that are related to the format and the repositories providing packages as well as on the policy for accepting packages into those repositories. They tried to solve too many problems at once, probably, w/o even realizing what kind of problem they are attacking. Here's a major problem with Conda ecosystem, the way I see it: package developers are encouraged to produce very "precise" dependency requirements because this makes installation process go faster and ensures less wiggle room for untested dependency combination. This, in turn, results in proliferation of package versions, all of which declare themselves as incompatible and non-interchangeable with many others. On top of this, beside providing just the Python packages, Conda, ambitiously, wants to also provide shared libraries, which blows up the minimum installation to an enormous size, including the number of packages that need installing. So, being a Conda package maintainer, you will face a dilemma: do you want your users to be able to install the package in acceptable time (at least under 10 minutes...) or work towards compatibility with multiple different versions of other packages the users might need for other reasons in the same environment (and your installation times start shooting through the roof). In my experience, the longest successful Conda install took about a day (i.e. close to 24 hours). It may take few days for Conda to fail to install (due to imaginary or real conflict). Now, add to this the existence of multiple sources for different packages (i.e. "channels"). Different channels may, and actually quite often do provide different unrelated other packages in their channel, but the users cannot specify the channel when requesting to install a package on per-package basis. So, the solver needs to guess somehow which channel to prefer as a source of any given package. The problem is that while the package version found in that channel might be the same, the requirements for that package might differ between channels. So, while aiming at providing a lot of possibly useful features, Conda programmers programmed themselves into a corner with no good way out. Several things might help, but none will ultimately solve the situation. * Package bundles (or meta-packages, like they are called in some Linux distributions) will cut significantly on the work a solver needs to do. A lot of the basic (shared libraries) packages are, essentially, always the same in each install. * LTS curated channel snapshots. I.e. provide a very large collection of packages that are verified to all work together, but unlike the meta-packages need not be installed together. * Incorporate channel into version specification, especially when specifying package dependencies so that maintainers were required to specify the source of their dependencies when submitting their package. * There are a bunch of (very) infrequently used options in the current solver that can be safely removed, but they add to the complexity and the time the solver spends in various "impossible" scenarios.
- Dudester230602 3y ago"Dependency resolution is something programmers usually take for granted. Be it cargo, npm, or whatever package manager you use, no one is actually surprised when this black-box figures out, all by itself, the specific set of packages that should be installed." I am personally still surprised npm/node etc. both work and are used unsarcastically.
- Gunax 3y agoNever thought of dependencies as a SAT issue before... Makes sense once I think about it. I always imagined dependencies as a tree. If I have dependency x which uses foo 1.0 and dependency y which uses foo 2.0, isn't it possible to get an unsatisfiable rule? Why do we need a universal version of foo? Cant we just use both foo 1.0 and 2.0? When being used by x, we use foo1.0 and when being used by y, use foo 2.0
- duped 3y agoDependencies are a graph, not a tree. Say z depends on x and y, which version of foo do you use? Requiring a single version makes certain things a lot easier to deal with, because not all dependencies are actually inherited by another. If foo is an executable then sure, you can have multiple versions and it's no issue. If foo is a library then you might need to fail because there is no solution that works. Or you might not, it depends.
- droelf 3y agoNPM actually works like this afaik (you can have multiple versions of a given package in tree). Vendoring is another way to work around the single-version per name requirement. It's mainly problematic with shared types and when they "escape" over the API boundary (ie. are exposed somehow).
- karatinversion 3y ago> Why do we need a universal version of foo? Cant we just use both foo 1.0 and 2.0? When being used by x, we use foo1.0 and when being used by y, use foo 2.0 You need language-level support for your packaging for this, so that the compiler or runtime can bind a call to foo.bar() to foo 1.0 when the call happens in x, and foo 2.0 when it happens in y.
- wofo 3y agoAuthor here. As others have commented, some package managers allow multiple versions of the same dependency. Conda doesn't, and I wanted to stick to the Conda rules in the article ;)
- menaerus 3y agoDoes anyone know how Linux system package managers solve this problem, SAT or something more? It's also more difficult problem because of having to maintain ABI compatibility.
- PhilipRoman 3y agoThe Arch Way of doing things involves only ever providing a single version of each package.
- chriswarbo 3y agoBut which version should they provide? The dependency resolution still needs to happen somewhere ;)
- PhilipRoman 3y agoThe latest one obviously; any issues are left to the user as an exercise
- menaerus 3y agoThat's quite surprising but I haven't delved into this topic basically never. Is it true that yum (dnf), apt and zypper took the approach by having SAT solvers under the hood? If so, why does arch take a different route?
- rcxdude 3y agoBecause it's easier for the developers. Arch's philosophy very much based on reducing the workload for maintainers. So they package things as close to vanilla as possible (and expect the users to make any adjustments, like enabling any services), package only one version, and have a rolling release which doesn't support older versions or even partial upgrades, so maintainers only really have to worry about making things work with the current version. They were one of the first distros to go exclusively systemd for the same reason of it allowed them to ditch their current init scripts and supporting both is a huge burden.
- 3cats-in-a-coat 3y agoI think the need for a Magic Solver suggests weaknesses in how we package and load dependencies. Every piece of running code should have its own space for dependencies. If most of them happen to coincide, then good: we can save a lot of RAM. But if not, also good.
- epage 3y agoThe problem is when you are dealing with dependencies that are "vocabulary terms". These are required to be common. In Rust, one example would be serde and I'm assuming pydantic would similary apply. For Rust, these can also impact build times so people actively work to keep down what dependencies can be duplicated which are semver incompatible versions.
- droelf 3y agoYeah, and what happens if you pass a NumPy array around and the different NumPy versions cannot agree on internal functions or memory layout ...
- menaerus 3y agoI don't understand the purpose of package managers in general during the development where you have the total ownership of your code, 3rd party dependencies and whole build process. I had a privilege to use conan on semi-large codebase - what a clusterfuck that is. I can't imagine a single scenarios why I would be using it.
- 3cats-in-a-coat 3y agoIndeed but this comes back to opting to write spaghetti apps by default. For example we're talking in English right here, and it's a rich language we can express a whole lot in. It has no version although it's slowly updated over time. But it's mostly the same version around the world, with some local differences. We still understand each other. Despite our minds may be vastly different, we can express each other in this language. Same goes even for language models now. Imagine we had something as common as JSON, but more in-depth so it can be used as direct representation in memory, and this was the lingua franca not only between servers, but between libraries, apps, and even objects within an app. Then our dependency entanglement issue would also not exist. What we did was a choice. It was a lazy choice, ignorant choice. A choice that introduces instant tech debt. And dependency conflicts are one of the results of this choice.
- enriquto 3y agoUnpopular opinion: code should have no dependencies. Thus, tools that ease the proliferation of dependencies encourage bad practices and must be avoided. If you want to use somebody's else code, copy it verbatim into your project. Avoid unstable third-party code that needs to be updated every year (or worse, every week).
- eschneider 3y agoAlas, sometimes that unstable third-party code is from some guy two desks down, and you need to use it. And it's unstable because it's not-quite-done yet. Dependencies are life. Sure, eliminate the ones you don't need, but on most larger(ish) projects, you're going to have dependencies, be them internal or external.
- gadflyinyoureye 3y agoI kind of agree with you. This is why I’m trying to move my team to Go for new development rather than TypeScript. I end up having to spend about one day per month fixing dependencies. Our versions are pinned but upstream someone just said, “use latest”. Then everything breaks and I have to upgrade our project from TypeScript 3 to 4 or 5. Then I have to figure out interdependencies. It’s a mess.
- logophobia 3y agoIf your dependency has a security update, how are you going to get that if you copy-paste the code? The thing these dependency managers do well is that they notify you about these types of issues. .. That said, people need to be very careful about what they add as dependency. Having 1000+ transitive dependencies is just asking for security issues.
- magicalhippo 3y agoWe effectively do the copy-paste. Dependencies are manually committed as a subdirectory to a "deps" directory, and the build scripts updated to search for it in that subdirectory. To update, we simply download the new version we want, replace the code in the subdirectory, do a build, make some tweaks if needed to build scripts and code, and run tests. Once it builds and tests are fine, we commit both the updated dependency and the other changes in one go. To get notified we use email (ie subscribe to updates) or just manually keep track. A nice side-effect of this is that it's trivial to study the changes between the old and the new version before committing. While subtle subterfuge would be hard to spot, blatant stuff like including a bitcoin miner or whatever is trivial to catch.