4 ms·
As a pip maintainer, I've long been looking at the tradeoffs of uv's cache, it's the biggest item that makes warm installs faster for uv vs. pip. As pip caches
by notatallshaw 1mo ago
As a pip maintainer, I've long been looking at the tradeoffs of uv's cache, it's the biggest item that makes warm installs faster for uv vs. pip. As pip caches the original distributions and then has to unzip them each time, uv caches the unzipped distribution and hard links to it if it can.
But it has always had two major issues:
1. No way to reproduce exact distributions for a "download" command (there is no uv equivalent of "pip download")
2. For people with a lot of different environments the cache grows significantly more than pip
I'm interested to see, at least anecdotally, if this significantly improves 2, then we can perhaps have a two layer caching strategy without the significant disk space cost.
- user5994461 1mo agoAlso as a pip contributor. The only advantage of uv is to have support for parallel async extraction. If pip extracted multiple files/wheels in parallel without being blocked by the GIL, pip could easily match or outcompete uv. I'm personally not looking forward to any deduplication/hardlink in pip. Hardlinks are very dangerous and system dependent. It's very dangerous for empty files (init.py, empty.log yet not written). When the user edits one file, all files are modified simultaneously, all venv ever created by the user can be broken by editing one file, which is quite catastrophic. It's also dangerous for small files with repeated content, for example random settings files that would contain a "1" or "true". Again, when the user edits one file, all files are edited and they were supposed to be different! Hypothetically, a simple deduplication of binary files (.dll .so) should achieve 50% of the savings without significant drawbacks I'd venture to say that pip extraction is more optimized than uv in at least one way. We have optimization for empty files (0 bytes) because there is nothing to write and checksum. uv doesn't seem to have the same optimizations, though I could be wrong, I just had a cursory look and my rust is not great. uv should probably review their treatment of empty files, it's counter productive to do any file system operation open/read/write because there is no content, it might be counterproductive to use any cache/comparison/hardlink if it takes more operations than doing nothing.
- zanie 1mo ago(I work on uv) > The only advantage of uv is to have support for parallel async extraction. This isn't true, the biggest speed ups are for the warm cases where we've already unpacked the files into the cache, as notatallshaw mentions above. We can also make low-level optimizations during resolution, e.g., in version parsing, that are not possible in pure Python code. > If pip extracted multiple files/wheels in parallel without being blocked by the GIL, pip could easily match or outcompete uv. I'm a bit confused by these claims about the GIL? The expensive IO operations release the GIL. > When the user edits one file, all files are modified simultaneously This is why we default to reflinks or copy-on-write semantics when creating environments, not all file systems support it but it's becoming more common. > Hypothetically, a simple deduplication of binary files (.dll .so) should achieve 50% of the savings without significant drawbacks We also explored this (see https://github.com/astral-sh/uv/pull/19694 https://github.com/astral-sh/uv/pull/19694) and the linked pull request has a table comparing to this strategy. > We have optimization for empty files (0 bytes) because there is nothing to write and checksum. Interesting, I would be very surprised if this made a significant difference? but I'll take a look.
- user5994461 1mo ago> I'm a bit confused by these claims about the GIL? The expensive IO operations release the GIL. FYI: The extraction of files is largely python code that doesn't release the GIL. (cf. the zipfile class from the python interpreter has large layers of abstraction with a massive overhead in python code). Regardless, python packages are thousands of tiny files, so pip never gets to release the GIL for any meaningful duration. If you were writing an app that only extracted large GB files, you could take advantage of some I/O operations and some zlib operations freeing the GIL for a bit. Unfortunately pip is the opposite use case, lots of tiny files. > Interesting, I would be very surprised if this made a significant difference? but I'll take a look. Optimizing empty files was actually quite worthwhile for pip, because about 10% of python packages are empty init files. This might not give the same result for uv though. pip is fully linear, every single open/read/write/stat operation we removed was a direct performance gain. uv does parallel async IO, you could very well remove 10% of filesystem calls and barely affect the overall duration. :D
- zanie 1mo agoYou might be interested in taking a look at the `uv download` sketch I started on last week https://github.com/astral-sh/uv-dev/pull/875 https://github.com/astral-sh/uv-dev/pull/875
- notatallshaw 1mo agoThat's excellent news for uv, and I think covers one of two major use cases I often see where uv does not cover standard packaging workflows. This one being downloading to an offline wheelhouse and installing from that. The other one being having a shared named global environment ;o). P.S. I'll have to remove this as an important feature nab has that uv doesn't when I make the announcement nab is no longer experimental, aha.
- ALLTaken 1mo agoHowto use it? sorry for the noob question. I tried `uv --preview-features content-addressed-cache` and I get an error: `uv --preview-features content-addressed-cache error: 'uv' requires a subcommand but one was not provided`