4 ms·
Also as a pip contributor. The only advantage of uv is to have support for parallel async extraction. If pip extracted multiple files/wheels in parallel without
by user5994461 1mo ago
Also as a pip contributor. The only advantage of uv is to have support for parallel async extraction. If pip extracted multiple files/wheels in parallel without being blocked by the GIL, pip could easily match or outcompete uv.
I'm personally not looking forward to any deduplication/hardlink in pip. Hardlinks are very dangerous and system dependent.
It's very dangerous for empty files (init.py, empty.log yet not written). When the user edits one file, all files are modified simultaneously, all venv ever created by the user can be broken by editing one file, which is quite catastrophic.
It's also dangerous for small files with repeated content, for example random settings files that would contain a "1" or "true". Again, when the user edits one file, all files are edited and they were supposed to be different!
Hypothetically, a simple deduplication of binary files (.dll .so) should achieve 50% of the savings without significant drawbacks
I'd venture to say that pip extraction is more optimized than uv in at least one way. We have optimization for empty files (0 bytes) because there is nothing to write and checksum. uv doesn't seem to have the same optimizations, though I could be wrong, I just had a cursory look and my rust is not great. uv should probably review their treatment of empty files, it's counter productive to do any file system operation open/read/write because there is no content, it might be counterproductive to use any cache/comparison/hardlink if it takes more operations than doing nothing.
- zanie 1mo ago(I work on uv) > The only advantage of uv is to have support for parallel async extraction. This isn't true, the biggest speed ups are for the warm cases where we've already unpacked the files into the cache, as notatallshaw mentions above. We can also make low-level optimizations during resolution, e.g., in version parsing, that are not possible in pure Python code. > If pip extracted multiple files/wheels in parallel without being blocked by the GIL, pip could easily match or outcompete uv. I'm a bit confused by these claims about the GIL? The expensive IO operations release the GIL. > When the user edits one file, all files are modified simultaneously This is why we default to reflinks or copy-on-write semantics when creating environments, not all file systems support it but it's becoming more common. > Hypothetically, a simple deduplication of binary files (.dll .so) should achieve 50% of the savings without significant drawbacks We also explored this (see https://github.com/astral-sh/uv/pull/19694 https://github.com/astral-sh/uv/pull/19694) and the linked pull request has a table comparing to this strategy. > We have optimization for empty files (0 bytes) because there is nothing to write and checksum. Interesting, I would be very surprised if this made a significant difference? but I'll take a look.
- user5994461 1mo ago> I'm a bit confused by these claims about the GIL? The expensive IO operations release the GIL. FYI: The extraction of files is largely python code that doesn't release the GIL. (cf. the zipfile class from the python interpreter has large layers of abstraction with a massive overhead in python code). Regardless, python packages are thousands of tiny files, so pip never gets to release the GIL for any meaningful duration. If you were writing an app that only extracted large GB files, you could take advantage of some I/O operations and some zlib operations freeing the GIL for a bit. Unfortunately pip is the opposite use case, lots of tiny files. > Interesting, I would be very surprised if this made a significant difference? but I'll take a look. Optimizing empty files was actually quite worthwhile for pip, because about 10% of python packages are empty init files. This might not give the same result for uv though. pip is fully linear, every single open/read/write/stat operation we removed was a direct performance gain. uv does parallel async IO, you could very well remove 10% of filesystem calls and barely affect the overall duration. :D
- zanie 1mo ago> FYI: The extraction of files is largely python code that doesn't release the GIL. (cf. the zipfile class from the python interpreter has large layers of abstraction with a massive overhead in python code). TIL. I ran some benchmarks and confirmed this is the case for many small files as you'd see in wheels — the GIL is released in a meaningful way for larger files though. Thanks! > This might not give the same result for uv though. Yeah, I built a prototype and ran some benchmarks. It makes a big difference if the entire wheel is empty files but for any real world examples it's within noise of the baseline.
- notatallshaw 1mo ago> We can also make low-level optimizations during resolution, e.g., in version parsing, that are not possible in pure Python code. As a complete aside, uv can and does do this, but for this particular optimization I'm not sure how much absolute time it ends up saving compared to pure Python in real world resolution scenarios. uv's total memory usage isn't that much leaner than pip's, and for parsing speed it turned out that the library pip uses, packaging, was just very unoptimized at the time uv launched. This has been significantly addressed since then: * We did a lot of work to make version parsing twice as fast: https://iscinumpy.dev/post/packaging-faster/ https://iscinumpy.dev/post/packaging-faster/ * Since that blog post I made typical version parse three times faster on top of that: https://github.com/pypa/packaging/pull/1082 https://github.com/pypa/packaging/pull/1082 * Also since that blog post version filtering has gone through multiple optimizations and in some cases will be more than 30x faster e.g. https://github.com/pypa/packaging/pull/1105 https://github.com/pypa/packaging/pull/1105, https://github.com/pypa/packaging/pull/1111 https://github.com/pypa/packaging/pull/1111, https://github.com/pypa/packaging/pull/1120 https://github.com/pypa/packaging/pull/1120 At this point large dependency resolves in pip are spending very little of their time doing things in packaging, like version parsing. The main non-IO time spent in large resolves is now in the core resolver, resolvelib, which I hope to one day replace with my experimental resolver nab: https://github.com/notatallshaw/nab https://github.com/notatallshaw/nab. Nab scales to large resolves much more efficiently than resolvelib (in fact I've cross-ported some of the algorithmic efficiency gains to uv already ;o)).
- optionalsquid 1mo ago> The only advantage of uv is to have support for parallel async extraction. If pip extracted multiple files/wheels in parallel without being blocked by the GIL, pip could easily match or outcompete uv. That sounds like it would be relatively easy to demonstrate: You could tweak uv to perform downloads/extractions sequentially and then compare its runtime with pip. Has any such comparison been done?
- zanie 1mo agoI haven't done the comparison but it shouldn't be so hard - https://docs.astral.sh/uv/reference/environment/#uv_concurrent_downloads https://docs.astral.sh/uv/reference/environment/#uv_concurre... - https://docs.astral.sh/uv/reference/environment/#uv_concurrent_installs https://docs.astral.sh/uv/reference/environment/#uv_concurre...
- notatallshaw 1mo agoFun fact! The original request to add a concurrency variable in uv was requested by me because, among other things, I wanted to measure the difference between uv and pip: https://github.com/astral-sh/uv/issues/3311 https://github.com/astral-sh/uv/issues/3311 I'm not really following this performance discussion as I think it's gone off the rails.
- jvolkman 1mo ago> The only advantage of uv is to have support for parallel async extraction Maybe the only advantage in a particular area (installing)? Because there are many other advantages. The rich lock file for instance allows for much better cross-platform tooling. I can build a linux Docker container from a macos build host, for example, without VMs or any other emulation - simply using the cross-platform details in uv.lock and the correct tooling. I can even cross-compile numpy and other native wheels (linux -> macos, macos -> linux). Other locker tools provide similar cross-platform information (Poetry, PDM), but pip is still lacking.
- amelius 1mo agoImho deduplication belongs at the filesystem level, so the user won't (directly) see it or even know about it. Modern file systems like btrfs have the api for it.
- scheme271 1mo agoThe dedup functionality in something like zfs or btrfs isn't all that great. It tends to be extremely memory hungry and to slow things down significantly. E.g. ZFS needs around 1-5GB of ram per TB of storage and writes need to be compared to a hash table to dedup properly. Using hints or knowledge at the app level is a much better experience if the app can tell the FS that two files are identical. The FS doesn't have to worry about hashing blocks within the file, correcting alignments, etc.
- amelius 1mo agoYes, but package managers are not that great either. Better keep them as simple as possible. And you don't have to do deduplication in an online fashion; you can do it overnight, if you want, as just a simple example.
- Timon3 1mo agoDo you specific issues with uv that makes you distrust the implementation? Otherwise that reasoning is pretty weird - it's possible to write good software, even when the existing options aren't good. Why would we ask them to limit themselves to what might make sense for worse code?
- amelius 1mo agoYou're missing the point. No normal user program can ever do what a filesystem can do: deduplicate in a way that is hidden for the user of the filesystem. Unless you want to change everything into a black box managed by the package manager, making everything confusing for users and also maintainers.
- tedivm 1mo agoThere are a ton of advantages to using uv over pip. I don't use UV because it's faster, although I do appreciate that. I use it because it's smart enough to manage virtual environments for each environment and tool, it can isolate to different python versions trivially, and it handles locking in a way that is actual sane.
- VeejayRampay 1mo agoit's very bizarre for a python developer like me to read comments like this uv has been absolutely black magic for everyone in the way it just works and works much faster, but here you are telling us that actually no it's somehow not that good
- winstonwinston 1mo agoI can say that I have used pip only and it does a package manager job well, it just works. I’m not sure at what uv excel at but reading comments it seems uv is not only package manager but also venv manager.