7 ms·
> I wonder why Linux distributions such as Ubuntu still download the entire new packages on an upgrade. A lot of upgrade time and bandwidth could be saved by on
by nuclear_eclipse 15y ago
> I wonder why Linux distributions such as Ubuntu still download the entire new packages on an upgrade. A lot of upgrade time and bandwidth could be saved by only sending the differences. And it would reduce load on the mirror sites.
Speaking as someone who has worked on this problem for my own projects [1], I think I can answer this.
Fedora/Yum already supports downloading a binary diff between rpm packages to reduce download size, but this also requires keeping a cache of previous rpm's to run the patch against. There are multiple reasons why you can't rely on binary diffs against the files actually stored on the system, most namely for files like /etc/* that are more than likely modified since installation.
But the real problem with binary diffs is that unless you're doing what Google does to ensure that people stay up to date, the number of binaries you need to diff against grows very quickly, and there are a lot of edge cases to take care of.
For example, let's assume some package A has been released as version 1, 2, and 3. When A has a new release 4, you obviously want to build a diff against release 3, but then you also most likely need or want to build a diff against 2 and maybe even 1 to take care of people who haven't already upgraded to 3. And even if you build a diff against every single version ever released, you will still always need to provide a full version of the package as well for two cases:
1. New installations, or reinstallations, of the package.
2. When the user has cleared their package cache to save room.
And even beyond that, creating diffs involves a lot more effort and knowledge on the part of the packaging team because they not only need to know how to build those diffs, but they also need to keep track of old package versions to build those diffs against.
The end result is that you trade download bandwidth and time on part of the server and end users for a lot of effort, time, and storage space on part of the packagers and distro mirrors. For mirrors that are already encroaching on 50GB for a single release of Ubuntu and/or Fedora, adding a whole bunch of binary diff packages will most likely grow the repository size by at least 30-50%, if not more, depending on how many old versions you diff against.
The question then becomes: does this trade off actually make sense, or does it present further roadblocks for contribution from packagers and donated mirrors?
[1]: If you would like to see how I handled this sort of task, I have a Python library I wrote to handle the client side updating. I know it's not the entire piece of the puzzle because it doesn't cover generating the updates, but it might be useful for someone else. http://github.com/overwatchmod/combine http://github.com/overwatchmod/combine
- rjh29 15y agoYou mentioned storing diffs against every previous version. In principle, couldn't you, when a new package is pushed to the repository: 1. Diff against the previous version and store the diff. 2. Delete the previous version. You could upgrade from any previous version by applying all the diffs in sequence, and you only need to keep one full version around. You could also discard diffs after a certain date because, as you point out, the worst case is that the full version is used instead. I guess this would increase disk space for the mirrors, but even 100GB wouldn't be a lot of disk space, and the savings for (presumably more expensive) internet data transfer would be a lot bigger?
- nuclear_eclipse 15y ago> You mentioned storing diffs against every previous version I didn't say you need to. Eg, with my software project, I wrote an update system that supported diff upgrading against the two latest versions, and anyone still running an older version had to download a full update. > You could upgrade from any previous version by applying all the diffs in sequence At that point you also need to make sure that you aren't downloading more in the process of applying a series of patches than you would need to download for a full update, which also means you need to start being aware of multiple update options, which balloons the complexity of your update code.
- true_religion 15y ago> At that point you also need to make sure that you aren't downloading more in the process of applying a series of patches than you would need to download for a full update Well you wouldn't really need to but it would be a good sanity check on the client side. The common case could be that people are simply keeping up with the stable edge w/o patching binaries themselves out of band---that's the case with some desktop linux variants.
- leif 15y ago> At that point you also need to make sure that you aren't downloading more in the process of applying a series of patches than you would need to download for a full update, which also means you need to start being aware of multiple update options, which balloons the complexity of your update code. The obvious way to handle this is to store full package "snapshots" on "major" version releases (probably upstream releases for packages with lots of local patching, or "one level up from the bottom" releases) and diffs in between. That is not a lot of code if you already have a sane way of managing release numbers within your package manager, which you hopefully do.