5 ms·
I know what you mean, but — as the article explains [0] — Git’s packfile format uses delta compression, which is essentially “storing diffs”. [0] https://githu
by tomstuart 4y ago
I know what you mean, but — as the article explains [0] — Git’s packfile format uses delta compression, which is essentially “storing diffs”.
[0] https://github.blog/2022-08-29-gits-database-internals-i-packed-object-store/#delta-compression https://github.blog/2022-08-29-gits-database-internals-i-pac...
- srvmshr 4y agoIn my mental model of git internals: I envision the initial commit stores a full copy in object store - and thereafter the states are captured as diffs. If a new file is added at a later commit stage, of course the "blob" will be stored, but subsequent commits will add the deltas to it - not a full copy.
- masklinn 4y agoThat's not the case. The actual model of git internals is that every commit always stores full copies (of modified entries, unmodified entries are dedup'd through the magic of content addressing). Then as an optimisation of the packing process git will perform delta compression between objects based on a bunch of heuristics. Until you pack you get full copies, and depending how "hard" you pack you may or may not gets huge savings. Furthermore, Git's pack files are ordered, so depending on the selection it makes it can diff later files from earlier ones (your model) or the reverse, or even both at the same time. You can see some of the early packing heuristics at https://github.com/git/git/blob/master/Documentation/technical/pack-heuristics.txt https://github.com/git/git/blob/master/Documentation/technic...
- cryptonector 4y ago> That's not the case. As far as the actual internals are concerned, yes. But as a model it's not wrong. It's like epicycles aren't wrong as a model either. GP's mental model of Git is mine as well, and it works very well for me.
- masklinn 4y ago> But as a model it's not wrong. As a model it's wrong, and useless, and possibly harmful. Because the "states are captured as diffs" model does not tell you anything true that's useful. And at worst it gives you incorrect ideas of the performance characteristics of the system e.g. that small changes are bad, because lots of changes means lots of diffs to apply which is slow. By forgetting that model, you're strictly less wrong, and helped no less about reasoning about git. > It's like epicycles aren't wrong as a model either. But they were. "Adding epicycles" has literally become a saying about trying to pile on more crap to try and (fail to) make an incorrect model work. And epicycles were at least trying to keep a simple model working, the "snapshot model" is more complicated than git's actual behaviour to start with.
- robertlagrant 4y agoCalling things harmful considered harmful.
- cryptonector 4y agoEspecially without an explanation. them: harmful me: how? them: really harmful me: sure, but, can you explain it? them: like a bajillion harmful
- cryptonector 4y ago> As a model it's wrong, and useless, and possibly harmful. It's no more wrong than epicycles. It's not useless either: it works very well for me. For users (as opposed to Git developers), it's hard to see how it could be harmful.
- johannes1234321 4y agoThe harmful part comes from assumptions based on it. Systems like CVS or Subversion indeed store diffs of file conceptually and storage wise. This has notable co sequences: when doing operations on a large span of revisions all takes long as a bunch of diffs have to be applied in sequence. This then leads to reluctance of small commits for the wrong reasons. Wrong assumptions about the workings leading to wrong decisions is harmful.
- cryptonector 4y agoThis is a good mental model. The internals are the internals. But the model users see is that a repository is a bag of commits (deltas) that organize into a tree, with branches and tags being names for specific commits. With this model in mind anyone can become a Git power user.
- stu2b50 4y agoFrom a user point of view, you should not see commits as diffs or deltas, but as a full snapshots of the contents of the repository. Delta compression is just an implementation detail for storage optimization. Conceptually each commit is an immutable snapshot in its own right.
- cryptonector 4y ago> From a user point of view, you should not see commits as diffs or deltas, but as a full snapshots of the contents of the repository. That tends to lead to a merge-centric workflow. The alternative makes it easy to understand rebase. Ergo it's good.
- stu2b50 4y agoI don't see any particular reason that'd be the case. Even with rebase, the diff model causes some of git's behavior to no longer make sense. For instance, the fact that rebase causes an entirely new series of commits to be made - this is because rebase is just a layer over cherry-picking commits, and that cherry-picking commits calculates the diffs between two commits then applies the diff as its own, new commit. If commits were diffs, then that process doesn't make any sense. It only makes sense when you think of commits as snapshots, and cherry-pick ing as taking a commit, calculating the diff with the current commit, and then applying that diff to make a brand new commit.
- cryptonector 4y agoYes, rebase is a series of cherry-picks. Each commit picked is treated as a set of diffs that are applied with conflict resolution, then committed, then off to the next. The mental model of commits as diffs works, obviously. > If commits were diffs, then that process doesn't make any sense. It only makes sense when you think of commits as snapshots, and cherry-pick ing as taking a commit, calculating the diff with the current commit, and then applying that diff to make a brand new commit. Sure it does. It works for me. You can use whatever mental model of Git that you works for you.
- JoshTriplett 4y agoGit actually arranges to delta starting from the current version to older versions, so that it has less work to do to show you recent versions, and does more work as it works its way backwards in time.
- eminence32 4y agoWhen you rename a file in git, you can get a glimps of its internal model -- renames are not explicitly tracked. Since git only stores full copies of each object, it has to employ heuristics to figure out if a (deleted file, created file) pair should be considered a rename. This is why you sometimes see a percentage next to a rename in a "git log" -- this is git's internal metric about how similar/dissimilar the files are.
- rossmohax 4y agoIs there any value in first class renames other than some performance boost ?
- dlkmp 4y agoI'd argue yes. I remember times where I was puzzled that some config file had been deleted (in a commit touching lots of files) only to find out after some analysis that is was simply moved/ renamed.
- masklinn 4y ago> Git’s packfile format uses delta compression, which is essentially “storing diffs”. FWIW it's closer to lossless compression e.g. LZ77's behaviour: git's delta compression has two commands which are LITERAL and COPY, and technically it doesn't need to copy linearly from the base object, it would be unlikely[0] but a delta could start by COPY-ing the end of the base object, then add literal data, then COPY the start of the base object (or COPY the end again, or whatever, you get the point). [0] outside of handcrafting, as the packing process would need to find such a situation which it would estimate useful
- dunham 4y agoMany years ago, I wrote a some code that made pack files with out of order copy commands. (Just messing around in Go trying to learn how pack files worked.) I was too lazy to implement the diffing algorithm, so I used the suffixarray from the go standard library and greedily grabbed the longest match at each point in the input. It's been a while, but I think I got a 12% improvement on my test case. (Probably wouldn't do that well in general.) I'm assumed the git code just runs off of a diff (longest common substring), but I never checked.