4 ms·
People are disagreeing with the author, not because they didn't necessarily read the article, but because they don't agree about how things should be defined.
by smallnamespace 6y ago
People are disagreeing with the author, not because they didn't necessarily read the article, but because they don't agree about how things should be defined.
At the root, this is a disagreement about semantics and philosophy, not about git itself. I'm going to refer to Aristotle here: we think we have knowledge of a thing only when we have grasped its cause, and there are four general 'causes' [1]:
- The material cause: 'What is it made of?'
- The formal cause: 'What is the ideal of this thing?' , e.g. what's its abstract nature?
- The efficient cause: 'How did this thing come to be?'
- The final cause: 'What is its purpose?' How is it actually used? What role does it play in the world?
Here we can see that commits are used (at least in the git internals) as 'snapshots' — they refer to bytes, not changes in bytes. That's pretty close to the formal and efficient causes — the abstraction inside of git is closest to a snapshot, and that comes from the history of what Linus wanted when he wrote it.
But! The underlying storage uses deltas (which are diffs) to save space. That's the material cause.
But also, when we actually use commits, git often creates diffs for us as a convenience (cherry-picking, rebasing), and hides the fact that they're snapshots under the hood (final cause).
So there's an inherent tension between the different ways to answer 'what is a thing?'. For commits, this is especially bad, since there's an even split between 'causes'.
This tension never goes away because the most useful definition really depends on the context.
[1] https://plato.stanford.edu/entries/aristotle-causality/#FouCau https://plato.stanford.edu/entries/aristotle-causality/#FouC...
- iudqnolq 6y agoThis is exactly what I'm talking about. A person posts "this is literally how this works", and someone replies "philosophically I would prefer to think it works differently, therefore you're wrong".
- hnjst 6y agoYour distilled summary of this form of objection relying on wishful thinking made my day, thanks a lot!
- efaref 6y agoThe true zen of source control is that they are both.
- cesarb 6y ago> The underlying storage uses deltas (which are diffs) to save space. Not necessarily! The base git storage stores each object individually, not as deltas ("disk space is cheap"); it's only after a "git gc" that they are stored as deltas to other (potentially unrelated) objects. The original implementation of git didn't even have the delta storage (pack files), it was added later as an optional optimization. So answering to "what it's made of?" with "deltas" comes with a huge caveat, that it's often partially or completely untrue.
- haberman 6y ago> But! The underlying storage uses deltas (which are diffs) to save space. That's the material cause. This does not make the "commits are stored as diffs" story much more true: 1. This is only true of pack files, but pack files are only created once the repository exceeds a certain size. 2. Nothing about the pack file format requires that deltas follow the chronology of commits at all. The deltas could be stored in reverse order or even random order compared to the chain of commits. 3. The deltas in a pack file do not correspond to a change in a given commit, they are just the data to create a particular snapshot. If you find that a commit's file blob is stored in a pack file as a delta, that does not tell you anything about whether the file changed in that particular commit. You have to look at two commits and diff them to determine which files actually changed. If a person wants to think about version control in an abstract way, then yes the two views (commits vs diffs) are somewhat interchangeable. If a person wants to understand what actually happens when you run a Git command, the answer to that question is less open to interpretation.