3 ms·
This is the thing though. You're talking about snapshots which actually have duplication removed... in my mind this really fits more with the 'diff' model. I've
by hmsimha 6y ago
This is the thing though. You're talking about snapshots which actually have duplication removed... in my mind this really fits more with the 'diff' model. I've already done the exploratory diving-into-git-internals thing years ago, so I could develop a better understanding of how things actually work.
But for newcomers who want to understand how git is working, it really makes more sense to tell them it's 'like a diff. Not exactly under the hood, but think of it like a diff for now'. This is what I've been telling people as I've mentored a number of people in getting acquainted with git over the years, and if they're curious enough to look under the hood, they'll get a better understanding of the internals.
As a programmer, what you're working with is essentially the diff. This is the easiest way to think about things initially. The fact that git is storing blobs under the hood, shallowly deduplicating blobs but still storing large chunks separately that may contain duplicate data, until it generates packfiles which do a deeper deduplication/compression, is really not that helpful. Telling people it's more like zipping is a bit disingenuous because it doesn't really explain how things are compressed more efficiently over the course of many changes.
If I have a 1MB code file and make 1000 commits of one-line changes then sure, git is initially storing large blobs representing those, but then will compress over the change set when it generates the packfile.
Compared to making a zip of the file for every change (say these are 100KB compressed) and now you have people thinking the 1000 one-line changes generate 100MB in the .git directory.
You may think that a 1MB file with many smaller changes is a fabricated example, but consider that dependency lockfiles (package-lock.json I'm looking at you) can easily grow to this size, and contain this many changes.
- ako 6y agoNot really: if you do a checkout of a snapshot into an empty directory, you expect the entire state at the time of the snapshot, not just the diffs.
- goerz 6y agoIt may depend on the background of who you're talking to. Programmers may be very comfortable with diffs, but non-programmers (in my case, physics graduate students) usually aren't. On the other hand, everybody is familiar with snapshots: even high school student will end up with "report_v1.docx", "report_v2.docx", etc, which are snapshots at the file level (and work reasonably well as long as you have a consistent scheme and don't need branches). I've also routinely seen less-technical people organize their research / paper writing by making a weekly snapshot of their work folder ("project-2020-04-1"). Telling these people that git basically does the same thing for them automatically with a tree-like "labeling scheme" that allows for branches tends to go over quite well, in my experience. For actually programmers, I'd be inclined to give them a more technical introduction to git's internals. I'd still point out that git stores compressed snapshots, not diffs (especially if they're older and may have previous SVN experience)
- hmsimha 6y agoThose non-programmers are likely going to have a worse understanding of what is happening when you zip/compress something anyway, but I concede this is probably the most straightforward path if they have some understanding of what a zip is, and can't understand what a diff is. But even then I question if they should be using git, since `git diff`, `git show`, basically everything git exposes, is going to show them diffs.
- sagonar 6y agoA storage with pure diff would be impossible to recover if you get a error in any commit. It would also be much slower to examine the data, and newer version control do not use pure diff. The version control system Mercurial had description about these problems on the homepage, "behind the scense", which was good reading. I am not sure if GIT is the best solution, but at least a "pure snapshot" is okey, but where a diff storage must in practise include some snapshot logic as well.
- KptMarchewa 6y agoDiff based, but with snapshot "control frame" every N commits, like video? Obviously joking though.
- gowld 6y agoAs a programmer I care about diffs only when I am comparing two versions. A commit creates a new version. "Snapshot" is a distraction.
- goerz 6y agoAlso (for less-technical audiences), I don't exactly dwell on the de-duplication. It's just "Git makes snapshots and puts them into .git in some efficient way. Don't worry about it. Or, if you want the details, read the Git SCM book."
- bosswipe 6y agoThe diff mental model doesn't work for things like `git checkout <commit>`.
- hmsimha 6y agoI actually haven't had a problem with this, though perhaps it's because I understand what's happening at a deeper level. You're generally referencing commits which exist somewhere in this family of commits you can view with `git log --graph`. You can easily think of checkout as the path of diffs to get there. Files at commits are still whole objects, mentally, but the thing we care about as programmers working with multiple versions are the diffs. I have had it break down a bit more when working with stash though, because now the object you're referencing can exist outside of that graph-like commit family.
- formerly_proven 6y agoThe "snapshots which are stored as deltas, if that works" part is unrelated to the diffs the git porcelain generates for you when you do a git-diff or git-show. The former is purely an implementation detail of the storage (albeit an important one), while the latter is entirely virtual, calculated from the snapshots every single time you view the data. That's why operations like git-diff and git-blame can take some time on large trees or histories (and why e.g. git-blame has various options to tweak how it tracks files across revisions, because that is not something git does), while git-log is fast.