11 ms·
Picturing Git: Conceptions and Misconceptions
- karaterobot 5y agoI read most of this long article, and I found it useful, but: It's unsurprising that people's mental model of git is incorrect. Git is not something people study at a conceptual level, it's something they learn recipes for in order to work on some project. Recipes like "how do I save all this work I just did" and "oh shit, everything is hosed, please give me a magic spell I can paste into my terminal to fix it". I don't really blame people, since git itself does nothing to teach you how it works. Git it is the definition of something you have to deal with in order to do something more important to you. Some people want to dig deep and understand how the system works: it's nice to sit near that person and ask them for help sometimes. Saying "you should really understand more about git" is like saying "you should really study the tax code, it's important and it affects you whether you like it or not." True, but deeply irrelevant!
- ilikepi 5y agoIn a literal sense, sure, git _the tool_ doesn't do much, though I think this is slowly improving as it evolves. For example, there is an experimental `git switch` command[1] under development to provide a simpler interface for changing branches. For me, the biggest leap in developing my own mental model was reading Scott Chacon's book Pro Git, and that is now available online for free on the official Git website[2]. [1]: http://git-scm.com/docs/git-switch http://git-scm.com/docs/git-switch [2]: http://git-scm.com/book/en/v2 http://git-scm.com/book/en/v2
- deleted 5y ago[deleted]
- Certhas 5y agoI think it's the other way around. The fact that git does not provide a clean analogous way to intuitively interact with it just demonstrates that the git interface is horribly broken. This is not essential complexity, it's just bad design that stuck. Take a look at https://gitless.com/ https://gitless.com/ If you just look at a summary of the commands, you will have an accurate mental model of what's going on: gl init - create an empty repo or create one from an existing remote repo gl status - show status of the repo gl track - start tracking changes to files gl untrack - stop tracking changes to files gl diff - show changes to files gl commit - record changes in the local repo gl checkout - checkout committed versions of files gl history - show commit history gl branch - list, create, edit or delete branches gl switch - switch branches gl tag - list, create, or delete tags gl merge - merge the divergent changes of one branch onto another gl fuse - fuse the divergent changes of one branch onto another gl resolve - mark files with conflicts as resolved gl publish - publish commits upstream gl remote - list, create, edit or delete remotes To me this clearly demonstrates that the problem isn't that people aren't learning git, it's that git is bad to learn. Stash + Index + Working Tree isn't the right abstraction to present to people. Just say there is a working tree, and tracked and untracked files and snapshots. Done. Branches aren't particular commits but particular working trees on top of particular commits. Working on a feature and want to look at the main branch, but not ready to commit the changes yet? Well just switch to the main branch, then switch back and pick up where you started. No need to know about an additional data structure called the stash. Unfortunately this did not pick up enough steam. And because a lot of tools expose concepts from gits broken interface you have to learn the git interface anyway...
- zauguin 5y agoHaving used `gitless` a while ago as my main interface I strongly disagree. Having a distinction between my working tree and things I'm actually considering to commit is a luxury you only really start to miss when it's gone. IMO gitless makes it way too easy commit too much. Also it's "feature" of keeping uncommitted changes local to the branch is just weird. If I want to make a branch specific change, I create a commit. This has the big advantage that it actually forces the user to add a message what the change is about, so if something else comes up I know what was going on when coming back to it later. It's not like this has to be a formal commit message, after all the commit can be dropped again later. Otherwise you end up being surprised by old experiments when switching to branches you haven't used in a while. If I just switch branches then the most likely reason for that is that I want to move the changes.
- Certhas 5y agoInteresting. I haven't met many who have. Your two usecases basically never arose for me. If I switch back to a branch and there's random stuff there, then I can just revert easily. So it's an extra operation at a different time to get there. The other usecase for switching branches temporarily where it's one less command, is more important to me though. The crucial thing though is that both behaviours can be accessed easily but we are dealing with one less data store/stateful thing, because we don't need the stash. As for the first point, fine grained control for what goes into a commit, that's definitely a power user feature, but an important one of course. Again there are ways to achieve this without introducing new state (the index), for example by allowing to amend the last commit. I wouldn't claim that gitless is a 100% complete git replacement for expert users. It just shows that git has way too much state exposed to users, and has confusing commands to make that state interact. Obviously we all learned git and use it successfully, so it's obviously not broken or anything, it's just worse than it could be (and the constant chorus of "it's so simple, just a DAG!" is a bit grating if you have to teach beginners regularly). The gitless authors did do some research with users that backs up the claim that this is conceptually easier to use: https://spderosso.github.io/oopsla16.pdf https://spderosso.github.io/oopsla16.pdf
- 5y ago
- cerved 5y agoidk man, if you're a software engineer I think the onus is on you. There are plenty of great and free resources, like the pro git book. Every month there's a thread where a bunch of people come in and bemone how git is complicated blah blah. Every month lots of people point out that git is much easier to use if you just bother to conceptually learn about it's internals. it's like coming into a forum for accountants where people bitch about having to learn tax code. please...
- tomxor 5y ago> I don't really blame people, since git itself does nothing to teach you how it works. Git it is the definition of something you have to deal with in order to do something more important to you. Some people want to dig deep and understand how the system works: it's nice to sit near that person and ask them for help sometimes. The official git handbook, freely available on the official git-scm site is not terribly long, and explains the internals on a conceptual level quite well. I think the problem is most people learning git land on some wordpress site of someone trying to flog a condensed and uninsightful shortcut to getting started with git for ad clicks, which only involves a series of commands without explaining the effects of those commands - This, combined with peoples expectation that an SCM should take no thought whatsoever causes most people that use git on a day to day basis to not really understand it at all. Git needs to be introduced as powerful data structure, kind of like how SQL is not a DB, imagine someone explaining SQL without ever refering to the DB tables, rows and fields... only talking about git commits is like only talking about the result of a single query. You must understand the data structure to easily use the interface, otherwise the interface will be very confusing or you will be limited to "recipes"... after that you are just learning new variations on how to manipulate and navigate that structure (yes the graph), and from this perspective peoples complaints about the historical inconsistencies we have to put up with in git porcelain are moot.
- MathMonkeyMan 5y agoI was going to write a blog post conveying my mental model of what git is (having had one too many conversations along the lines of "no, git is not a ledger of diffs"). So, I started reading through <https://git-scm.com/book/en/v2/Git-Internals-Git-Objects https://git-scm.com/book/en/v2/Git-Internals-Git-Objects> again to make sure I didn't have anything wrong. But now there's no point in writing a blog post. Maybe I'll write one that just links to <https://git-scm.com/book/en/v2/Git-Internals-Git-Objects https://git-scm.com/book/en/v2/Git-Internals-Git-Objects>. It even has nice diagrams, which I think are essential for this kind of thing.
- MathMonkeyMan 5y agohere: https://www.davidgoffredo.com/git https://www.davidgoffredo.com/git
- cryptonector 5y agoI use this when I want to teach someone Git: https://gist.github.com/nicowilliams/a6e5c9131767364ce2f4b3996549748d https://gist.github.com/nicowilliams/a6e5c9131767364ce2f4b39...
- tempodox 5y agoI find that a good introduction. “All operations on a repository involve adding commits and/or manipulating the name resolution table.” It may be simplified, but that statement alone, taken in context, is worth its weight in gold.
- cryptonector 5y agoThanks! It's simplified, but really, not that much.
- ScreaminScott66 5y agoSimplified? Really? Heres a example only about one screen in.... "a bag of "commits" identified by a cryptographic hash value and which organize into trees via parent commit references in each commit" Say what?
- OneEyedRobot 5y ago>Some people want to dig deep and understand how the system works I'd say that's definitely the case but also a problem. Sophisticated users mixed with people who just want to do a few simple things is a bad combination. I seem to remember that ClearCase had the same issues.
- dreamcompiler 5y agoThe tax code is a completely inscrutable mess but git's internal model is one of the most simple and elegant structures in modern computer science. It's just covered over with utterly stupid commands and terminology that obscures the beauty of the underlying architecture. I used to despise git because it was so hard to learn. Then as an exercise I started writing my own code to read and write its underlying files and it finally dawned on me how simple the whole thing was. Git's a very unusual piece of software; it's mind-bogglingly useful, the basic data structures and algorithms are perfectly matched to its job, and it has a UI that's a train wreck.
- darekkay 5y agoThat's the blog post I always wanted to write. So many people spend little to no time to actually learn Git, because "it's just a tool to help you doing your "real" work" (=coding) . Or because "Git is too difficult/confusing/broken". Or because "I don't need anything except commit/push/pull". I find those arguments somewhat true, but I still feel that people are missing out when they don't learn a tool they use daily properly. In best case, it makes them less efficient. In worst case, they get into "unsolvable" issues and/or pollute the Git history with useless commits, making blames more troublesome for the whole team.
- jimbob45 5y agoIt is just a tool to help you do your real work and famously gets in the way. You use SVN and it covers 99% of use cases much more simply than Git manages.
- 8372049 5y agoWhen you write things like that, you come across as trolling. Their point wasn't git vs. another tool, their point was the importance of knowing intricate details about your tools in certain circumstances. If SVN is wonderful for you: Great! But that's not really relevant to the issue of using git effectively.
- sagonar 5y agoIf svn covers 99% of your use cases, then you need more experience with distributed version control systems. Able to commit locally, examine changes work with them and then push is a something you might not need or require if you think about version system like SVN. But if you have learned Git or Mercurial or some other distributed system you would never go back to svn.
- maccard 5y agoI have lots of experience with git (5 years of usage, 1 of those years was writing tooling in and around git) and pretty much the same experience with perforce, and I much prefer the centralized model of perforce to all the extra fun that comes with git
- wokwokwok 5y agohm. vexing. I feel like this is mostly accurate, to my knowledge, but reading this: > I do not claim that this way of looking at Git represents absolute “facts” in any hard and fast or literal sense. But I contend that if you conceive of Git in the way that I’m going to suggest, if you substitute these conceptions of Git for any misconceptions you might have now, you’ll be a much happier and more fluid Git user. …vexes me. “Think of git like bowl of peanuts and marshmallows” and other pointless, wrong, metaphors about how git works are a dime a dozen. Yet, here is someone who is clearly quite familiar with git, and they go to pains to point out they are simplifying and may not be correct in their explanations. Its good to be humble, but ffs, git is too frigging complicated if the best you can get is a “probably wrong simplified mental model of how it works so you can be a bit more productive with it”. I dont care; - a simple meaningless metaphor that lets you be more productive? OK. - a accurate description of how things actually work? OK. …but pick one. What I do not want is a possibly wrong complicated explanation of how git maybe works.
- ori_b 5y agoAnd it's not even necessary -- the git data model is simple. It's simple enough that you can generate valid commits in about a page lines of python, with no libraries. The rest is just packing files for efficiency, and finding the difference between hash lists for syncing and pulling. The command line interface to git is insanely complicated, confusing, and unnecessarily difficult to use, but this isn't a result of the git data model. It's definitely possible, to give a complete and accurate description of the data model, even using examples from `git cat-file` to walk through the commit history by hand. I've also got a simple demo that generates a complete repo with a commit. You can manipulate the resulting repo from git. There are 65 non-comment lines of code. Here it is: https://orib.dev/ugit.py https://orib.dev/ugit.py
- Quenty 5y agoThis article presents an accurate picture of how things work at a high conceptual level. It glosses over certain details, because git is very complex. For example, git has, if I recall correctly, 4 staging areas, of which represent different sources when it comes to a merge conflict. However, this detail can mostly be ignored because it’s not relevant to the high level conceptual ideas this article is trying to present. I would argue most things in technology are complex, and mental models are intentional ways to take something complex and turn it into something more simple. This article does not create meaningless metaphors.
- bironran 5y agoUgh. So many concepts. So many things to remember. Why? Git is simple. SIMPLE. But only, IMO, if you go bottom-up and not top-down. There are only 6 critical concepts in Git and each is simple enough to be described in a single sentence. 1. Commits are immutable blobs that have one or more parents. Graphs, not trees. Anyone who uses trees for git commits misses the whole point and makes their (and their collaborators) lives complicated. 2. Tags are (mostly, best practice) immutable pointers to commits. Tag are "this is this thing FOREVER*." 3. Branches are named, mutable (by design) pointers to commits. Branches are "this is this thing FOR NOW. Later it'll be something else." 4. HEAD is special "branch" that moves around automatically. 5. Origin is the local snapshot of the remote. Origin is "what did it look like when I last looked." 6. (fundamental but not critical) Remote is the current remote state (queried by RPC). 7. Index (aka stage) is where you put changes you want to make into commits. (this is somewhat simplified). Index is "My current and immediate plan. Scrub as needed." That's (mostly, for non advanced use cases) it. Everything else are commands to query or manipulate the various state. Every action (until it becomes instinctual knowledge) should follow the same recipe: 1. Figure out the current state (current commit graph, relevant branches). 2. Figure out the target state (desired commit graph, new branches positions). 3. Mutate using ANY command you want. I think that's the issue really. Inexperienced dev / people who don't understand git look at commands as "this is how to do a thing". No. In Git there isn't "how to do the thing". It's exactly like writing code - so many ways to achieve the goal, just choose your own. It might be efficient and elegant, or bumbling and ugly, but it'll get there.
- megous 5y agoReminds me of: Bad programmers worry about the code. Good programmers worry about data structures and their relationships. -Linus Torvalds Anyway, it's not for everyone to get to understand git this way, I guess. Some people will just react "just tell me how to do X in git!"
- bironran 5y agoLike a lot of things Linus, it's very pretentious and aloof but right at the core of it. Code matters a lot and bad code can tank performance, stop evolution and introduce security issues. But with Git, this is a mostly truthful statement.
- mynameismon 5y agoRelated: A video[0] from The Missing Semester, a course by CSAIL MIT, which covers Git in one of it's lectures. Personally, the entire series is a must watch, but if time is limited, the first 20 odd minutes give an absolutely fantastic introduction to Git. [0]: https://missing.csail.mit.edu/2020/version-control/ https://missing.csail.mit.edu/2020/version-control/
- Amin699 5y agoThe purpose of this article is to present a simple way of looking at what Git really is and what it really does. I do not claim that this way of looking at Git represents absolute “facts” in any hard and fast or literal sense. But I contend that if you conceive of Git in the way that I’m going to suggest, if you substitute these conceptions of Git for any misconceptions you might have now, you’ll be a much happier and more fluid Git user. When posed with a puzzle as to what happened, what will happen, what you should do in order to make a certain thing happen, the answer might suddenly be obvious, where previously it wasn’t.
- jldugger 5y ago> Picturing Git: Conceptions and Misconceptions Based on the title, I was expecting a more in-depth study of user misconceptions about git, similar to the famous CogSci paper "Two Theories of Home Heat Control." Except with like, diagrams. And now I want someone to make that happen.
- avip 5y agoWhile you're on wait there, you can read the excellent hg init https://hginit.github.io/ https://hginit.github.io/ for inspiration.
- rendall 5y agoI avoid using words like "simple" when writing about technical topics. "It is simple!" is not inclusive language and, love git or hate git, it is not inherently simple. It takes effort to understand, and experience to avoid its pitfalls.
- thewebcount 5y ago> The problem with how people use Git, I’m suggesting, is that their analogical or metaphorical conception of Git doesn’t work — it doesn’t fit the way Git actually behaves — if, indeed, the conception exists at all. No, the problem is not with "how people use Git". The problem is with git. We've known for years how to make clear, concise interfaces that help people understand what's going to happen. Git does not have a clear, concise interface. That is its biggest problem and will continue to be until it is changed to have a clear, concise interface.
- cerved 5y agogit has a very clear, concise and stable interface if you understand how git works. it's designed this way, intentionally. people should stop complaining about it and either learn how to use it, switch to another tool or just write their own interface
- recursive 5y agoIt's not possible to switch to another tool unless you only ever work on code that you wrote, and never need to collaborate with anyone else.
- jatone 5y agono, its absolutely possible, as demonstrated by the numerous repositories that switched from cvs and svn to git. what you describe is a lack of desire to switch to another tool by your coworkers because you've been unable to make an adequate case for doing so.
- cerved 5y agoNo it's totally possible to use your own VCS and sync changes to another VCS. I don't know which tool you'd prefer but I used a git repository for my own work and synched it to a TFVC repository until my company switched to git. It sucked, because TFVC sucks but it's totally doable
- squaresmile 5y ago
- _aleph2c_ 5y agoGit isn't going away, so we might as well master it. If you would like to know more about how to manipulate the git graph, take this excellent (and free) training: https://learngitbranching.js.org/ https://learngitbranching.js.org/ To slowly level up, you can watch video demonstrations from Dan's git school. Dan provides 48, 30 minute training videos: https://www.youtube.com/watch?v=OZEGnam2M9s&list=PLu-nSsOS6FRIg52MWrd7C_qSnQp3ZoHwW https://www.youtube.com/watch?v=OZEGnam2M9s&list=PLu-nSsOS6F...
- toiletaccount 5y agoComplaining about git is like complaining about any other unix tool, if you don't read the docs you're in for trouble. Sometimes a lot of it.
- ScreaminScott66 5y agoWhat gets me is how complicated Git is when all I want to do is check out the production code for ONE file, make a change, and check it back in. With git, I need to clone a repository, create a branch, do a add to move my changes to a staging area, commit the changes to update the local repository, then some combination of pull requests and merge (haven't figured out that part yet) to eventually get the changes back into the remote repository. And now this article tells me that not only to I need to know all those commands, but I need to understand the structures behind them? Geez, I just wanna change some code, test it and check it in.