24 ms·
Git concepts simplified
- dcre 13y agoI love Git, but I'm pretty sure there's no way to explain it simply.
- ulisesrmzroche 13y agoThere's a tree with branches and that's about it.
- jordigh 13y agoI have a much easier time explaining Mercurial. And I will keep using Mercurial and keep improving Mercurial until we all realise how much better it is than git because it's easier to explain, it's just as fast, and no less powerful.
- jlgreco 13y agoI don't get this attitude. The DAG is exceedingly simple and I think can be taught reasonably in only a few minutes. From there you really only need to teach a very minor amount of the UI, and teach the user how to perform "I want to do this to the DAG"=>"This is what I type" translations on their own. I've seen all of this done well in sub-hour presentations. Mercurial on the other hand has a pain in the ass datamodel (so much so that most introductions to it that I have seen do not even approach the topic), so you actually do have to learn all of the UI commands to get an idea of what can be done and what cannot be done. It is far more complex than git. I really cannot think of a simpler VCS than git. I've used plenty, but never got off the ground faster with anything else.
- volaski 13y agoGit is simpler than any other VCS, but VCS itself is not simple concept to understand. That's what he's saying, and I agree. I use git all the time, but it's not easy to understand the whole version control thing as a beginner. Also Git is not just about DAG, you need to be able to understand the decentralized nature, etc. In that sense, some people may find SVN easier to understand.
- dcre 13y agoExactly what I meant. If you're working with people who need to learn what a DAG is in the first place, that adds some overhead. After that you still need to understand that your working directory is just a scratchpad, and branches aren't actually like real-life tree branches but instead they're pointers to the ends of those real-life branches, etc. Not exactly simple. Conceptually speaking, a centralized VCS -- one where you check out files, make changes, and check files back in when you're done -- is vastly simpler than Git. Sure, it's also much less powerful (and I would choose Git a million times over such a system), but it's definitely simpler.
- damncabbage 13y agoI agree with the model being great. The git command itself is abominably baroque in its user interface (inconsistencies and strange defaults abound), but I've gotten over that with more effort than I'd like to admit. I love Git's plumbing; I just hate its porcelain.
- jlgreco 13y agoEh, I guess I just sort of see hating the porcelain like hating the parens in Lisp.
- Perseids 13y agoOne of the great reasons to use Haskell instead :D
- tootie 13y agoPeople who refuse to see the ugliness in git are the same people who think it's manly to live on the command line. Using git through Eclipse looks a lot like using SVN through Eclipse. If I right-click on a file and go Team > Replace With > Remote and select 'origin master' is saves me from trying to remember the obscure list of flags I need to pass to the command line to do the same thing.
- jlgreco 13y agoRefuse to see the ugliness? No, I really don't see it. Sure, a few flags could be cleaned up, but the beauty of the rest of git more than offsets a few weirdly named flags. Meanwhile git through eclipse causes nothing but trouble as far as I have seen. Making it seem like SVN is exactly the problem, git isn't like SVN so if it seems that way, something is going wrong. Pretending git is something that it isn't will bite you in the ass sooner rather than later. The disappointing part is that there isn't any technical reason why git integration in eclipse couldn't be good, it just isn't currently.
- tootie 13y ago
- jeltz 13y agoAs soon as you understand that git is just a DAG then it becomes simpler to understand and reason about than any other VCS I have used.
- ezquerra 13y agoI find that your comment about mercurial's data model a bit ridiculous. The reason why most introductions to mercurial do not mention its data model is because it is _not_ important. You really do not need to care about it at all on your day to day use. I've been using mercurial for years and I've never had to ask myself what is mercurial's data model. IMHO the reason why git forces you to understand its data model is because its UI is terrible. It is a failure of the tool when you need to understand how it works internally to use it. If git's UI were better its awesome data model would be something that only git devs would need to understand.
- jheriko 13y agoreally? its not that complicated, it just has lots of bad defaults in my experience (i use mercurial when i start new projects, not because its fundamentally better - but because it has sensible defaults and i don't need to configure it our get bitten by gotcha x for the nth time - e.g. i can revert a merge without reading a document or configuring anything - a vital feature of source control imo)
- aeon10 13y agoThere seems to be ALOT of 'Git Explained' stuff. The best way to learn Git in my opinion is start using it! Google stuff as you go.
- recursive 13y agoI thought I would try that, but I'm probably still not really thinking in Git. I never branch or merge for example, since I never think of doing that. I understand this isn't really normal git usage.
- dsego 13y agoMe neither. But I think I should. Every time I want to try something new, take a new direction in code or add a new feature I should branch, so I can safely add or remove pieces of code. Things not related to that new feature, like bug-fixes, should be done in the main branch and then get pulled into the feature branch. I'm just too messy with my commits to do that, because it requires making frequent smaller commits, instead of huge ones once in a while.
- recursive 13y agoI am familiar with that approach, but it's really not how I work. If I'm not done with a feature, I don't check in. And I've never needed to work on two features at the same time.
- develop7 13y ago> Google stuff as you go. Are you implying git man pages are useless to newcomer?
- th3dude 13y agoUseless? No. Overwhelming at times? Absolutely. I think that Google is probably the best way to go, because not only will you get the man pages as high ranking hits, but you'll also get great hits from sites like StackOverflow and blogs that can explain things better.
- tekp2 13y agoI was expecting this: http://tartley.com/?p=1267 http://tartley.com/?p=1267
- octo_t 13y agothats my favourite ever git article.
- nilkn 13y agoI really did not expect that. That's one of the best unexpected funny things I've experienced this week. I don't think I would have found it as funny if I had stumbled on it randomly; the discussions here provided the perfect context for this link.
- pessimizer 13y agoThis did it for me: http://www.sbf5.com/~cduan/technical/git/ http://www.sbf5.com/~cduan/technical/git/
- tootie 13y agoThere is an inverse relationship the number articles title "x explained simply" and the actual simplicity of x. I honestly don't understand why the developer community refuses to admit the obvious that git is unholy clusterfuck of a product. It has a nice data structure inside it? Name another end-user product for which you are even vaguely aware of what data structures were used.
- angdis 13y agoI think part of the problem is that git, like other version control systems, does not enforce any particular way of working (workflow). When authors try to explain how to use git, they often have a very particular workflow in mind and don't necessarily describe exactly what that workflow is. This can cause major problems when somebody tries to apply advice to their own workflow. I like the graphical approach. It helps to conceptualize what is going on so that someone can apply it to their own situation.
- tootie 13y agoThat's not it. Following any reasonable workflow still requires a set of arcane commands and flags that only make the slightest bit of sense if you know how git is implemented. I was able to use SVN successfully without ever knowing a thing about the implementation. I've generally had the same experience with HG and even CVS and VSS back in the day. Git adds sophistication over a prodcut like SVN, but adds vastly more complexity.
- dasil003 13y agoConversely, SVNs data model remained opaque after years of daily use simply because it solves the problem at hand poorly and is not defined very well.
- DougWebb 13y agoI'm curious: did you use branches and labels in SVN? I've come across many svn repositories that don't use the trunk/branches/tags layout, and as a result the developers keep completely separate repositories for slightly different versions of their projects instead of creating branches. I've even seen new repositories created for each release version of the project. If you're using svn you need to understand the implementation to get why copies are cheap, so that you can understand how to use branching and tagging appropriately.
- jheriko 13y agoa lot of this is generally (d)vcs and applies equally to mercurial or even svn... also if you think x pages of anything is a simple explanation then you missed a trick or two. e.g if you have to explain why your arrows are pointing backwards you are doing it wrong, instead of using the standard notation for graphs and lists and stuff which are not generally well known, use what most people will understand on inspection.
- Jugurtha 13y agoI prefer the Fox News tutorial on the subject. 'Repo' means 'reciprocity' or 'reposotory' if you didn't know.
- Jugurtha 13y agoWow! Did people forget what sarcasm looks like ?
- grapeot 13y agohttp://static4.businessinsider.com/image/52330e4deab8eaef7ac8a35b-960/fox-news-interview-github.jpg http://static4.businessinsider.com/image/52330e4deab8eaef7ac... I also wanted to post this screenshot and then saw your comments...
- MBCook 13y agoI'm... slightly afraid. I can't wait to send this to my coworkers.
- ezrasuki 13y agoYeah right.
- WestCoastJustin 13y agoGitolite (where this is hosted) is actually pretty cool too. For those who don't know what gitolite is, it is software to works in tandem with git-daemon, that basically allows you to run a centralized git sever with access rules. I created a screencast about it @ http://sysadmincasts.com/episodes/11-internal-git-server-with-gitolite http://sysadmincasts.com/episodes/11-internal-git-server-wit...
- aylons 13y agoMoreover, if you integrate gitolite with redmine (easy to do with a plugin [1]), you get a great corporate-leve, self-hosted, easy-to-use team management tool for programmers and alike. [1]https://github.com/ivyl/redmine-gitolite https://github.com/ivyl/redmine-gitolite
- beagle3 13y agoThere's also gitlab, which is a github style app you can run locally. It used to use gitolite internally for access control, but it is now using its own access control system.
- csense 13y agoI've used both. Gitolite does everything on the command-line, including access management. Gitlab is essentially an open-source clone of Github's web UI. Of the two projects, I think Gitlab is harder to deploy and far more resource-intensive on the server, but easier for users.
- EdwardDiego 13y agoAnd there's also gitosis, which IIRC is considered superceded by gitolite, but it's still working well for us.
- btbuildem 13y agoThe 90's called, they want their bitmaps back..
- bb0wn 13y agoI found the git book to be more than adequate enough for explaining the structures and practice patterns of git. http://git-scm.com/book http://git-scm.com/book
- gbog 13y agoNice. I'd love to see more details about the index, the way to see differences between branches with log and diff, and I think stash should be mentioned.
- b0z0 13y agoVery nice, and to see it in action, this is also really cool: http://pcottle.github.io/learnGitBranching/?demo http://pcottle.github.io/learnGitBranching/?demo
- Pxtl 13y ago... that is not simple. But I'm figuring it out anyways. I commit my changes to my own repo and keep building changes under HEAD, committing as I go. If I get a branch from a buddy and I want to add it to my code, I either merge or rebase depending how I want his commits to be intertwingled. Because GIT is decentralized, there's no difference between merging in a buddy's branch, and import the latest changes from Origin into my branch. So I fetch the changes from origin/master and then rebase or merge my repo on top of that. That's my "Get Latest" command, basically, right? Assuming I'm working on "master", I fetch then rebase or merge origin/master. To check in, I tell the origin server to take my stuff and then rebase or merge its master with that. I still feel like this is a rather baroque approach to the problem... managing oodles of local commits separate from rebase/merges seems bizarre, above and beyond the decentralized approach that makes my own repo, my peer, and the "origin's" stuff all equivalent. The decision of when to merge vs. rebase is still confusing to me.
- tbatterii 13y ago> The decision of when to merge vs. rebase is still confusing to me. i always thought it was "never rebase commits you have pushed" at least you know about rebase though. :)
- MBCook 13y agoThat is certainly the rule 99.99999% of the time. It's possible, but it can be a mess for everyone else. Before you push, you can choose to do it either way.
- dcre 13y agoOne small point: on check-in, I don't think it's exactly that origin is also rebasing/merging just like you did. It's more like origin is just taking whatever you have and copying it exactly. The merging/rebasing process itself only happens locally. Regarding merge vs. rebase, here's my approach: rebase to keep history a straight line when it's just your changes and it's just a few commits. If it's too many commits you tend to have more conflicts and it's usually easier to merge.
- stormbrew 13y agoA suggestion for anyone hoping to write accessible tutorials: > Hg folks should read this section carefully. Among various crazy notions Hg has is one that encodes the branch name within the commit object in some way. Unfortunately, Hg's vaunted "ease of use" (a.k.a "we support Windows better than git", which in an ideal world would be a negative, but in this world sadly it is not) has caused enormous takeup, and dozens of otherwise excellent developers have been brain-washed into thinking that is the only/right way. If this is one of the most important concepts, starting it off with a slew of negativity about things the reader may be currently using (windows, hg) is probably not the best way to get them to keep reading.
- ender7 13y agoThere's also a bit of "people in glass houses..." to this comment. I don't think git really wants to start a fight when it comes to poor design decisions.
- ranman 13y agoout of curiosity could you enumerate some of the poor design choices in git?
- acqq 13y agoSee https://news.ycombinator.com/item?id=6452379 https://news.ycombinator.com/item?id=6452379 and https://news.ycombinator.com/item?id=6451748 https://news.ycombinator.com/item?id=6451748
- beagle3 13y agoNot sure what ender7 refers to. As far as I know, there are a few things in git that its maintainers think are missing - e.g., Linus mentioned that having a "generation" in a commit, which is 0 for the empty commit and 1+max(generation of ancestors), would have sped up some merge operations. Finding the history of a single file requires traversing the entire commit tree. But there is no "fundamental" problem or bad design choice - and in fact, this data can easily be cached without adding it into the protocol. EDIT: saw some complaints listed in this thread - they all have to do with UI, not with a design choice that is limiting use. Wrappers like "eg" and "legit", and macros, can be used to fix the UI (except no one agrees on what a better UI looks like).
- wiremine 13y agoGit's learning curve feels similar to the learning curve of a programming language like Python. Once you understand you're not going to pick it all up in an afternoon (just like a language) and that there will be lots more to learn down the road (like a language), git feels great.
- guard-of-terra 13y agoProgramming language is 80% of my work effort but version control is more like 5%. It should be an utility, I don't want it to be a world of its own right, I don't have needs for insanely powerful version control system because my needs are sane and limited - and that's where git is not so cool.
- Perseids 13y agoYou can use the same argument about debuggers, dependency managers, editors, etc. If you are happy with whatever tool you are using, why change?
- guard-of-terra 13y agoVersion control is social and there you can see a few maniacs ruining it for the rest of the team. Debuggers aren't much harder than pour and drink. Dependency managers are pain in the ass (unsolved problem in CS) but you don't wrestle with them every day. I don't use very many features of my Eclipse and I don't use terribly many commands in vim. I also use arrows, I kid you not.
- EdwardDiego 13y agoGit's learning curve is far worse than Python's, simply due to the limitations of its abstractions - e.g., a branch isn't really a first class entity in Git, yet most people using the centralized server model for Git will tend to think about their daily Git work in terms of branches. The problem only manifests when you're trying to do things like "show me only commits from branch <X>". Or, "show me when branch <X> was created from master."
- delinka 13y agoSee also Git for Computer Scientists at http://eagain.net/articles/git-for-computer-scientists/ http://eagain.net/articles/git-for-computer-scientists/
- 0xdeadbeefbabe 13y agoIt's possible, at least after I see the concepts this way, that git has the simple design Hoare was talking about when he said: “There are two ways of constructing a software design: One way is to make it so simple that there are obviously no deficiencies, and the other way is to make it so complicated that there are no obvious deficiencies. The first method is far more difficult.” ― C.A.R. Hoare
- mrcactu5 13y agogit init git add . git commit ":-)" git push origin master 6 months later I still look it up... this tutorial is for me!
- anandabits 13y agoThis is very nice as a review of Git, but probably best in that context rather than an initial presentation of the concepts. I really enjoyed the Source Control Made Easy series by Jim Weirich. It presents the same information in an easily digestible, step-by-step approach. Highly recommended for those trying to understand how Git works and how to best make use of it. http://pragprog.com/screencasts/v-jwsceasy/source-control-made-easy http://pragprog.com/screencasts/v-jwsceasy/source-control-ma...
- obilgic 13y agoCan someone please convert this into a nice pdf? Chances are I will do this type of reading when I am offline...
- acqq 13y agoI don't use Git and would like to know how the following problem is solved in Git. Say you have a project which is a hundred megabytes big. And you have to develop almost in parallel three or four "generations" of the project -- let's say. v1, v2 and v3. In parallel means you'd like to be able to build any of the three versions without having to take the version out of the repository first. You can't say that v1 is obsolete, as soon as some bugs are reported in v1 you have to fix them in v1, v2 and v3. And every bigger version is "newer" but some features can be added in v2 and v3 some just in v3 etc. How can you work on such a big project and have a single repository where all three versions are present, and work on these three versions in parallel (having sources which are compiled in different base directories)?
- taspeotis 13y ago> (having sources which are compiled in different base directories) With my nascent Git understanding, I think you would just have multiple branches for v1, v2... and then clone the repository multiple times so you have multiple working copies. Check out v1 in the first one, v2 in the second one. Although changing between related branches is usually quite quick in Git. Also, a fresh checkout of ~100mb is not a lot. At least for an SSD. This also relies on having a centralised Git repository for you to push/pull changes to. But I believe Git allows you to synchronise multiple repositories on disk. You're rarely developing two things at once in any given instant of time... why not just quickly check out the branch you want?
- MBCook 13y ago> and then clone the repository multiple times so you have multiple working copies. This is probably not what you want. First, you should know that switching between branches in Git is insanely fast. In general, it won't get in your way. If you clone the repository, each one is a full git repository. That means you'll triple the storage on the disk. Worse, you'll have to do 3x as many pulls to keep all 3 repositories up to date. > You're rarely developing two things at once in any given instant of time... why not just quickly check out the branch you want? It often comes up, but that's what we do. We may have a dozen branches on our machines (the thing(s) we're working on, recent things we worked on, the one that's been sitting for a while we're waiting on an answer to pick up again) and we can switch our project within a second or two on a simple rotating hard drive.
- lisper 13y agoGreat article. Just one rather glaring omission: only a single mention of the index, and that only in passing. I have used git for years, and I still don't understand what the fleeping index is supposed to be for. What can you do with the index that you can't do with a branch? And why is it called the index? (And why is it git add -a but git commit -A? Or maybe it's the other way around?)
- Cyranix 13y agoNot to be cruel, but I don't understand how you've used Git for years without understanding what the index is. There's more than a few learning resources for Git online to satisfy your curiosity. Anyway, to give an abbreviated explanation: The index is a staging area for your commits. When you use `git add`, changes in the working directory are staged (prepared) for the next commit. If you pass the `-a` flag to `git commit`, Git will stage all changes to files that it is already aware of. (Recall that new files are untracked and must be manually added to the index the first time they're committed; `-a` won't add those files because Git doesn't already know about them.) Why have a staging area instead of just creating a commit directly from all the changes in the working directory? It's basically a sanity measure for organizing commits if you're ever anything less than a perfect developer. If you make a bunch of changes and later realize that there's more than one "unit of work" represented in those changes (however you choose to define those units), you can selectively add files to the index to create commits that make sense. You can even use the interactive mode of `git add` to selectively stage changed sections within a single file. If you care about the benefits of sensible commits -- bug hunting with bisection, ability to run `git revert` to undo a logical unit of work -- then the index is your friend. A few random pages on the index that I pulled up: [0] http://www.gitguys.com/topics/whats-the-deal-with-the-git-index/ http://www.gitguys.com/topics/whats-the-deal-with-the-git-in... [1] http://git-scm.com/book/en/Git-Tools-Interactive-Staging http://git-scm.com/book/en/Git-Tools-Interactive-Staging
- lisper 13y agoWell, I'm being a little facetious. I do understand what the index is and what it's used for (but not why it's called the "index" instead of, say, the "stage"). What I don't understand is why the index exists as a separate abstraction. You could have the exact same effect by, for example, doing a git stash, and then popping changes out of the stash into your (now clean) working directory. The WD in effect plays the role of the index, and you get the same result, but with fewer abstractions, fewer commands, and less confusion. But I hold Linus in high enough regard to take very seriously the possibility that the index is a reflection of some deep wisdom that I have missed. That's the real reason I raise this every now and again.
- forrestthewoods 13y agoIf your "simple" explanation of how to use something is 27 pages then it is many things and simple isn't one of them. My favorite thing about Git is how it's forced Perforce to add features and lower price. Thanks Git!
- andrewflnr 13y agoThis has been danced-around in the comments, so I'm just going to say it: I don't need the concepts of git simplified, I need a better explanation of how git's bizarre command set maps onto the obvious DAG/filesystem operations.
- ams6110 13y agoSeconded. And, at 30 printed pages, I'd hate to see the non-simplified explanation (no, I didn't actually print it).
- Sir_Cmpwn 13y agoIt's the stuff behind the curtain that makes me love git. It doesn't just seem like version control software - it seems like the software is an interface to a much more powerful version control engine. Git just makes sense under the covers.