3 ms·
Flashback: Discussion of merging 2009: Git: Bram Cohen vs Linus Torvalds http://news.ycombinator.com/item?id=505876 http://news.ycombinator.com/item?id=505876
by notaddicted 14y ago
Flashback: Discussion of merging
2009: Git: Bram Cohen vs Linus Torvalds
http://news.ycombinator.com/item?id=505876 http://news.ycombinator.com/item?id=505876
which refers to
2007: A look back: Bram Cohen vs Linus Torvalds
http://www.wincent.com/a/about/wincent/weblog/archives/2007/07/a_look_back_bra.php http://www.wincent.com/a/about/wincent/weblog/archives/2007/...
which refers to
2005: Re: Merge with git-pasky II.
http://www.gelato.unsw.edu.au/archives/git/0504/2153.html http://www.gelato.unsw.edu.au/archives/git/0504/2153.html
Where Linus says:
For example, it seems like most SCM people think that merging is about
getting the end result of two conflicting patches right.
In my opinion, that's the _least_ important part of a merge. Maybe the
kernel is very unusual in this, but basically true _conflicts_ are not
only rare, but they tend to be things you want a human to look at
regardless.
The important part of a merge is not how it handles conflicts (which need
to be verified by a human anyway if they are at all interesting), but that
it should meld the history together right so that you have a new solid
base for future merges.
In other words, the important part is the _trivial_ part: the naming of
the parents, and keeping track of their relationship. Not the clashes.
For example, CVS gets this part totally wrong. Sure, it can merge the
contents, but it totally ignores the important part, so once you've done a
merge, you're pretty much up shit creek wrt any subsequent merges in any
other direction. All the other CVS problems pale in comparison. Renames?
Just a detail.
And it looks like 99% of SCM people seem to think that the solution to
that is to be more clever about content merges. Which misses the point
entirely.
Don't get me wrong: content merges are nice, but they are _gravy_. They
are not important. You can do them manually if you have to. What's
important is that once you _have_ done them (manually or automatically),
the system had better be able to go on, knowing that they've been done.
- natep 14y agoI see that git has been updated to 'pass' the indent-block test, because it produces the correct output, but the resulting indentation is not correct. I have git.mergetool set to bc3 (Beyond Compare 3), so I tried running 'git mergetool' for each of the failed cases. In the adjacent case, bc3 merged things correctly, and all I had to do was accept its merge. In the indent-block case, I just had to fix (some of) the spaces, before accepting the merge. The only case where I had to do some real work was in dual-renames, but even then, it was fairly trivial. So, I agree with you. To me, it doesn't matter that git (or any other tool) sometimes gets content merging wrong. It _is_ gravy, and can be handled by other tools (bc3 in my case). What external tools can't do is manage your history.
- haberman 14y agoI agree with Linus that creating a solid merge base is far more important than clever merging. But I think there is still room for improving on Git in this respect. A lot of what I'm about to say is inspired by this video from the Camp guys: http://projects.haskell.org/camp/unique http://projects.haskell.org/camp/unique Git forces you to treat your history as a single linear sequence of commits. This is an unnecessary restriction if some of the changes in that sequence are totally independent of each other. For example, if two changes touch two completely different files and are unrelated, why should you be forced to sequence them in one order? Here is a practical situation illustrating this limitation. Once in a while I'll want to patch a coworker's in-progress change into my working directory where I also have changes. Perhaps I want to build a binary with several experimental in-progress changes in it. Suppose my coworker's changes are totally independent of mine (say they touch completely different files). I can do this in Git by applying his patch to my working directory (or doing a "git merge" with his branch). But now suppose I'm done with the experiment and want to back out my coworker's change, so my working directory is left with only my change. If I haven't made any more tweaks to my change in the meantime then I'm ok, I can just "git reset --hard HEAD^" to discard my coworker's change. But what if I made further changes to my change in the meantime? There's no easy way with Git to manipulate the two changes independently within the same branch, even though there are no actual dependencies between them. Sure you could create a separate branch for the merged thing. Every time you want to change your part you switch back to your branch, make the change, then switch back to the merged branch and merge again. But who wants to be that disciplined? Who should have to be that disciplined when the computer could do the work of knowing that the two lines of change are independent of each other? Git's ability to create stable and verifiable SHA1's is important, and I think that any future SCM will need to have this capability. But I don't think this implies that you have to treat the history in a strictly linear way. You could create SHA1 checkpoints when a particular person wants to publish and/or sign a tree and its contents, but still allow the individual commits to be treated in a more flexible way. The SHA1 checkpoints could be like barriers; each change is either part of the checkpoint or not, and the checkpoints could know their parent checkpoint(s) so that there is still a verifiable history available for auditing. I hope an approach like this could make large projects like Linux more intuitive to follow. I always found it unfortunate that the graph of commits for any project with lots of merge activity is totally indecipherable. For example, here is a screenshot of Git's own Git repository: http://i.imgur.com/RyQm3.png http://i.imgur.com/RyQm3.png If independent changes could be viewed independently, and if every merge didn't have to be an explicit commit, perhaps this could be easier to follow. There are definitely lots of unanswered questions here and I don't claim to have all the answers. My point is just that I don't think Git is necessarily the last word in distributed version control.