9 ms·
Git is Inconsistent
- etherealG 15y agoI agree with you completely, but want to know how this can be fixed in git? Surely there has to be something about the merging algorithm that can be changed to fix this, and if that's the case we can just patch it and move on. What is the specific problem with the algorithm that causes this?
- pmjordan 15y agoI assume this will reduce the quality of the merge algorithm from a stand-alone point of view, which is presumably a very hard sell.
- etherealG 15y agoI don't know this is true for sure, perhaps introducing the patch would increase it's quality. If someone offered such a patch we could discuss, instead the article only shows the broken test case. It's almost a darcs plug without and reasoning.
- etherealG 15y agosee the link below posted by tonfa, seems this patch isn't worth it anyway :)
- jojo1 15y agoHmmm, nobody seems to care: http://article.gmane.org/gmane.comp.version-control.git/105748/ http://article.gmane.org/gmane.comp.version-control.git/1057...
- daviddavis 15y agoI wonder how mercurial compares in this aspect. Also, I'll keep using git because for sure, it's a helluva lot better than SVN or CVS (which my company was using when I got there).
- tonfa 15y agoSame as git, and you'll probably get the same reactions. """ In other words, we're already at the point of significantly diminished, possibly negative returns on effort. The last few percent will always require some level of human-equivalent intelligence. I think effort here is much better spent elsewhere, like researching general AI or playing on waterslides. """ http://thread.gmane.org/gmane.comp.version-control.mercurial.general/26109/focus=26110 http://thread.gmane.org/gmane.comp.version-control.mercurial...
- anonymoushn 15y agoThe least we can do is detect a situation like this and make it a conflict.
- etherealG 15y agoThanks so much for this link, this is exactly the kind of analysis I was hoping for. Clearly this is all a bit FUD, and darcs which gets this right, is trying too hard. I wonder how fast the general merge algo that darcs is using to get this right is? <trollface>
- tonfa 15y agoMatt's point is that while some algorithms will fix this particular case, you can still come up with a different edge case which makes it break. The whole "prefect merge tool" was very popular five years ago (during git's and mercurial's infancy), but it didn't lead anywhere. Simple merges strategy are "good enough" in practice.
- ob 15y agoMatt's point is that they've chosen a system that makes it really hard to get that last 10%. "We have tried to draw spirals using cartesian coordinates, what we have gets us 90% there, but there are infinities and edge cases involved in getting a perfect spiral. The equations describing them would get so complicated it's just not worth it." What we have in BitKeeper is the equivalent of polar coordinates... it makes drawing spirals much, much easier ;)
- yuvadam 15y agoTo quote Johannes Schindelin [1] : This all just proves again that there can be no perfect merge strategy; you'll always have to verify that the right thing was done. [1] - http://thread.gmane.org/gmane.comp.version-control.git/105748 http://thread.gmane.org/gmane.comp.version-control.git/10574...
- Andys 15y agoAmen. There's no way I ever do this in a real code base without checking that the result is what I intended.
- ams6110 15y agoYes, I also always look at diffs after a merge, and also before I commit. Several times I've caught changes that I really didn't want to go back to the branch.
- lambda 15y agoIt's obvious that there is no perfect merge strategy. There will always be ambiguous cases, cases in which the merge algorithm doesn't have enough information to make an informed decision, or cases in which there are changes that effect lines not caught by doing line-by-line diffs. I think that the point that Zooko and Russel O'Connor are making is that there are cases in which the merge algorithm does have available to it the information necessary to make a better decision (that is, the entire history of changes, rather than just just the two changes being merged and their common ancestor), but in Git, that information isn't being taken into account. While you are never going to have a perfect merge strategy, the argument is that you can have one that is better. Some people, however, feel that the Git algorithm is good enough, and doing it the Darcs way would be slower without much benefit other than for fairly artificial examples (you have to be doing something where you move a block of code, and then re-introduce that same block back in the original place on one side of the merge, while patching that block on the other side of the merge). Personally, I've found Git's merge strategy adequate for everything I've used it for. Git has support for multiple merge strategies, so if someone wanted to implement a better but slower one as an opt-in, they could do so.
- tytso 15y agoI've contributed a tiny amount to git (the high-level "git mergetool") so I can't speak for all of the git developers, but I've spent enough time hanging around for them to say that the general feeling they have is that git's algorithm which is "3-way merge, and then look at the intervening commits to fix any merge conflicts" is good enough. You can always try to spend more time trying to use more data, or deducing more semantic information, but past a certain point, it's what Linus Torvalds has called "mental masterburation". For example, you could try to create an algorithm that notices that in branch A a method function has been renamed, and in branch B, a call to that method function was introduced, and when you merge A and B, it will also automatically rename the method function invocation that was added in branch B. That might be closer to "doing the right thing". But does it matter? In practice, a quick trial compile check of the sources before you finalize the merge will solve the problem, and that way you don't have to start adding language-specific semantic parsers for C++, Java, etc. So just because something could be done to make merges smarter, doesn't mean that it should be done. It's a similar case going on here. Yes, if you prepend and postpend identical text, a 3-way merge can get confused. And since git doesn't invoke its extra resolution magic unless the merge fails, the "wrong" result, at least according to the darcs folks, can happen. But the reason why git has chosen this result is that Linus wanted merges to be fast. If you have to examine every single intermediate node to figure out what might be going on, merges would become much slower, since in real life there will be many, many more intermediate nodes that darcs would have to analyze. Given that this situation doesn't happen much in real life (not withstanding SCM geeks who spend all day dreaming up artificial merge scenarios), it's considered a worthwhile tradeoff.
- adambyrtek 15y agoGood point, and another argument for maintaining a reasonable test coverage. I'd even argue that a merge strategy that is too clever (like the one you described) could be more risky than a dumb one. It could lead to a resolution that is valid from a compiler standpoint, but wrong semantically, which makes it even harder to discover.
- scorpion032 15y ago"Make things as simple as possible; no simpler" ~ Albert Einstein.
- davidmathers 15y agoHere's the short version: I am the original sentence. Alice commits a change in her repo: I am a different sentence. Bob commits a change in his repo: I am the original sentence. I am the original sentence. Now Alice pulls Bob's commit. What should happen? The argument is that in certain cases it can be known which of Bob's 2 sentences is the original and which is the copy (due to context provided by an intermediate commit) and that therefore a correct VCS will figure out that the original is on the bottom: I am the original sentence. I am a different sentence. But git doesn't look at history so will always produce: I am a different sentence. I am the original sentence. I don't care. If you force me to care then I actually prefer git's behavior. Git is consistent: a merge will always produce the same result for the same files. I don't want history to matter. The problem is not actually solvable. So git doesn't try to solve it. I think that's why it's called "the stupid content tracker." EDIT: Is there anything worse than "smart" features that only work, say, 80% of the time? The closer they get to 100% the worse it gets, because then you start relying on them and they break right when you stop paying attention.
- Peaker 15y ago> Git is consistent: a merge will always produce the same result for the same files I thought the point was that if you pull the exact same commits in different order the merge will produce a different result for the same files, meaning that in git the history does matter. Whereas darcs/etc will always produce the same result, such that history does not matter?
- riffraff 15y agomore precisely, I believe, the history of the merges counts, not the history of the files per their original edits.
- davidmathers 15y agopull the exact same commits in different order Sort of. The OP doesn't write clearly. He's also confused about how git works. What he means is.. Say Bob has 2 commits (B1-B2) and Alice has 1 (A1) Scenario 1: Alice merges each of Bob's commits in sequence (i.e. she replays his commit history onto her repo: A1-B1-B2). Scenario 2: Alice merges only B2 (A1-B2). The point is that, with git, Alice's repo will be different in each scenario. Because in scenario 2 git doesn't examine commit B1 and use that info to try and figure out what the content in commit B2 "means". With darcs, on the other hand, her scenario 1 repo will be identical to her scenario 2 repo. The flip side is that in scenario 2 git will always produce the same result for the same B2, because B1 is irrelevant. With darcs a change in B1 will change the result. NOTE: "git pull --rebase" actually does "replay commit history" instead of "merge" when pulling code into your repo (result: B1-B2-A1). I use it as my default. The outcome is the same as darcs, the difference is that everything is explicit.
- mml 15y agohmm. i was hoping the article discussed git's mind-bogglingly horrible user interface. can't have everything i guess.
- hasenj 15y agogit's UI is great; as long as you understand how it works. The good thing is: "how it works" is really simple. You should treat it like a language (just like all system/unix tools), not an "app".
- Peaker 15y agoI think git is one of the best tools we have, but its UI is really bad: checkout and reset do completely different things when given files or when not given files. reset on files should really have been called unadd. reset on refspecs should really have been jumpto, moveto or something else indicative that the current branch ptr is moved to a new refspec. --soft and friends could have been --no-update-index or --no-update-files. checkout on files should really have been called overwrite. checkout on branch names should have probably been switch, setcurrentbranch or a name indicative that the current branch is being changed. pull and push are symmetric names for asymmetric behavior. pull could have been a flag for merge (-f meaning fetch first). reset --hard was for a long time the only way to move a branch ptr to a new position along with the files, but it has the potentially unintended consequence of also irreversibly deleting working tree changes. If you use it to delete, that's fine, but since you had to use it to move the branch ptr, it is simply wrong to have irreversible damage as a side effect. Especially in an RCS which is used by many as the fail-safe against their own user mistakes. There's no easy way to see which branches are tracking what. And until recently it was a big PITA to even make the current branch track a remote branch. Deleting remote branches has awkward syntax (pushing an empty string to a branch name) and then you have to use a specialized command (remote prune) if you want the deletion to be propagated to other repositories. Another annoyance: Git doesn't let you push a detached head to a new remote branch, so you have to create a temp branch ptr to the detached head position and later delete it. Git also doesn't have good support for versioned sub-projects. submodule is sub-par, and requires a multitude of extra commands even in the cases that should have been seamless.
- mebigfatguy 15y agoI agree that i can live with it as others had said, but it would be interesting to know how the vcses that apparently resolve this issue, actually resolve it.
- nevinera 15y ago>There are still some people who still think nothing is wrong with git; that it is okay for the result of a merge to depend on how things are merged rather than on only what is merged; that is it okay for two git repositories that pull the same patches to have different contents depending on how they pulled those patches. I don’t know what to say to those people. Such a view seems like insanity to me. Git merges files, not file-histories. Git's behavior is simple, clear, and easy to understand. I can see why you might expect merges to be transitive like this (it would be an elegant property, if it were true), but why does it matter to you? In what way do you use merges that could rely on this expectation?
- JoeAltmaier 15y agoThere are so many theories of "source control" that none of them are "simple clear and easy". They take study, and if you start from a different place, a new paradigm will be hard to learn and internalize. That said: An elegant property? Are you kidding? That is intrinsic to most tools that dare call themselves "source control". Git requires extraordinary explanation if it behaves in an extraordinary fashion.
- nevinera 15y ago>That is intrinsic to most tools that dare call themselves "source control". Bullshit. 'Merge' is one of the most complicated operations in every versioning system. I'm pretty confident svn is 'inconsistent'. Or is that too niche?
- JoeAltmaier 15y agoMerge is a tool. The issue is, can you reproduce source accurately, from a variety of starting points, and be sure you have some canonical thing (release x.y). Do I understand you right? This is not the expectation for a source control tool?
- nevinera 15y agoYou can reproduce any repository state that you (or anyone else) have stored. You can do it simply, reliably, and quickly. That is not in doubt. Merge is not a tool for reproducing a canonical state, it's a tool for combining two or more of them, an entirely different topic. Any other straw men you'd like to hold up real fast?
- saalweachter 15y agoIs there any reason to assume that merges should be associative? Hell, of the four normed division algebras, only three are associative; just because you can say "operations on octonions should be associative" doesn't mean that you can necessarily create a system of octonions where it's true. For what it's worth, "git pull --rebase" does enforce a specific order to changes (local changes always happen after remote changes) which will produce the same results regardless of when user Bob pulls user Charlie's changes: regardless of whether Bob pulls change c1 after commiting both b1 and b2 or after commiting b1 and before commiting b2, the final commit order will always be "a, c1, b1, b2". Of course, if Bob commits and pushes b1 before Charlie commits and pushes c1, the final commit order will be "a, b1, c1, b2", but how could it ever be otherwise?
- pjscott 15y agoThere are ways of making a DVCS that allow all merges to be associative, and all patches commutative except when there's a causal dependency between them, e.g. if patch A creates a file, and patch B edits that file, then they cannot commute. I believe darcs makes these guarantees, and making a correct implementation is relatively straightforward. (Making it fast is more complicated, but definitely doable.) Ultimately, though, what you really want is for the VCS to just do what you mean. That's a lot trickier than providing mathematical guarantees about patch reordering and convergence.
- gnosis 15y agoDoes anyone know how bazaar would handle this?
- tonfa 15y agoJust try or check the source. If they use patience or some kind of cdv merge, I expect they would get the same merge in both direction.
- KirinDave 15y agoNot to be grumpy about it, but git's shortcomings are well-known and most people don't run into them on a daily basis. Some DVCS, like Darcs, might behave better, but they all seem almost comically slow even for medium-sized repos. If I have to sacrifice git's speed for certain types of correctness (that don't trouble me on a daily basis), I will be VERY reluctant to make that choice.
- Groxx 15y agoSuper-simple-summary: Git doesn't use history to determine merge behavior (edit: in this circumstance). Git behaves like applying patches. Darcs uses the history to make "intelligent" patches. It's a matter of taste. If you look at Git as having a history, therefore should use the history, yes, it's incorrect. But if you look at it as a patch manager, it's behaving as it should, and Darcs is frighteningly unpredictable - the numbers on the patch might not match the numbers of the lines it modifies. I side with Git on this. I can generate patches from Git that will work anywhere, and use them 100% identically within Git as manually applying them. The same cannot be said for Darcs.
- ob 15y agoOf course Git uses history. It doesn't _have_ to, but it does. As a matter of fact, as soon as you use diff3, you are using history (that's where the GCA comes from).
- Groxx 15y agoKnow which situations it does use it, similar to this setup? Apparently not for moves, any other potential gotchas? I prefer patch-like behavior, because it can be predicted by looking at the patch.
- tzs 15y agoThe article mentions that some systems do have the associativity property--that is, extra rungs in the merge ladder do not affect the result. I can see how that can be achieved in the case of fully automatic merges. When merging B2 into C1+B1, you'd effectively un-merge C1+B1, merge B1 and B2, and then merge C1 and B1+B2. But how would that work if C1+B1 had a conflict that had to be manually resolved? Assuming merging B1+B2 into C1 has the same problem (a fair assumption) will I have to do the same manual fixes again? Or are they smart enough to look at the failed automatic C1+B1 merge, and generate a patch to that from the manual fixes I did, and then try to use those to resolve the merge of C1 and B1+B2? I suspect there will be cases where this is just not going to work well.
- ob 15y agoThere are two things most commenters in this thread have missed: 1) The article talks about auto-merges. If the code is "too close" by some definition of close, you get a conflict that needs to be manually merged. The article does NOT talk about manual merges. 2) The article is titled "Git is Inconsistent", it doesn't claim Git is WRONG, it claims it is INCONSISTENT. It does different things depending on how you merge and when. I think consistency in a DVCS is a desirable goal. It should not matter whether you pull A then B, or pull B then A, or whether given a series of commits, you pull after each one, or just once at the end. The end result should be the same. That it is a rare occurrence only makes it worse. You will mostly trust the auto-merge algorithm until you hit the corner case and it will be very expensive in terms of time/money to fix the mistake. Git's brilliance/stupidity is precisely that it only tracks contents, so although it could get the right answer it makes it very expensive to do it.
- davidmathers 15y agoThe article is titled "Git is Inconsistent", it doesn't claim Git is WRONG, it claims it is INCONSISTENT. Ok. The claim that git is inconsistent is wrong. From OP: The problem with git’s merging is that it doesn’t satisfy the “merge associativity law” which states that merging change A into a branch followed by merging change B into the branch gives the same results as merging both changes in together in one merge. There is no such concept in git as "merging both changes in together in one merge". I have modified a shell script written by Simon Marlow that illustrates, using git, how merging two patches separately can give different results than merging two patches together. The shell script doesn't do what is claimed. It can't because git has no facility for "merging two patches together". Git can only do 2 things with patches: 1. generate a patch 2. apply a patch But! git has a function which is equivalent to combining 2 patches in a single merge: git pull --rebase The shell script does not use this command. It first applies 2 patches separately. It then applies 1 patch separately. There are still some people who still think nothing is wrong with git; that it is okay for the result of a merge to depend on how things are merged rather than on only what is merged; that is it okay for two git repositories that pull the same patches to have different contents depending on how they pulled those patches. I don’t know what to say to those people. This is just incoherent. I have no idea what to say in response because I have no idea what the intended meaning is.
- __david__ 15y agoAfter reading this it strikes me that git is imperative--it stores files as they were when you checked them in and merges what you tell it in the order you tell it. Darcs, however, is more declarative--it stores patches. And not just patches but patches with dependencies. This set of patches describes how the current state of the repository is constructed. So when you merge you're really just adding new patches to the repo and it knows exactly what to do to make it work. The interesting thing is that git has all the information there... It could go through the relevant history, diff everything and put the resulting patches in a darcs-like data structure and then commute patches with darcs' patch theory. But in the end I'm not sure I'm ready to call darcs' style right and git's wrong. Both of them have a fairly easy to understand object models and they both have merges that act in accordance to the internal philosophies of those object models.
- dmoney 15y agoOff topic, but the link to the shell script and the images in the article use Data URIs, which I didn't know existed: http://en.wikipedia.org/wiki/Data_URI_scheme http://en.wikipedia.org/wiki/Data_URI_scheme