10 ms·
Resistance Against Git Merge Hell (2015)
- stevage 2y agoSo why can't the tools that display history just filter out all the merge commits? I don't really understand why this requires merges to be done differently at the time.
- adityaathalye 2y agoRebasing-only is a manual way to "linearise history". Whereas, a git log is a query target. And the `git-log` tool is one's friend. I bash-alias these for quick use: git log --oneline # `gl`, to see "linear" history with branch annotations git log --oneline --merges # `glm`, to see merge commits only git log --oneline --graph # `glg` see the train tracks too, if I need to And I know that git log has me covered, if/when I need to narrow / slice history more.
- fargle 2y agothe way i look at it is this: we actually use git for at least two different but related things - a public change history (e.g. commit history of master, dev-1.0, whatever) - a personal change history, especially useful when prototyping a fix or feature on some independent branch. you might even call this independent branch "master" but it's your personal version of master in a clone on your laptop. it's a time machine, you can go backward/forward. you can make a change and then reverse course. you can commit every work-in-progress that compiles as you work get it running or passing tests. whatever you like. you can pull or merge or rebase as you wish. when submitting it to the public "master" branch, you are effectively saying "i want to commit this package of changes on top of the master branch. do you really want to see "fhgiry pulled from master into fix-splork-feature on date x" and https://xkcd.com/1597/ https://xkcd.com/1597/ in your public history of master? every branch you merged from or rebased on and every failed attempt and spelling mistake and WIP and addressing of review comments? or do we want to see a series of commits that just record the end result in a series of nice little patches that simply add a finished changes like: "splork: fix canoe paddle explosion issue by decreasing default gamma" the common ubiquitous guidance to "never rewrite history" is nearly always valid for the public history. converting a messy personal history into a neat series of patches/changes/commits that apply cleanly onto the latest public master should not fall under that same guidance. i'd say it's closer to "never rewrite public history"
- Charon77 2y agoThere's this handy git config 'pull.rebase' in case you miss it. A lot of the merge commit was triggered when you pull and merge instead of rebasing locally. https://git-scm.com/book/en/v2/Git-Branching-Rebasing https://git-scm.com/book/en/v2/Git-Branching-Rebasing
- cess11 2y agoI have never seen something like the example consisting of pretty much only merges, and most places where I've worked didn't use rebase or squash or whatever it's called. It has also been uncommon with 'fix' and 'did a thing' commit messages, usually people type in what they actually did, except in one code base but that was someone who really didn't enjoy software development.
- giantg2 2y agoMost merge issues can be avoided with proper architecture and planning. If your files are so large that you have multiple devs working in the same file at the same time on a regular basis, then it's probably time to get more modular. I honestly hate Git compared to SVN. The Git tool seems technically better. However, the easier aspects of the tool seems to have promoted less planning, more stepping on each other's toes, and what I see as non-value added work (merges, rebaseing, retesting, etc).
- srvaroa 2y agoIf a pizza team struggles to work in small independent changes, and end up having to deal with long lived branches, merge festivals etc. then the problem is really not in the tool chain.
- giantg2 2y agoYeah, it just seems that the tool chain has made it easier to fall into those poor practices. But I probably just work for a really shitty company.
- deleted 2y ago[deleted]
- juped 2y agoWhat a disjointed writeup of a terrible idea. We only use git _because_ its history is a graph rather than a sequence, freeing us from the necessity to ever acquire locks on fellow humans. There's a reason git.git never does anything like this, despite the history being completely constructed by Junio Hamano (who applies patches with git am, rather than taking a set of commits from someone's remote as is typical on Github). A clean history full of useful information looks like git.git's history: nonlinear.
- jFriedensreich 2y agoTo be able to work with large repos i found a few more important strategies: - in your git client: enable "only follow first parent" for the base branch which only shows merge commits and only open merge commit history as needed/ when relevant, i'm only aware of fork.app doing this nicely but should be an option in more git clients - hide all branches by default and only show your base branch and your own work, just let go of caring what others are doing, the only way to interact with other branches should be a review system at which point you can selectively pull in the relevant commits as needed to try locally. (exception is maybe a technical manager or team lead who should of course care somewhat if branches get abandoned or not cleaned up properly, but this is a completely separate workflow) - Look at saplings notion of public vs draft commits, this is a game changer and works just as well for git, just not supported by the tooling in the same way. Don't try to argue and adopt workflows that are the same for both, they are completely different phases of work with completely different needs. In a nutshell: a) public commits are the ones that were pushed to a branch that is shared with anyone else. This maybe main, a feature branch worked on by multiple colleagues etc. These commits are considered immutable, never amended or rebased so the only option to bring them up to date with another branch is a merge. b) BUT every commit that is not part of one of these branches is considered a "draft" these are never merged into, always amended or modified and mutable draft state. It is considered rude to expose your internal working struggles like merging in master 100 times a day to avoid big merge conflict buildup or reverting changes into your work submitted to review, you submit a clean stack of pull requests, one per reviewable batch of changes - rebasing and amending has a few other big advantages. One of them is that syncing to the base branch does not pollute history and can be done as often as possible without a down side. This allows keeping all work in sync without buildup of big hard to resolve conflicts. There is also tooling to do this for all your open work at once and to resolve conflicts smarter (.eg git rerere) - merges should be set to allow only squash + rebase as this works the same for devs working with clean commits as well as devs keeping their commit history in their branches. this way at least the base branch is mostly clean
- Ferret7446 2y agoVery few Git users use a true distributed workflow that merge is designed for; naturally, most Git users would have no use for merge. Merge is useful when there are multiple "canonical" repos which occasionally merge changes in between each other. Think web of trust vs central authority.
- Aethelwulf 2y agoDo people really look at their git commit history like this? Why?
- sschueller 2y agoThe map? Sometimes it helps to get a visual queue (at least for me) but most of the time I don't need it.
- Snild 2y agoYes. To figure out what happened, and which path it took to get into my repo. Or to see how branches have diverged. I also recommend it to git newbies, as a tool to understand the state of their repo when something has gone wrong (e.g. they did a bad merge or rebase).
- kaffekaka 2y agoBut with rebase, the history does not show what happened. It shows the changes to the code yes, but not how they came into being.
- arrowsmith 2y agoWhy would I care how the changes came into being?
- nicoburns 2y agoIf you only rebase feature branches before merging to main/master/trunk then you still get most of the history.
- neallindsay 2y agoPeople often bring up this objection, but I don't want a complete history—I want an easy-to-understand history. If I make a variable in commit A and think of a better name for it in commit C, why wouldn't I use rebase to squash C into A? Some sense of purity of history? Or more directly related to the rebase vs. merge debate, if I fix an issue across three functions, and in the meantime someone has removed one of those functions on the main branch, rebasing eliminates the "history" of me fixing that removed function and I think that's good. It makes my commit simpler. Rebase can certainly be used to simplify the history too much, but that will always be a judgement call. That shouldn't keep us from editing our branches in ways that are clarifying instead of confusing. We never capture the full complexity of what we went through writing code in our source control. It would be bad if we did.
- wobblyasp 2y agoClean is such a weasel word. Does it matter if you visual tooling produces a series of squiggly lines? The information is all there. You can grep it to find specific commits. This isn't a hill worth dying on with your team if it's their MO already.
- ClumsyPilot 2y agoExactly - it’s subjective - detailed vs clean - if you want every detail of reality, it’s not clean, it’s messy. Ever seen inside of an engine? It’s greasy, dirty, etc. a clean engine doesn’t run. This is putting lipstick on a pig Which can be fine, but let’s not pretend there is a perfect approach, needs differ
- IshKebab 2y agoThe information is all there but the point is that it's an unreadable mess.
- stevage 2y agoTo me it is weird to talk about readability when not talking about a specific tool. Can a given tool not make a history more readable by filtering out merge commits?
- tilsammans 2y agoThis doesn't look dirty or confusing to me at all.
- hardlianotion 2y agoOh. I wanted this to be about the London Tube Map.
- jfengel 2y agoThe Tube could definitely stand for some rebasing. Its history is rather convoluted. Also removal of code smells. And actual smells. /I love the Tube. So damned useful.
- madeofpalk 2y agoOf course, "What went wrong with the Tube Map?" from Jay Foreman https://www.youtube.com/watch?v=jaEhvWXmLyk https://www.youtube.com/watch?v=jaEhvWXmLyk
- hardlianotion 2y agoMuch better.
- pjc50 2y agoThis is a very short advert for rebasing your personal branch. People seem to have very strong opinions about this. I don't much, either way; I have to use the gerrit flow at work, which mandates a lot of rebasing, and in general I prefer not having a commit which only records a merge. Some people seem to want highly detailed tracking of what an individual developer has done in their personal checkout. I will only note that the act of observing changes what is observed.
- klyrs 2y agoI'm an ADHD weirdo who will mix six issues in my personal branch, and use rebase, cherry-pick, and in a pinch, difftool to carve out clean patches when something is good and ready. If I were to give people an honest history of what was actually going on in my repo, and I used merge and revert to really keep it all when I cut a patch, it would be a forensic nightmare. There is some benefit to my tendency to do experimental work in a branch where I've started to work on something else. In that mudpit, I find serendipitous solutions to multiple problems, and I find incompatibilities between desired changes. And when multiple issues have overlapping changes, there's usually an optimal ordering of which to address first -- and that's not always obvious. I greatly prefer to do all the forensic work up-front, to make a clean history with a cogent story in the commit message. When I see other people's messy histories with every merge/revert, I need to do that forensic work every time I go back in history; without the benefit of recent first-person experience.
- juped 2y agoSame. But there's a difference between a few separate things here: - assembling your own chaotic actual work into cogent commits using interactive rebase and friends (but not moving the base, e.g. git rebase -i --keep-base): good and necessary. people doing reviews should also review commit history, not just a megadiff of all changes. - rebasing in the sense of "moving the base upon which your work is based", which for unfortunate historical reasons uses the same command: this is rarely ever useful, despite how much people seem to enjoy doing it. replace this with a no-op in nearly all cases; and in the few cases where you might want to (the upstream you want to integrate with changed incompatibly) merging a recent version tag into your topic is better. - rebasing in the sense of "as a maintainer adding completed work to an integration branch like master, rebasing the commits atop the integration branch rather than simply merging": this is actively harmful (as is "squash merging", which is this _plus_ destroying the commits crafted in point 1) and is also unfortunately what most people seem to mean.
- peanut-walrus 2y agoYou squash merge when stuff gets added to master (or any other shared/long-lived branch) and delete the development branch, nobody has absolutely any interest in what happened on your development branch. There you go, clean commit history if you care about that sort of thing. Rebase workflows are awful and unintuitive. Leave the rebasing to the git wizards who actually know what they're doing, in no circumstances should this be part of your day-to-day work.
- madeofpalk 2y agoThe beauty of this is that the rebase wizards are free to rebase all they want in their feature branches, but everything is squash-merged ito main and no one has to know.
- tome 2y agoBut one reason that rebase wizards might curate a a rebased branch of small commits is exactly so that the small commit structure remains on merge to main. That way they can track down any problems with it more easily in the future. (Ironically, I've found that this style of development makes it less likely for bugs to be introduced in the first place.)
- lolinder 2y agoThis is terrible for "rebase wizards" because about 80% of the reason why I rebase is to make my commits useful when someone inevitably needs to do a git blame in 5 years to understand the history of this code. A squashed commit can tell them "this was part of this feature", but a well-crafted history plus a merge commit with the PR name and number can tell them that and show what specific code changes were related to one another in one atomic step towards a feature.
- Faaak 2y agoWhy would you squash merge if you have two different atomic commits ? Makes bisecting + reverting a pain... I'd avocate for a rebase instead
- loloquwowndueo 2y ago
- planede 2y agoproblem statement: git history, as normally presented, is hard to follow and contains too much noise. article's proposed solution: simplify the git history by destroying information that are irrelevant. That's what rebase is. The problem is what is or isn't relevant depends on context. I think the right way to go about it is to simplify the presentation and otherwise improve tooling to get information out of the git history. git already has some ways to filter the history, but it lacks very feature rich query language, like mercurial. git guis should also step their game up.
- robertlagrant 2y agoI agree. I'd prefer a messy history but have a way to just see MRs into master/main as the main sequence of changes. Then be able to zoom in futher if necessary to see how an MR was arrived at.
- planede 2y ago> I'd prefer a messy history but have a way to just see MRs into master/main as the main sequence of changes. git log --first-parent gets you there. > Then be able to zoom in futher if necessary to see how an MR was arrived at. Yeah, an interactive UI would be nice for this, maybe there are some, but I really only use the CLI.
- INTPenis 2y agoI was guilty of merge hell until I learned about rebase. I think it's a natural instinct to want to protect your production environment so I'm sure others have made the same mistake as I did. Main is my default branch, I don't want anything pushed into the default branch to result in a deployment to production. So naturally I create a new branch called production. The relationship makes sense to me because we develop in main, might even deploy to staging from main, but it's not until we feel we're ready that we merge main with production. And to a user of git this requires a manual step where you explicitly specify the word production, git checkout production. But that resulted in merge hell until I learned I could just rebase main onto production. And the command structure is even exactly the same, standing in production branch I either do git merge main, or git rebase main. This comment was a message to my younger self.
- 01HNNWZ0MV43FF 2y agoYour comment says 6 minutes ago but when I go to reply it says three days... And I do remember reading it three days ago Does hacker News fudge timestamps on comments when it boosts a post?
- Uvix 2y agoProduction shouldn't have its own branch - it shouldn't even be a new build at all. The same binaries that were deployed to staging should be redeployed to production once you approved. Otherwise, your "staging" environment has the wrong name... and the testing there doesn't represent what goes to production.
- toast0 2y agoIn a system with production branches and staging, you build for staging from the production branch. When you're in the release process, the branch won't match what's currently deployed, but that's ok. The point of a production branch is not to indicate what is on production at this instant, but to be a record of what was deployed to production or at least was intended to be, for changes that get canceled before deployment.
- larsnystrom 2y agoThe more I work with git, the more I wish there was a rebase (including squash/fixup) which kept the original commits, but hides them. I’m not sure how that would work in practice, but there is value in keeping all change history, and there is also value in having a readable commit history, but git does not let you do both.
- OJFord 2y agoWork with it a bit more, discover reflog, and you'll find that's exactly what happens (until gc) ;) It could be a bit more visible somehow though, I get the sentiment. Maybe it's more of an add-on to git's role though, at least without plenty else also becoming more visible/GUI-like too.
- dieortin 2y agoIf it’s only saved until gc then it isn’t something you can rely on
- dllthomas 2y agoIt's saved while it's in the reflog, and then saved until gc. How long things stay in the reflog is configurable. The bigger deal is that things are never shared simply for being in the reflog - which is probably correct for its intended use but doesn't really fit what's asked for up thread.
- tome 2y ago> I wish there was a rebase (including squash/fixup) which kept the original commits, but hides them That's exactly what rebase does. (OJFord said that too, but buried the lede slightly, so I thought it worth saying in a single sentence.)
- GrantMoyer 2y agoThat's effectively what merge does. If you want to think of your branch as one linear history, then git merge creates a single commit in your branch which represents a collection of commits from some other space (and there's no rule that you need to use the default merge commit message). Then you use `git log --first-parent` to view your branch's simple linear history.
- gsliepen 2y agoIdeally, your main branch is always in a known good state (every commit results in compilable code that passes the tests you have so foar), which makes it bisectable. Keeping all commit (including broken ones and subsequent fixes) in topic branches and then merging them without some squashing and rebasing gets in the way of that.
- chx 2y agogit bisect is extremely, extremely powerful. I have found an almost security hole with it in the dropping of a BC layer which just no one would've expected to do that. And because it was all deleted code tracking down the origin bug any other way is fair impossible.
- dllthomas 2y ago> Keeping all commit (including broken ones and subsequent fixes) in topic branches and then merging them without some squashing and rebasing gets in the way of that. It seems to me that, with respect to bisecting, a merge workflow with git bisect --first-parent is equivalent to a squash workflow with a bare git bisect. Am I missing some way in which that's not the case?
- gsliepen 2y agoYou might not want to squash everything into one commit before merging, you can still have multiple commits in one (fast-forward) merge, as long as each of them is in a good state. This is made relatively easy by using the `--autosquash` feature of `git rebase`. One issue I've seen a few times is that some commit in the middle of a topic branch is the problem, but if you didn't rebase it then that commit itself would look fine on top of the topic branch's parent. However, after merging it's now also on top of other commits, and the interaction with those was the problem. That makes it very hard to find such a problem. Rebasing the history of the topic branch before merging will make finding it much easier.
- dllthomas 2y agoAh, I wasn't trying to say those were the only two options, or that either was what you were suggesting. I also typically prefer other things. I just find that "always squash the entire branch" is a common reaction to "history is messy" and I wanted to surface that (per my understanding) it doesn't actually improve the situation (vis-a-vis bisect in particular, assuming you're passing the correct arguments for your situation) over merging (no-ff, I neglected to specify...) branches where some of the commits do not build.
- IshKebab 2y agoI definitely agree with this. If you merge to this extent then you've basically given up on having an understandable history. Rebase is much better where possible. I think a lot of the disagreement about this is really people talking about different things. Some people say "always squash; nobody cares about the trivial typo fix commits and whatnot" and other people say "never squash; you lose important history" and really they're both right... you should squash when it's not important to preserve the history. Obviously people are going to disagree about when that is but in my experience if a PR is big enough that you think it should be more than one commit then it's too big, unless it's a big feature branch that has been worked on by multiple authors. Similarly with rebase vs merge, if it's a small single author PR then definitely rebase. For big feature branches you may want to use merges though I would still suggest rebase is better. You just need to make sure everyone is using the safety flags when they force push.
- echelon_musk 2y ago> end up with git commit history which looks like London tube map I've saved you a click. TFA has nothing to do with TFL beyond this line.
- breckenedge 2y agoSurprised to not see a mention of the —no-merges flag.