7 ms·
The reason to squash commits is more than just keeping your commit history read-able, it's about making easy to revert a feature and being able to keep history
by davewritescode 6y ago
The reason to squash commits is more than just keeping your commit history read-able, it's about making easy to revert a feature and being able to keep history in a way that makes it simple to revert a change if you run into issues.
If I rollout a rewrite of an endpoint and run into a weird issue in the QA environment, I'm a simple git revert away from fixing the issue. If I had spread that endpoint across 25 commits, I'd have to actually debug the issue in QA and figure out what I broke. Reverting quickly lets me debug the issue on my time instead of keeping our test suite broken.
It's not always feasible, and I try not to be a stickler when people on my team don't do it but the fact is, if you're working in a world when you're delivery code quickly into real environments, having a back-out strategy is paramount.
- twic 6y ago> If I had spread that endpoint across 25 commits, I'd have to actually debug the issue in QA and figure out what I broke No, you could just revert the entire lot in one go.
- Mazzen 6y agothat's not a valid argument. How do you identify "the lot"?
- magicalhippo 6y agoBy the ticket/issue number?
- deleted 6y ago[deleted]
- gmfawcett 6y agoThere's an obvious upper bound. If you can't identify the last in-prod version that worked, and the first in-prod version that failed, then you have bigger problems than deciding whether to squash your commits.
- twic 6y agoBy story number, by author, by timestamp, by looking for merge commits, by reading the commit messages - by common sense, basically. If someone has just pushed broken code to master, they usually know exactly what they've just pushed. This doesn't seem like a real problem to me.
- falcolas 6y agolast_prod_tag < the_lot < current_prod_tag
- blandflakes 6y agoAnd if the merge of their commits interleaves with other commits from... teammates?
- falcolas 6y agoThanks to the complexity of how components integrate, even seemingly disjointed components, their commits are going to matter when it comes to troubleshooting an issue. This gets a bit more murkey with mono-repos, but even microservices can combine to create complex production issues.
- blandflakes 6y agoThis sort of sidesteps the issue, IMO, of how hard it is to isolate a problem. Yeah sometimes things collide, but pretty often in my experience there has been one change that needs reverting or patching, not six changes that interleaved. Or rather, interleaving them at best increases the difficulty of the blame game.
- alexmingoia 6y agoProgrammers can have the best of both worlds. Use granular commits on a local branch and squash merge into shared branches. That way one gets clean shared history while preserving local work history.
- giancarlostoro 6y agoI would equally argue there is nothing wrong with the local work history branches being remotely available. If I have multiple branches for one ticket nobody ever seems to care, they just care about which branch is in the PR. Besides, if you have a reasonable web UI for git, it shows it all merged together as one big changeset.
- dastx 6y ago> squash merge into shared branches Why not just rely on a merge commit instead?
- CrunchyTaco 6y agoMy gripe with merge commits is they don't integrate nicely with `git blame`. If I'm looking through historic commits (to understand why a change was made, or perhaps to debug an issue) I'll often `git blame` the line and diff that commit. If the commit is super granular, I can't get the context of the whole change: I need to dig for the merge commit then look at that, which is faff. If there's a way that I don't know to show merge commits in blame rather than the actual source change commit, then I'd be all over it. Until then, single (whole) units of change per commit.
- Aeolun 6y agoWhy would you want one author for a merge commit? That merge can have many commits by many different authors.
- SkyPuncher 6y agoThey can, but 1. The large majority of PRs I've reviewed have a single contributor. Additional contributors are rare. When they do happen, they're often a minority contributor or simply consulting on a PR. It's net neutral when all PRs are squashed in the same pattern. 2. Even with multiple contributors, most features have one leader. It's much easier to talk to that person (and have them delegate) than it is to piece together multiple contributions.
- BizyDev 6y agoYou can also have merge commits and revert only the merge commit in this case... But yes, generally history is much cleaner when squashing commits.
- strken 6y agoOne can branch and then merge features, and get easily revertible commits and a full history at the same time. Of course, there are caveats for reverting a merge - finding the right parent adds another failure-prone step, and your team can get in all sorts of trouble if they try to work with the branch without reverting the revert.
- grumple 6y agoYou very rarely want to rollback a major feature. I’ve never seen it done except right after a deployment, in which case you can revert the merge commit instead. Bisect to find a breaking change is a very common operation. Squashing is bad if you expect to work on your project long term or if others may have the same commits as you from working on the same feature branch.
- golergka 6y agoThe fix is very easy: forbid fast-forward merges, and then you can always revert the merge commit of particular feature branch. As for cleanliness of history, everything should be as simple as possible, but not simpler. Squashing and rebasing is destroying history, which often could be valuable, as OP shows.
- m12k 6y ago> If I rollout a rewrite of an endpoint and run into a weird issue in the QA environment, I'm a simple git revert away from fixing the issue I'm firmly of the belief that any benefits that squashing brings would be better achieved with better tooling, rather than by re-writing history and throwing away potential debugging information. In this case, what you need is for git to make it easier to revert 25 commits in one go
- bluGill 6y agoIf git has a concept of branches like mercurial has it would be a lot easier as you can actually see what branch your commits attached to. (there are pros and cons of both approaches - this is a con of the git approach, I'm not knowledge about enough about esoteric details to comment on if git actually made a bad choice or just a compromise)
- m12k 6y agoIf you use the standard merge commit message in git, then you can still tell what branch things came from when being merged. As someone that has used both mercurial and git, the trouble with named branches (and having to remember to remove them when merging to master) is one of the reasons why I prefer git.
- idoubtit 6y agoKnowing the name of the branch is not enough to find the commits. With Git, a branch is just a kind of moving tag on the last commit. The problem mentioned in this thread is rolling back a feature that was merged. The only solution I know is navigating the log to find the first commit on the branch from which to revert. Don't forget there may have been several merges from and to master, as well as commits shared with other git branches that should not be reverted. In a Mercurial branch, each commit is tagged with the branch name. Unamed branches, à la Git, are called "bookmarks".
- nemetroid 6y ago
- withinboredom 6y ago> I'm a simple git revert away from fixing the issue. I don’t think this is true /just/ for squashed commits as you can simply revert the merge. Squashing is also bad, because you’ll lose the history of the bug fix that prevented the deployment. Or at least someone will have a hell of a time deciphering the PR to reintegrate the original change and your bug fix. After a few days, weeks, or months, the argument loses even more water because code will likely depend on the commit in question. War story: there was a time someone accidentally deleted a multi-gb table in production. The table would take hours to delete and replicate globally, so the entire company spent at least an hour deleting the feature from production to stop the database errors. There wasn’t any reverting of original commits. No one had time for that.
- matheusmoreira 6y agoA merge commit works just like a squashed commit except it keeps all history. This is precisely why it's sensible to avoid fast-forwarding since that operation discards the fact that a branch existed in the first place. It's better to always have merge commits. They can be reverted just as easily. I don't understand why merge commits aren't the default in git.
- Aeolun 6y agoIf you can fast forward the fact that the branch ever existed was irrelevant, since the branch is a direct child of what you based it on. Just with a different name.
- matheusmoreira 6y agoIt's not irrelevant. The branch represents a feature, a topic. It groups the commits you're merging into one logical set. This grouping of commits is exactly what will let you revert the feature later if it causes problems.
- LanternLight83 6y agoI've always been frustrated by losing my topic branches once they're merged and deleted, but can't bring myself to clutter my local branches keeping them around, to the point of tagging them just to keep track- I like the sound of your argument and will look into how fast-forward effect my commit history vs. a merge commit, thx c:
- ezst 6y agoThis is where you realize that what's killing git is that git has no concept of branches whatsoever. Such a merge isn't "merging A into B (plus shove metadata as string into the commit message)", it is "merging A and B together", which is topologically identical, but semantically very distinct. That's why Mercurial (esp. with evolve and topics) has forever my preference over git.
- Karunamon 6y agoWouldn't it be preferable, then, to avoid squashing your commits, and then tag your releases? This gets you the best of both worlds, you can still back out to a known good version in case of issues, and you can still bisect to narrow down the exact commit that caused your problem. Bisect is the killer feature of git, for me. Squashing releases takes that superpower away.
- kelchm 6y agoCompletely agreed — squashing commits isn’t desirable in every circumstance, but if all the commits pertain to a single feature then it makes complete sense to me.
- dcolkitt 6y agoI wish git had a builtin notion of two different types of commits: working commits and release commits. I really like making tiny, continuous commits as I work. It's a great flow. git-revert becomes a Ctrl-Z on steroids. I don't what to clutter up the "official" history, with all these tiny changes, many of which don't even compile. That breaks git-bissect and all kinds of other flows. So the only option is to squash commits. But there's something deeply uncomfortable and unsettling about permanently re-writing history. Plus, it's nice to have a history of those working commits as an artifact. If I'm trying to unpack the reason that I did something 9 months ago, then seeing a replay of the code changes is super-useful.
- deleted 6y ago[deleted]
- IggleSniggle 6y agoYou could get this behavior pretty easily with a double merge: “main” which contains the squashed commits and “history” which contains all commits as well as “main” merged back in post-squash. I do see your point though. Edit: the only reason to merge “main” into “history” is to enforce convergence.
- madhadron 6y agoI go one step further and work entirely in get rebase -i where I build up a stack of tiny, incremental commits. This also lets me get small changes reviewed and committed on a nearly daily basis instead of building up days or weeks of work. I've been wishing for a git GUI that lets me drag hunks around such a stack so I don't have to keep moving them by hand.
- amenghra 6y agoYou want this for git: https://www.mercurial-scm.org/wiki/GroupExtension https://www.mercurial-scm.org/wiki/GroupExtension
- mpweiher 6y ago> two different types of commits: working commits and release commits ?? git tag ??
- Nursie 6y ago> If I rollout a rewrite of an endpoint and run into a weird issue in the QA environment, I'm a simple git revert away from fixing the issue. If I had spread that endpoint across 25 commits, I'd have to actually debug the issue in QA and figure out what I broke. Surely this is why we have things like release tags and snapshots of previous versions? Unpicking individual features is rarely simple even if you have got your git repo into an immaculate state, as things are interdependent.