16 ms·
Fortunately, I don't squash my commits
- Joker_vD 6y agoDo people out there actually squash commits? Granted, I didn't change many work places in my career, but at no place where I worked people squashed commits. What's even the point of it? It's not like people routinely read the commit history, and when they do, they really would like a complete story, not 20 gargantuan commits that contain 3 years of development.
- rmrfrmrf 6y agodepending on what im doing, i'll do squashing, fixups, and commit reordering in an interactive rebase prior to merging to at least reduce the number of noise commits.
- aurelianito 6y agoYes they do. And it is frustrating. Specially frustrating when you ar the new guy and the history of the code does not tell what happened.
- eptcyka 6y agoI don't think having a history full of "Fix the shaver, maybe yaks don't need 6mm trim" with subsequent "Fix shaver again, yaks need as low as a 3mm trim" with some more intermediate commits help understanding what happened either.
- janee 6y agoyep exactly this. I found early on while having to work with perforce and using timelapse view (git blame but with a nice gui) that lots of intermediary commits like this generate enough noise to discourage even the most determined sleuth from using commit history to solve a crime. Well scoped and well sized commits, squashed, and in my personal preference rebased, provide a commit history that's segregated on actual tickets/stories/features/fixes that I find are navigatable very far back. I'd say squash usage is very much case by case. But yeah equating a feature like squash to large commits is a fallacy. It's a useful tool imo that can be misused, but that doesn't mean the tool is "bad".
- maccard 6y agoThat's much better than looking at a file in a 300 file commit and seeing "merged from XXXX" with no information as to why that one line was changed. I'd much rather spend 30 seconds parsing through the 10 yak shaving commits than have to go trawling through old commits on a file to find the most likely owner of it to ping on slack.
- OJFord 6y agoThat's not what would happen though, by default the resulting commit message contains the messages from all those squashed into it.
- saagarjha 6y agoYou’re creating a false dichotomy. When people say “squash your commits”, the don’t mean “squash your entire repo into one commit”, they mean “get rid of you ‘typo’, ‘typo fix’ commits”. Your commits should still be small and self-contained.
- maccard 6y agoActually I think you're creating a false dichotomy. How many people are commiting dozens of "typo/fix build" commits and then squashing just those? In reality people are squashing the iteration process of "add an X, add a Y, remove the X because it didn't work, and add a Z instead of an X and Y" into "Add a Z" If you're simply talking about removing "fix typo" commits, then just don't. Just ignore them. You don't need them, but someone might. They're not hurting you?
- saagarjha 6y agoWhen that happens, I usually stick in a comment about it. Generally anything worth committing gets merged anyways so you still get the "Revert X" commit in there too.
- blandflakes 6y agoI feel like any discussion of squashing or not is only half the story - people need to write good commit messages. A lot of people I've worked with commit too granularly for that single commit to be useful, so squashing gets things grouped more usefully. However, if they still make low effort commit messages, then the history is useless. But it always was useless, squashed or not!
- fs111 6y agoI always squash my commits and at my current place that is even enforced via phabricator
- gray_-_wolf 6y agoYou can always land with `--merge`. That is what I do if it actually makes sense to preserve individual commits.
- objclxt 6y agoPhabricator has a quite different (I would say better, others would say worse) development flow philosophy to GitHub/PRs. Phabricator’s preferred model, which is heavily influenced by Facebook, is to forgo feature branches entirely and just stack many small changes on top of each other, landing as and when you want (this doesn’t preclude you working on feature branches locally, of course, because Phabricator doesn’t care what your local checkout looks like). Because of this, Phabricator considers each diff to be discrete, and if you have multiple changes making up a single feature they should in turn be broken down into separate diffs. Personally, I think stacked diffs are the killer feature of Phabricator. Unfortunately I haven’t been able to find a similar flow with PRs (recently we migrated from Phabricator to GitHub for one of my projects), you end up fighting against the tool a lot.
- ericyu3 6y agoStacked diffs have been a pain for me as well, but recently I found a tool that makes it super easy to implement stacked diffs on top of GitHub! I started using it a month ago and it makes complex code changes so much easier to split up into manageable chunks. It's called ghstack (https://github.com/ezyang/ghstack https://github.com/ezyang/ghstack) If you want to learn more you can email me at ericyu3@gmail.com. Also happy to help you get it up and running - just put some time on my calendar at https://calendly.com/ericyu3/15min https://calendly.com/ericyu3/15min
- Kihashi 6y agoI find the individual commits on feature branches to be more noise than signal after they are merged. They can be useful during review, sometimes, but mostly I want a cleaner history. > they really would like a complete story, not 20 gargantuan commits that contain 3 years of development That sounds like maybe we split up work differently. 3 years of development for me or my team would likely have hundreds+ of merge/squash commits, not 20 large ones.
- pbalau 6y agoIf you look at the PRs list, you will have a clean and tidy history.
- alkonaut 6y agoYep, but many tools (most notoriously git bisect, famous from the article) doesn't understand "--first-parent" which is ridiculous. The only way to have a clean history that all tools accept is clean, is basically to outlaw merges.
- mcny 6y ago> That sounds like maybe we split up work differently. 3 years of development for me or my team would likely have hundreds+ of merge/squash commits, not 20 large ones. With the people I've worked with, I'd say most don't commit at all until they think the code is "ready" and they commit all at once. In the teams I've worked with, squashing vs not squashing isn't the question. I just want them to commit/push as soon as they've hit a stopping point or at least once a day. Maybe the people you've worked with are good with git but I am not that good with git. I'm still stuck on 6: Resolve a merge conflict on git exercises because I made one too many commits and now the exercise says I have too many commits. https://gitexercises.fracz.com/ https://gitexercises.fracz.com/ Previously on HN: https://news.ycombinator.com/item?id=24671638 https://news.ycombinator.com/item?id=24671638
- mkesper 6y agoYou need to squash your commits, then: https://git-scm.com/book/en/v2/Git-Tools-Rewriting-History https://git-scm.com/book/en/v2/Git-Tools-Rewriting-History
- FeepingCreature 6y agoI squash commits. Well, I rebase so that changes are logical rather than historical. This is because when people read the commit history, which actually happens regularly, they would actually like a complete story, not 15 commits of "fix audit", "fix review" and "fix typo" - let alone refactors where previous work in the PR is thrown out, which you can by definition never care about. Commits so small that they're nonfunctional also break bisect.
- twic 6y agoNobody is suggesting you should make commits so small they're non-functional.
- xorcist 6y agoAre you simply suggesting the removal of those non-functional commits to be called something else than "squashing"?
- twic 6y agoI'm suggesting you don't make them in the first place. If you do make them by mistake - and everyone does sometimes, certainly including me - then sure, commit --amend or squash them before pushing them up. The squashing the article is talking about is collapsing functional, distinct commits into one when merging to the trunk. It would be useful if we had distinct terms for the two kinds of squashing; the git command commit --fixup and its associated interactive rebase operation suggest the name "fixups".
- xorcist 6y agoThe difference between fixups and squashing is only in how the commit message is generated. Both are (interactive) rebase operations. Both are established terms in git, and it is probably not very useful to change that nomenclature now.
- 6y ago
- s_dev 6y ago>Do people out there actually squash commits? Depedends on the git flow being adopted by the commiters. HNers constantly espouse how clarity is more important than cleverness. It may be the case that a single commit offers more concision than multiple commits which could be perceived as just more noise.
- ZephyrBlu 6y agoAt my workplace we generate ~150 commits every week or two. There are a lot of trash commits and the history is completely unreadable. I'm sure that's barely any compared to large companies, so I really question the value of commit history unless message guidelines are well enforced.
- maccard 6y agoWe go through ~600 commits per day. In my experience, there's nothing worse than looking for when a value was changed from 10 to 100, and finding it in a "Merged from X" commit with the history of why the value was changed from 10 to 100. I don't often (ever?) browse the history, but I _do_ regularly search the history using tooling, (git log | grep <something>), and more is more in that case, even if the history isn't perfect.
- saagarjha 6y agoThat just sounds like a bad squash.
- maccard 6y agoIf the original commits are: - Change X to 10 by Foo - Change X to 100 by Bar - Change X to 50 by Baz And that gets squashed into `Change X to 50 by Baz`, you lose the context as to why Foo and Bar changed it, and the values. If I need to go investigate an issue with X, I'd rather thave the history of all the changes.
- saagarjha 6y agoCommits by multiple people should rarely, if ever, be squashed.
- ChrisMarshallNY 6y agoI squash, but it's not something I do regularly. In my case, it's generally because I'm preparing training material. When I do that, I have two branches: dev and master. master is the one that is exposed to students, and dev is the one I use for prepping, testing and staging. Sometimes, I may even have the dev branch in a private repo, or one that is associated with a different GH ID, like so: https://24ways.org/2013/keeping-parts-of-your-codebase-private-on-github/ https://24ways.org/2013/keeping-parts-of-your-codebase-priva... The main "gotcha" for me, is to make sure that I merge the master back into dev, after doing the squash. Makes life a lot easier, for the next squash.
- flohofwoe 6y agoI squash commits on large PRs on a case-by-case basis, especially when they have a lot of very small and silly commits (fix this, fix that, wip, ...) but implement a very specific feature which required a lot of experiments that didn't end up in the final PR. And yes, I use the commit history very frequently, so it's important for me that this doesn't contain too much noise. If I need to bisect a bug that was introduced by the PR I still can do this in the original branch.
- dmatech 6y agoIdeally, that would give you the best of both worlds. But some places delete the old branches after they're merged. It would be nice if there were an easy way to hide or rename old branches so that the in-use ones stand out.
- stepsrabbit 6y agoIf you always work in feature branches and commit often (which you should IMO) it makes perfect sense to squash on merge. It gives you a nice history that only contains relevant commits instead of having a bunch of “added tests”/“fixed XYZ”/“remove debug log”/etc commits. With some discipline this makes the commit history actually worth reading and makes git blame a useful tool.
- SXX 6y agoI guess It's my development culture is flawed, but as one-man-team I always time constrained. So I sometimes don't have sufficient time to write too detailed commit messages. So I end up with "one commit per feature" rather than "one commit per logical change". So I actively using squash during interactive rebase when I prepare to merge completed branch since my commit history sometimes looks like this: Backend: new set of APIs for XYZ Backend: implement feature X (+ some long comment) Frontend: implement feature X fix for feature backend fix for backend API XYZ front fix backend fix Yeah again I know I could do better, but yeah I use squashing for this reason. A lot.
- Joker_vD 6y agoWell, uh, yes, that's what the commit history generally looks like (although we do generally put "JIRA-9999: " in front of all commit messages, for context). So what?
- SXX 6y agoWhat I wanted to say that since I don't have that many people other than me looking into my code there no reason to preserve real development history of every feature and IMO 3 "feature commits" are preferable than my development mess of 50 "fix that fix this" commits.
- np_tedious 6y agoThe argument for squashing is that minor updates like spelling, renaming, or test fixes during initial development can really clutter the history if they each have a commit. Many would rather see the actual change "Update X to use Y instead of Z" in the history, and minor details like "Fix mock in XTestCase" or "Perform renames from code review" within that single commit. I'm sort of agnostic on this issue, but I do feel the article's author kind of overstated it here. Say 5 to 10 commits of this size were squashed together. Git bisect would've taken him 90+% as far and he'd have to read code or manually trial and error changes just slightly more. The binary searchable problem space would be slightly smaller, and the linear manual effort space slightly bigger. Less good, but really not that big of a deal.
- twic 6y agoFor me, it's the other way round: minor updates like spelling, renaming, or test fixes during initial development can really clutter the major updates. If i am making a commit that makes a complex but important change to some significant application logic, i want that commit to contain that change and only that change, so that when i have to re-read it a year later, it's completely obvious what i did and why. Bundling a load of refactoring and cleanup in there is a significant speedbump for my understanding. Years ago, a sage pointed out the argument for squashing is really an argument for better tools. Imagine if you could flag commits as being of two types - major/minor, significant/insignificant, feature/refactoring, foreground/background, melody/rhythm, etc. Then imagine if the tools would by default hide, roll up, or otherwise de-emphasise the commits of the latter kind. This whole apparent dichotomy would go away in a flash. This idea is floating around in the Wiki world. I believe it was Ward's Wiki that introduced a 'minor edit' checkbox in the editor; if a change was marked as a minor edit, it wouldn't be show on the recent changes feed. You can imagine other ways to get somewhere similar. For example, you could have a special kind of commit that just groups a previous run of commits, and the tools could show that and hide the members of the group by default. There are probably many other ways to do this.
- tasogare 6y agoInteresting idea, however this could also be easily abused for nefarious use.
- jdmoreira 6y agoI religiously squash my commits for each merged pull request. I seems like madness any other way to me... why have a bigger granularity than a single merge commit?
- adrianN 6y agoGranularity helps with bisecting issues.
- bJGVygG7MQVF8c 6y agoYes. But also: Granularity undermines the effectiveness of bisecting when your commits aren't ACID. These conversations are so often tedious and repetitive largely because engineers seem incapable or unwilling to apply systems thinking to development practices.
- jdmoreira 6y agoat least in github after you find the squash commit equivalent to the PR you can restore the branch and bisect the rest in there. This is much faster!
- DougBTX 6y agoInteresting, it seems like madness to me to have it the other way around. Why bother with a merge commit if all the changes are squashed into one anyway? Why not just fast forward merge as if it was a commit directly to the main development branch?
- swalsh 6y agoI do, all the time. I highly recommend others do it too. I like to keep to 1 commit per ticket (even a super large ticket that might take me weeks of development). But having a commit history like the fella in this story is nice too. So what you should do is work in a feature branch. On the day of deployment I'll squash, and cherry pick (in my case to a train). If the deployment goes bad, and my code is at fault, then all that is required to fix it is a single revert. If you want to continously deploy to production (with real users using it) you need a process to very quickly revert bad commits from it. At my last company with hundreds of developers, we would push to prod several times a day. Each deployment would have several devs changes included. A clean linear history is essential to making that process work.
- Joker_vD 6y agoWait, I must be missing something. I too develop in feature branches, and then they're merged into develop or main (with a merge commit, obviously). Reverting that is too only a single revert. The only problem is that sooner or late those feature branch do get deleted, and if you swashed, you lose all the non-squashed history. If you don't, it's still in the develop/main branch. What do I miss?
- mfontani 6y agoYou get a similar benefit (a single revert to revert it all) via "git merge --no-ff BRANCH". If the BRANCH is a fast-forward (i.e. it's been rebased to master/main & tested before merge) then you get both benefits of a clean history and an easy revert, for little downside. Keeping the history of each incremental change, even in a branch, is IMO too useful to give up. Note I'm not advocating keeping those "fixed typo in previous commit" fix-up commits: those should be properly fixed up _before_ merge, by judicious use of git rebase.
- lordnacho 6y agoThe only time I ever do it is if it's trivial. Say I've removed some vestigial code, then I find some more after I've committed.
- aequitas 6y agoDepends on the Git craftmanship of the team is my experience. If they commit every small fix with a useless commit message, history becomes messy real quick. Squashing can be seen as a solution here. Until you learn proper rewriting of history as in (interactive) rebasing. After which you can put you code changes in every commit you want and order them around as you see fit. But that also depends on the Git workflow that is used and how much branches are shared amongst team members. I had one job in the past that used Gerrit[0] as Git tool. One of it's features is that it creates a "pullrequest" for every commit in the branch you push. Which needs to be reviewed individually. This is really anoying if you're used to organising your work in lot of commits to record each step of your developerment. But from a project's Git history perspective it makes a lot of sense. As every commit is 1 change, one feature, one contained unit, that is added to the main branch. So instead of the main branch now containing countless commits with each developers complete history on a specific feature (where code is added in one commit to be removed in the next) it contains the features as distict commits, making them easy to bisect and revert if needed. This looks a lot like squashing but because you do it before you push your code you learn to put much more thought into that single commit and the commit message. [0] https://www.gerritcodereview.com/ https://www.gerritcodereview.com/
- saurik 6y agoI think what you really want is neither of squashing nor retaining: what you want is to take the messy history you had and then reconstruct a history that makes sense and commit that; thankfully, git makes this easy. I hate it when people leave messy thoughts in the commit history as I do use the history, and it is even worse if they just leave mega-commits. I want to see easy to review step by step organized thoughts designed to help the reader appreciate the series of steps needed to bring you from where you were to where you went.
- twic 6y agoI would say the best thing is to cultivate the habit of thinking and working cleanly, so you create a history of small, logical, incremental commits. But the second best thing is to think and work messily, and use rebasing to fake a history of small, logical, incremental commits!
- simias 6y agoI often do intermediary commits that don't compile or break something significant, only to make sure that I don't lose the code. When that happens I always squash my commit afterwards once the code is in a usable state. The alternative is making bisecting harder which is not something that I want. I want every commit to compile and be testable individually. That definitely doesn't mean that I think it's a good idea to have "gargantuan" commits. An ideal commit should contain a single, atomic change to the codebase. Not more, but also not less.
- rutthenut 6y agoI'm a bit confused on your comments here. You state "I often do intermediary commits that don't compile ... to make sure that I don't lose the code" And follow that with "I want every commit to compile and be testable individually."
- simias 6y agoI mean that I do many small "crap" commits while developing, once I'm ready to merge into the main branch I clean the history to get proper, atomic commits.
- StanLindsey 6y agoI'm the same as OP. I commit constantly when hitting a natural stop point or when taking a break and push up to my branch/pr. I hate having code locally that hasn't been pushed to the origin. But doing this means I have many commits that don't mean anything or are unfinished and in an uncompilable state. So I git rebase and massage the history to make more sense.
- clinta 6y agoEveryone who works on Linux must rebase and squash commits so that every commit builds and passes tests, and does a single thing. More about how Linux uses git: https://www.mail-archive.com/dri-devel@lists.sourceforge.net/msg39091.html https://www.mail-archive.com/dri-devel@lists.sourceforge.net... https://www.linux.com/news/why-linuxs-biggest-ever-kernel-release-is-really-no-big-deal/ https://www.linux.com/news/why-linuxs-biggest-ever-kernel-re...
- alkonaut 6y agoYou don't have to squash into ONE commit just because you squash. If I make 10 commits I can squash it into two logical commits (e.g. refactor, add feature) then I merge the branch with those two. If the branch is small and has 3 commits that are one logical change, then I might as well squash to main instead of merging. Commit history is readable regardless (I'd never use anything but --first-parent ever).
- mcv 6y agoI do read the git commit history, and specifically to find problems like these. Squashing commits indeed sounds like a terrible idea to me. There are also people who insist you always need to rebase your commits before pushing, in order to get a nice, linear history, and again I disagree. It's fine when you're rebasing a very short (single commit) history, but for a long history, it not only gets very tedious, but when it introduces new bugs halfway through that history, you may not notice. You will notice later, and then tracking that bug you end up in the middle of your own (rebased) history and you may wonder why you ever did something so stupid, when the actual cause of the bug was the merge of two histories at the end of your work. I do rebase sometimes, but only when the history is short and I can easily see what I'm rebasing. A history longer than 2 commits should not be changed. I never squash.
- reallydontask 6y agoFrom experience of working with new CI/CD systems where one doesn't quite understand what's failing due to ... reasons. It's really quite easy to get into the 50s or close to 100 commits in a branch, which can lead to a horrendously messy history. There are also people that like to commit before the try something locally and their branches are effectively a mess. All of the above is magnified in a monorepo environment, where you might have thousands of commits a day against master if it weren't for enforcement policies that force a squash.
- timvdalen 6y agoIt depends on the situation. If I'm merging a feature branch that has a lot of commits that effectively make up one feature (because a dev had to go back and forth or because there was a lot of feedback), I might squash those commits into one before I merge it into master. As mentioned elsewhere in the comments here, that also makes it a lot easier to revert said feature if something goes wrong.
- 3np 6y agoRoutinely. Common example that I can't see why anyone would take issue with: * Commit 1: Fix a bug * Commit 2: Fix linting issues with the fix discovered through CI * Commit 3: Remove dead/commented code introduced thrugh commit 1 * Commit 4: Update documents as required by change I'd always squash those into a single commit before merging into upstream.
- asveikau 6y agoI can see a good case for separating #2 from everything else. Or depending on the change, maybe squashing #2 and #3 together separate from 1 and 4. It's a good idea to isolate the bugfix so that reading the diff is clear. Stylistic changes such as #2 and #3 can be separated and called out as "should produce no behavior change" sort of updates.
- ratww 6y agoSquashing is kind a blunt tool that is useful sometimes, but can be perfectly replaced either by crafting your commits with more care or doing interactive rebasing. I once had to enforce to an developer that wanted to make 20 commits per PR that were titled like "wip" "wip" "wip" "fix bug" "fix mistake" "format code" in a PR that only changed like 10 lines of code in the end. In this case, the main problem for me was that git blame became useless because this developer was touching way more lines than necessary and then undoing it using the code formatter.
- rjkennedy98 6y agoYes, our company squashes commits (we have tens or probably hundreds of thousands on the main branch). Yes, I routinely look at commit history. Commit history is also used for doing analytics for performance reviews such as how many commits did you make, what percentage had tests ect. For instance my last performance review made note that I had test coverage on 90% of my commits. Also, we auto-stamp the PR on the commits so we can get the context. I don't understand how anyone could make sense of a large codebase when all the commits are added with the standard inane commit messages (e.g. "fix it", "typo", "name changes", "tests") that people do when building a feature. I routinely have to look at a piece of code, do a git blame, get the PR in the commit and figure out what was being done.
- CGamesPlay 6y agoI'm probably 75/25 on this. 75% of the time I'm gonna squash, because my entire commit messages is along the lines of "got to this point" or "fix frotzolate when x=7". 25% of the time my first round of squashing yields good commit messages that are isolated and complete, so it's better to leave those as separate commits when merging into the mainline. I also aggressively reorder my commits. When it gets to yak shaving you basically develop in a stack, so you end up with a half-functional commit to system A at the top, then a complete commit to system B, then the finishing commit to system A, so it's better to just rearrange it so that you have one system B commit followed by one system A commit. But I do routinely read the commit history (using git blame) and no I don't want to see the "complete story" of my past self having "got to this point"--I want the documented MR commentary.
- __s 6y agoWe do. PRs are generally smaller than 3 years of development. Sometimes refactoring will be split out so what could be a single PR with multiple commits instead becomes multiple PRs. This allows review to be focused on each PR & enforces that each commit maintain CI passing
- UK-Al05 6y ago# Implement feature # Whoops fixed issue in feature i just implemented # Add in whitespace # Remove whitespace # Forgot place to add in whitespace # Fix variable name for feature Vs # Implemented feature "x" Which ones easier to rollback and read.
- chousuke 6y agoPersonally I think of working in a branch as the process of creating a patch set. A patch in a series should only depend on its predecessors. You can submit more than one patch in a pull request. I would never send out a patch set for review that includes all the "oops" "typo" "iteration 50" etc. commits I make as I work on code. Those are pure noise. However, git is a tool for development, and during development I should be able to use git in whichever way is convenient for me: commit, rewrite and do whatever the hell I want with my local history. Not having that freedom is the primary reason why working with most non-distributed version control systems is such a pain. When it comes to actually merging patch series to master, I like not doing fast-forward merges, since you maintain a natural grouping of the applied changes.
- DrBazza 6y ago> Do people out there actually squash commits? Yes, squashing parts of a discrete piece of work together, where it makes sense, then makes it trivial to git bisect any future problems. Code committed to the mainline should always compile, and should be made of discrete changes. However, do what you like on your unpublished local branch. > What's even the point of it? It's not like people routinely read the commit history, and when they do, they really would like a complete story, not 20 gargantuan commits that contain 3 years of development. My company routinely reads the commit history. The commit is usually the 'why', and the code is the 'how'.
- Vinnl 6y agoYes I do (well, I rebase, not squash blindly), but I do read the commit history. Or well, not so much the complete history, but I do go looking for which commit introduced a particular line, what message went with that change, and what else changed with it. Of course, that only works after you leave proper commits. The reason people don't do that is because you only look at your commit history if you've kept it useful, but if you've never looked at your commit history because it's not, you don't know the benefits of keeping it useful.
- vonmoltke 6y agoIt really depends on where you work and how your company's repo is organized. For instance, where I work 20 squash commits would represent, at most, 10 minutes worth of commits to the monorepo. Not squashing on merge would quickly turn thousands of daily commits into tens of thousands. It's already impractical to find suspect commits via git, and we typically use our code review tool to find the change that broke something (and that change includes the developer's full branch history). Adding more commits would needlessly slow down the tooling that displays 'git blame' results (and likely other commands). I suspect there is a pattern in the comments here. People who work on small teams with granular repos that, individually, don't see a lot of daily activity think squashing is bad and erases valuable history. People who work with large repos that see high commit velocity (like a monorepo) think squashing (at least merge squashing) is beneficial and don't see the loss of information as problematic because it's hard to access it in the first place. Maybe I'm just projecting my own opinions on this; I'd like to hear perspectives that conflict with my assumptions.
- spdionis 6y agoYou are definitely correct. People who argue against squashing have not worked on 10+ years old actively developed repositories, or in big enough teams. A commit is a change. A change has a ticket. A ticket is a small piece of work that does not result in 3k changed lines. This ensures that the change rationale is fully documented and easily identifiable. Nobody needs those "fixed typo" commits. Nor the "implemented function A" commits. What IS the change that you're doing? What functionality? That's the most comfortable commit granularity to debug imho.
- mixmastamyk 6y agoAgreed, I regularly add comments later as well. Understanding often comes after the code works. I could obsessively fiddle with every commit like an artesenal snowflake, or I could click the squash checkbox on the request.
- mixmastamyk 6y agoThat's not how it works, really. A simplified workflow: You take bug or enhancement of a couple of points, work on it on its own branch, then merge req+squash the mostly noise commits into a develop/master at once. A chunk is approximately 3 days, not 3 years. For forward-looking projects, it saves time and works well. "Enterprisey" projects maintaining legacy branches are less well served.
- kelnos 6y agoIt's a mix where I work. Most people aren't at all tidy with their commit history, and I'll often see a PR with a fairly small final diff (maybe a couple hundred lines changed total) with several useless one-word commit messages like "fix". And then after the PR has gone through review, most people tack on extra commits with equally-useless commit messages like "addressing feedback". It's infrequent that I see PRs with commit histories that actually chronicle the history of the change itself; it's more a chronicle of the developer's changing thought processes as they try different things, go down blind alleys, change approaches several times. Most of the individual commits have test suite failures and some of them don't even compile, so it's impossible to bisect across them if an issue is later found. In those cases, I wish people would just squash (and some do, where I work). Yes, you lose information and separation, but I'd rather have one large working commit than 15 small broken commits followed by one working commit. Ideally people would curate things before submitting their PR, but I've found that most people just don't care, or don't understand git well enough to even attempt to do it. Sometimes I toy around with the idea of trying to teach people (what I consider) better practices, and insist they are followed, but we all have a limited amount of social capital in our workplaces, and I'm not convinced this is something worth spending it on.
- aurelianito 6y agoI never understood the need to squash commits (or rebase). If you do merge requests and use merge commits (like GitHub or gitlab do). A "nice" history is a small script away. It should even be a part of the GitHub/gitlab gui. Do not loose information about the development history!
- fs111 6y agoMerge commits are noise, a clean history is has no merge commits
- NateEag 6y agogit log --no-merges
- leejo 6y agoSee also git log --no-merges --first-parent I don't get the "no merge commits" argument - git log has dozens of options allowing you to bend it to your use case. To address the grand parent that "a clean history is has no merge commits" I would argue that a clean history is also a lie if you're working in a branching workflow (hint: you should be working in that most of the time). If you want to avoid (note: not eliminate) merge commits then make sure to rebase against master before merging, and then ensure merges are fast-forward merges.
- aurelianito 6y agoIf you squash all the merge requests, you get ONLY merge commits.
- Droobfest 6y agoThey can both be pretty useful. So keep all commits and decide upon viewing the log which ones you really want to see.
- OJFord 6y agoA 'merge commit' is the commit that ties together two strands. * merge commit |\ | * branch work | | | * branch work |/ * If you squash to merge, typically you're also going to rebase it (equivalently, if it's more familiar, cherry-pick the squash onto the branch your 'merging' it into). * squash cherry-picked / rebased | * both branch works squashed | * | branch work | | | | * | branch work |/_/ * (In this case the target branch could have been fast-forwarded, but this also works if there's some other work on the mean time:) * squash cherry-picked / rebased | * something unrelated | * both branch works squashed | * | branch work | | | | * | branch work |/_/ *
- ZephyrBlu 6y agoIf you use Pull Requests with squashed commits, you can do exactly the same process as described here by isolating the issue to a PR then restoring the branch and bisecting from there. It seems like a small price to pay for a clean history, given the rarity of occurrences like this.
- NateEag 6y agoIt seems like a needless price to pay, since the whole point of keeping history is to help analyze problems and past states. The author has only used bisect for a head-scratcher once, but I've used it often, sometimes to pin down bugs in code I'd never even looked at before. That's not feasible if the commits are huge.
- ZephyrBlu 6y agoYou do keep the history in the branch, it's just out of the way unless you really need it. Also, the method I described has nothing to do with commits. PRs and small commits aren't mutually exclusive.
- NateEag 6y agoYou have to go dig the branch out of the central repo, though, which is annoying and takes time. That was the price I was thinking of. You're certainly right that PRs and large commits are orthogonal.
- prussian 6y agoI don't see how squashing or not would have made this case hard to find. Once you find the bad case "squashed" one, one could always reset HEAD^ and just checkout files back and forth til you find the failing change.
- karolist 6y agoSquash renders git log and bisect nearly unusable, that should be common knowledge, but I guess not. Normally there's never a reason to squash, small fixes can be done with git commit --amend for last commit or git rebase -i using the fixup keyword, but again only for really trivial things like typos and such otherwise same problems with log and bisect.
- vemv 6y agoIn a given team, mandating the squashing of commits essentially means admitting that pull requests' branch histories tend to be far from semantically valuable. Which can be perfectly fine, as it's hard to ensure that all developers use Git in an optimal way (I'm thinking of intentful use of interactive rebasing), uniformly. The only problem I find is when squash proponents claim their choice is superior. It's not; it's only superior if you aren't willing to maintain great history as you work on a given branch. Of course there are also added benefits to a fine-grained history, as the article mentions. My experience having maintained a production app with a fine-grained semantic history for 5 years is overwhelmingly positive. Understanding root causes, confidently reverting things etc becomes a much quicker job - particularly important when production is red.
- josephg 6y agoRight. The problem is that are two use cases for commits: - a historical, fine grained log of changes - and a log of merged features. There's value in keeping both of these data sets. But the commit log as it stands can't easily serve both masters. Once you can see the problem for what it is, the solution is simple. Instead of conflating these two use cases into the concept of a 'commit', we need separate tooling for each of these use cases. The commit log should probably house the historical record. And then we need a way to mark a set of commits as belonging to a particular feature's development. That could be achieved either by adding special support in git or via convention using commit messages. Either way, I want to be able to see the commits in my repository grouped by the feature that they belong to. And I want that data set browsable on github, referenced against the corresponding github issues when thats appropriate.
- avar 6y agoIf you consistently use feature branches and have meaningful summary commit messages in the merge commits as they land on the mainline branch what you're describing is simply: git log --max-parents=1 git log --min-parents=2 E.g. try this on the git.git repository (not perfect there, since Junio doesn't use this pattern all the time), but it's good enough for a quick demo. The most useful convention in a DAG like git is to have the topology of your commits reflect your workflow.
- thombles 6y agoOne obvious requirement for bisect to work is that your code builds on every commit - and so does everybody else participating in the git history. Straw poll - do you enforce this? If so, for literally every commit, or do you use partial squashing to maintain this property? While that's certainly a _desirable_ property, I've never really been concerned if, say, the penultimate commit on a PR failed CI. It feels like it would be a hassle.
- alkonaut 6y agoWhat (I assume) is normal is that you have PR validation and CI builds that only maintain the stable shared branches. If the validation isn't extremely fast (seconds or minutes) then maintaining it for every commit is almost impossible. It would just get other side effects like people avoiding commits because they don't want to run an hous-long test suite more than once.
- yipbub 6y agoI've always been careful about this even though I've not had the habit reward me yet (3 years of exp). If I'm rebasing, I go through every commit and make sure it builds. TIL about `git bisect` and it's really vindicating.
- aeonflux 6y agoThere is no such requirement, you can just skip commits which cannot be tested. You just make a simple test-case just for current bug outside the repo.
- speedgoose 6y agoWe use JavaScript at work so even if a part of the application doesn't compile, we usually can run most of the tests. We don't bisect very often and so far, every commit has been good enough for the specific tests we wanted to execute in case of a bisect.
- chinigo 6y agoI don't "enforce" this but I do aim for it, and yes, I do partial squashing throughout my workday. As I work on a feature branch, I'll check in WIP commits as checkpoints, especially at EOD. I don't expect these to pass the full CI suite. But as the code starts to shape up, I'll unstage all those WIPs and start to group the changes into logical commits. As work progresses, I'll use `git add --patch` to split new lines of code into those existing logical commits. Sometimes I'll split one up, sometimes I'll group two together; it's still flexible and amorphous at this point. By the time I'm ready to merge upstream, these commits tend to be neat, focused, and functional, and I do check to make sure they pass the relevant tests (though I don't enforce a full CI build here). Then a rebase from master, push to CI, and then a no-ff merge commit into master to retain both the low-level commits and the ability to easily revert the whole lot. It might seem like a ton of busywork, but I find that staging atomic commits like this doubles as an excellent line-by-line review of the code I've written. It also forces that review step to happen throughout the process rather than all the way at the end when I've forgotten all that deep context.
- davewritescode 6y agoThe reason to squash commits is more than just keeping your commit history read-able, it's about making easy to revert a feature and being able to keep history in a way that makes it simple to revert a change if you run into issues. If I rollout a rewrite of an endpoint and run into a weird issue in the QA environment, I'm a simple git revert away from fixing the issue. If I had spread that endpoint across 25 commits, I'd have to actually debug the issue in QA and figure out what I broke. Reverting quickly lets me debug the issue on my time instead of keeping our test suite broken. It's not always feasible, and I try not to be a stickler when people on my team don't do it but the fact is, if you're working in a world when you're delivery code quickly into real environments, having a back-out strategy is paramount.
- twic 6y ago> If I had spread that endpoint across 25 commits, I'd have to actually debug the issue in QA and figure out what I broke No, you could just revert the entire lot in one go.
- Mazzen 6y agothat's not a valid argument. How do you identify "the lot"?
- magicalhippo 6y agoBy the ticket/issue number?
- deleted 6y ago[deleted]
- gmfawcett 6y agoThere's an obvious upper bound. If you can't identify the last in-prod version that worked, and the first in-prod version that failed, then you have bigger problems than deciding whether to squash your commits.
- twic 6y ago
- craig_asp 6y agoIt really depends. I've seen developers commit tiny changes all the time, which do not make sense - I wouldn't want to trace a small logic change over 10 commits. Also, sometimes you commit, then realise that there's a stupid mistake you've made and then you fix it with another commit. In those cases it does make sense to squash the commits so that you get a normal, logical, commit in the history. On the other hand, squashing perfectly legit commits together into a horrendous single commit is a pretty bad practise for obvious reasons.
- FrankSansC 6y agoIt is not related to squashing or not. One of the main rule when using git (or any other VCS I guess) is that each commit should be atomic. IMO squashing commits is something that you should do locally, not remotely. For example I'm currently debugging a fairly large C application on an embedded system. Each bug has it's own branch. When debugging / testing / trying to fix it I tend to do a lot of commits. When the bug is fixed I do an interactive rebase to pick / drop / squash commits in order to have only a clean one at the end (and then push it).
- kelchm 6y agoWhy do it locally when most modern repository software (GitHub, GitLab, BitBucket, etc) can do it for you when a PR is merged? My point of view is that if work is being done on an individual feature or bug fix, having a view of the individual commits might help give any reviewers important context on how a final solution was arrived at. It can also be helpful to be able to see the specific changes that have been made since a prior review. While in most cases I favor squashing commits when it comes time to merge into the parent branch, it seems like doing it manually just creates extra work and potentially throws away information that may have been useful to reviewers.
- bwhmather 6y agoCommit history serves two largely separate purposes: 1) To provide a code-level record of what has changed in individual files and why. 2) To provide a high-level record of what features were introduced and what bugs were fixed over a period of time. Often people will forget about one of them when arguing for a particular approach. Squashing can make 2 easier but annihilates 1. Rebasing gives you 1 but makes 2 difficult, or requires that you track high level changes in an external system. In theory, approaches using merge commits can give you both but they are often difficult to apply in practice.
- bkirwi 6y agoWorth mentioning that, while the `git bisect` algorithm is pretty easy to run by hand if you have a linear history, `bisect` also handles very complex branching and merging histories as well. Very clever and worth using!
- barrkel 6y agoSquashing gets you in the habit of force pushing, and force pushing greatly increases the risk of clobbering someone's work. I've seen a lot more hours wasted on lost commits than I find credible to be wasted in slightly more verbose history. Squashing also turns cherry picked commits into merge conflicts. I'm not a fan. I can always go back to the feature PR and read the ticket details and diff commentary when I want the history with a full diff. And of course very old history isn't very relevant because less of it remains in the present, so rot of ancillary systems isn't a huge concern.
- Gaelan 6y agoI'm no squash fan either, but git very rarely actually clobbers data. (As long as it's been committed at some point, and wasn't just sitting in the working tree.) `git reflog` can get back pretty much anything, as long as it happened relatively recently.
- falcolas 6y agoUntil a garbage collection process trims those. You may control such on your local machine, but you don’t on the remote repo.
- Gaelan 6y agoThat’s why I said “as long as it happened relatively recently.” Of course, the commits in question should exist locally on someone’s machine, so we shouldn’t need to worry about what the remote does.
- reallydontask 6y agoMost people at my previous place did squash merge on pr completion and we cherry picked quite successfully, so I'd be interested to know how you get to: > Squashing also turns cherry picked commits into merge conflicts.
- nailer 6y agoHijacking the top thread to make an important point: - "Squash your commits" folks - yes, it's good to be able to revert a commit and remove entire features and have a readable commit history. - "Make small granular commits" folks - yes it's good to be able to bisect and see where exactly some behaviour changed. Rather than repeat these points (which are both true), there's a better question to be asked: Should there be a way to have a readable history and keep individual commits? Eg, 'subcommits' etc? - git may not have these features now, but revision control systems come and go (rcs, cvs, svn, bk, hg, git) - Or maybe it does. Or we can implement something similar on top of git as it stands. Discuss.
- twic 6y agoI'm intrigued to see that there are (by my count) three of us making this point today. Perhaps this is an idea whose time is finally coming.
- leejo 6y agoYou can revert a merge commit by providing which side of the merge is mainline: git revert -m 1 $sha_of_the_merge_commit That provides the "revert a commit and remove entire features" use case, assuming the original feature was all encapsulated within the merge (branch) in question. You can get the "readable commit history" with: git log --no-merges You can have the best of both worlds. git has had these features for as long as I can remember, but devs have had the "NO MERGES ONLY CLEAN HISTORY" platitude repeated so often they must think git lacks an alternative.
- vollmond 6y agoThis response [0] by avar shows you can get both by using the min/max parents flags, to either see only merge commits or only code commits: [0] https://news.ycombinator.com/item?id=24687027 https://news.ycombinator.com/item?id=24687027
- ivanhoe 6y agoIMO the better title for the post would be "Unfortunately, I don't setup the clean environment for each of my tests"
- _-___________-_ 6y agoThe most striking thing in this story for me has nothing to do with commit squashing. Faced with something that suddenly stopped working, the developer went on a wild goose chase, first decoding their JWT -- presumably the same one that "used to work" -- then looking for possible framework bugs or misconfiguration, before looking at their own code, the single most likely place that the bug would be. They even linked to select is rarely broken but failed to internalise the main lesson behind it. Whether they squashed their commits or not would make a far smaller difference to the time taken to spot the bug than simply not assuming that it's everywhere else than their own code first.
- jerf 6y ago"went on a wild goose chase" Because 19 times out of 20, that "wild goose chase" finds the bug faster. As there is no perfect debugging methodology, sometimes whatever method you use will go wrong and you'll end up on the bottom of your list of things to check, or worse, right off the bottom of the list. (Those are some bad days.) That is not, itself, proof that your list is broken. After all, the author found this an unusual enough experience to write a blog post about.
- _-___________-_ 6y agoStarting with your own code is going to find the bug faster far more of the time than starting with "maybe the JWT got randomly broken or my JWT library is broken" or "maybe ASP.NET is broken".
- DarkWiiPlayer 6y agoOver the years I've figured out that you can not only bisect in your development history, but also in the program runtime. Start with "print('1/2')" or whatever around halfway through the runtime of your code and see if you get there, then move to 1/4 or 3/4, etc. This can be done with break points, of course, but in my setup adding a print is usually just faster.
- saagarjha 6y ago
- scoutt 6y agoSide comment: it's curious how the author went from "not my code", to "the problem was in the code". I say this because as a low level/system developer, I often have to solve problems for the people integrating my platforms from higher level languages. When something doesn't work they come to me, and it often goes like this (with quotes from the article): First stage: someone else's fault. "the problem is environmental: a network topology issue, a bad or missing connection string" Second stage: magical things, but not my code. "Not my code ... the problem occurs somewhere in the framework" Third stage: acceptance. "Global variables are evil. Who knew? ... It took me hours to find the bug, and ten seconds to fix it."
- dpryden 6y agoI think much of the difference here depends on your workflow. Where do you run tests? How do you manage build infrastructure? What is your process for code review? What is your process for validation and testing of shared branches? How do you manage releases? In my opinion, squash merges make the most sense when the following are true: 1. Developers cannot push directly to shared branches; all commits into shared branches are done via pull requests with mandatory code review. 2. Shared branches must build and pass tests at every commit. 3. The team builds features incrementally and uses the "Branch by Abstraction" approach (feature flags/experiments) to ensure functionality can be merged before it is ready to be enabled in production. 4. Changes that merge into the shared ("trunk") branch are released frequently. Now, you may argue against some of those practices, and if you end up in a different place on one of those, squashes may not make sense to you. But then the actual difference of opinion is elsewhere. It's not really about squashing, it's about how much code is a reasonable granularity to code review at one time. If you can keep the code review cycle tight, then a pull request branch basically becomes almost like a virtual pair-programming session. There may be some back and forth as the author and reviewer settle on a final form of the commit. But I don't see the commits that happen along the way as valuable in the least. (In my experience, the vast majority of intermediate commits are things like "fix typo", "make linter happy", "rename this function because code reviewer pointed out it was named inconsistently", etc.) Instead, my mental model is that a pull request is like a mutable commit: you keep mutating it until it's the commit you want, and then you merge it. Since my mental model is that a pull request is a kind of commit, it makes sense that when it merges in, it does so as a single commit. Again, I'm not saying that this is the only valid way to work. But neither is it an invalid way to work. As a meta-comment: When you see someone else making a technical decision that doesn't make any sense to you, rather than instantly assuming that they are a clueless incompetent, it's usually more instructive to assume that the decision does make sense to them, and to try to figure out what about their circumstance is different from yours to make that the case.
- vitus 6y ago> Not all bugs can be reproduced by an automated test While that may be true, I'd say in this case, order-dependent tests are evil. This case study pretty much fits the definition -- the test result changes depending on what other tests are run before it (due to shared state across tests). It certainly is possible to write automated tests for this -- reset that state to its default before each test! One thing that Google recently (in the past year or so) added to its CI setup is a job that goes through periodically and randomizes the order in which tests are run, thereby increasing the chance order-dependent tests are caught. (Early on, I remember being annoyed by all the triggers due to floating point error accumulation in a monitoring library -- do I really care about an error of 1e-10 when my noise is >10?) That's a long-winded way of saying that it's also possible to catch order-dependent tests by shuffling the order in which they run, although built-in support may vary depending on your testing framework. (from a brief search, it looks like rspec and junit both support randomization; xunit forces it). That said, I absolutely agree with the other conclusions, especially that binary search for debugging is immensely useful (even when dealing with Google-scale tens of thousands of unrelated commits due to monorepo), and definitely empathize with debugging often taking much longer than the bugfix.
- plextoria 6y ago`squash` is a tool and it's neither good nor bad; it needs to be applied where it makes sense. The title is indeed click-baity. It would have been more interesting to read that the bug/mistake was caused as a result of squashing. It's not the case and I take issue with the way the author describes his commits. The problematic commit is described "Extract CreateTokenValidationParameters method", without an explanation of why the refactoring is necessary or what problem is it solving. It looks like it is improving code readability, but it doesn't go as far as fixing the more glaring issue with global variables. Other commit messages seem to follow the same pattern. In other words, the commit: - has minor code readability improvements - contains no useful message for the future programmer (why is it needed?) - provides no new functional feature/improvement/business value - is later found to contain a bug I find squash/rebase/cherry-pick useful when reviewing my work and deciding what should go in the current pull-request. For example, a refactoring might be postponed for later if it is deemed too time-consuming or irrelevant to the current PR. Or, one can squash logically related commits together, add a useful message and merge them separately. The resulting commit log will still be bisect-able. For beginners: there's a very good article with tips for Git Commit Messages[0] that helped me have git histories I enjoy reading. [0] https://chris.beams.io/posts/git-commit/#why-not-how https://chris.beams.io/posts/git-commit/#why-not-how
- nimblegorilla 6y agoSeems like the real bug was in his handling of JwtSecurityTokenHandler. He claims to be an expert on dependency injection with two decades of automated testing experience. I wonder what was so hard about writing a test to cover this scenario?
- acdha 6y agoIt's not a good look to trash someone's career because they made a mistake. This happens to everyone everywhere — the only question is how well you handle it.
- nimblegorilla 6y agoPoint is that `squash` is a useful tool used by many other successful professionals. I'm allowed to disagree with the author's opinion of the root cause of his bug and the factors that made it hard or easy to debug.
- xorcist 6y agoThis article presents a very good argument clean and squashed commits are important, under a click-baity title. The ability to use git bisect effectively is one of the more important reasons to enforce a clean and readable commit history by squashing and rebasing before merge. A history where the majority of commits doesn't even compile ("sorry, updated test value was wrong", "oops syntax error", "forgot to update these references in the last commit", "big refactor wasn't complete") is a major headache not only to readers but to anyone using automated tools such as bisect. The author here was lucky the commit history was in a good shape. The size of the offending commit was reasonable too, so any trivial commits had been squashed away here. My personal issue with this particular commit is the useless commit message. "Extract CreateTokenValidationParameters method". Well, obviously. But why? What was this intended to result in? Why was particular change made and not something else? A more suitable commit message would have included something along the lines of "JwtSecurityTokenHandler methods belong conceptually with other code that configures JWT parameters. Break them out to CreateTokenValidationParameters because ..", that would have make much more sense and made the change easier to understand for someone else! After all, this is how the author describes the patch, when taking the time to do so in order to write a blog post. It isn't that hard.
- mumblemumble 6y agoI would say that this is a good argument for linearizing the history by rebasing, but certainly not for squashing. Git bisect can only narrow you down to the scale of your commits. If you do infrequent, large commits, it's not very useful. If you do frequent, small commits, it's great. If you do frequent, small commits and then squash them into infrequent, large commits, it's not very useful. There are plenty of reasons to squash (and, personally, I generally think they outweigh the arguments against), but this is not one of them.
- xorcist 6y agoSquashing is a form of rebasing. Nobody suggests commits should be too large to be useful, just that they should form coherent changes. Preferably self contained enough to be useful in their own. When you do frequent small commits, unless you are superhuman the majority of them false starts or contains errors. Squashing these useless commits makes the history understandable, and enables the use of tools such as bisect. That's what "please squash before merge" means. It does not mean you should do rebase instead of merge. There may be good reasons for that too, sometimes, but that's not the point here. Should you wish to enforce a linear history, it is still just as important to squash undesired commits.
- thrownaway954 6y agoi always squash my commits when merging a feature branch. i don't need all the frivolous commits i've made while saving my progress.
- jefftk 6y ago> Had this happened in a code base with a 'nice history' (as the squash proponents like to present it), that small commit would have been bundled with various other commits. This is a misunderstanding of what proponents of commit squashing advocate. Often, when working, you end up with multiple commits which represent a single change: remove debug logging and fix bug more logging debugging typo Squashing these commits together into one coherent change makes for a history that is much easier to understand. Squashing unrelated commits, however, doesn't help anyone. (I prefer the --first-parent approach: https://web.archive.org/web/20180710234754/http://www.davidchudzicki.com/posts/first-parent/ https://web.archive.org/web/20180710234754/http://www.davidc...)
- jasonkester 6y agoI can only sit in amazement as I read this thread, watching seemingly smart people advocating throwing away history for the sake of tidiness. It's maddening. Those are four commits. That's what happened. It does not matter that it's untidy. It's your history. Just leave it. It will save you a lot of work some day when you need to find the error you introduced when you accidentally removed one too many lines on that "remove debug logging" commit.
- saagarjha 6y agoThere’s “throwing away history” and then there is just clutter. If you added a feature and then ten commits to fix typos on top of it, just squash it into the original one and pretend like it was right all along. It keeps the repository sane and not full of “typo fix” commits, and, like, I don’t really mind looking for a change in a diff that is 20 lines instead of a dozen that are one.
- kordlessagain 6y ago> It will save you a lot of work some day I already spend a fair amount of time on other people's problems and they aren't doing much to save me, so what is the point of being obsessive with my own problems such that I may avoid spending a little bit of time on them later?
- 6y ago
- BeetleB 6y agoMercurial's evolve extension has fold, which is similar to squash. However, it still has all the individual commits if you wish to examine them - they're just "hidden" and the usual commands (logs, etc) will show just the squashed commit.
- aayjaychan 6y agoGit does hide the old commits as well. What git doesn't do is track which commits are replaced by which. So sharing mutable history and seeing how a commit evolved over time requires more heuristic than necessary.
- BeetleB 6y agoI was under the impression that those individual commits eventually get lost (i.e. may be in the reflog or not sent to the server, etc). With Mercurial's evolve, the hidden commits are always there. When you push/clone, etc they get sent around.
- aayjaychan 6y agoYes, unreachable commits in Git are GC'd eventually, but you can disable it. In Mercurial, hidden changesets are kept locally indefinitely, but they are not exchanged; only their obsolescence makers are. So you always know the meta-history of a changeset, but not necessarily their original content.
- kelchm 6y agoGreat discussion in this post. While I do generally prefer squashed commits when merging a branch, I absolutely seen that cause some pain in specific circumstances. I think part of the problem is that modern repository software has built an 'additional layer' on top of git (IE: pull requests) that most of us have become accustomed to using in our day to day workflows. Git itself doesn't really have a 'lossless' way to persist that extra information. Squashing commits at merge gets us close, but ultimately we’re throwing away potentially useful information by doing so. The main time I’ve seen it be problematic is when multiple people are working on a complex feature where PRs are getting merged into a long-lived feature branch rather than directly to the default branch. Say I need to base my work on someone else’s WIP branch for said feature. If they squash their commits when merging their work into the feature branch, I’m going to end up with merge conflicts that need to be manually resolved as the individual commits that I had based my work on no longer exist, as they’ve all been squashed into a single new commit. What if git supported a sort of 'soft squash' concept, where associated commits could (optionally) be grouped together with additional metadata. That would let the application consuming the git history (IDEs, CLI tools, etc) make the decision in how commit histories are presented to the user.
- Vinnl 6y agoI usually understand "squash" to mean "bundle everything in a PR into a single commit". Which can indeed break bisect workflows. There's a middle ground though: rebase your commits, but not necessarily into a single one. Before I submit a PR, I rebase my branch (which has lots of small commits, some of which undo previous work or are a work-in-progress), and make sure that every commit is as small as it can be without including work that is halfway done (so all tests still succeed), and have a clear description of what they're doing. I then have both a nice history, and I can bisect to find a problematic commit, or inspect the commit history of a single line to get more context about it.
- takeda 6y agoSame here, I especially love `git rebase -i <base>` it opens an editor and I have an option to edit specific commits, reword the commit, squash multiple commits together (don't use this one often) or do fixup (which basically merges commit to the previous one)
- corytheboyd 6y agoInteractive rebase is the only way I do this now too, it’s so intuitive!
- seba_dos1 6y agoI believe that if you don't know interactive rebase, you don't know git ;) Especially "git rebase -i -r" is an incredible tool.
- sixstringtheory 6y agoIf you haven’t yet, check out rebase’s --autosquash option, along with git commit --fixup or the commit message directives “fixup!” and “squash!”
- Vinnl 6y agoI just learned about that recently, but haven't been able to put it in practice yet. The problem I encounter is having to specify which commit to fixup to when calling `--fixup`, which I think means having to look at my commit history. Instead, my process so far includes just describing the commit I want to fixup in my commit message, in a way that makes sense to me when I'm going through the interactive rebase. Do you have a good way to deal with this?
- mpawelski 6y agoA lot of people seems to suggest that squashing PR's branches is better because you don't get commits that are broken and it doesn't work good with "git bisect". But IMO it's not a big deal, you just do "git bisect skip" for these comments and at the end of bisecting session instead of one wrong commit you'll get couple of, for example one commits that builds and two that you skipped and doesn't build. It's very likely that looking only at this 3 commits will help you find the issue because they might be part of much bigger branch which would otherwise be squashed to one huge commit with enormous diff. I much more prefer to require that PR's branches are always merged without fast forwarding. So there is always merge commit for PR. Then you can actually display list of this commit with "git log --first-parent" and they should always build because build server verifies it. Unfortunately "--first-parent" doesn't work for "git bisect" now, but it finally will in the next git release! [1] [1]https://github.com/git/git/blob/ab4691b67bc1a2cd8d9068fb03e3fbd6979247d6/Documentation/RelNotes/2.29.0.txt#L28 https://github.com/git/git/blob/ab4691b67bc1a2cd8d9068fb03e3...
- nneonneo 6y agoIs there a way to make git keep the unsquashed history, linking it to the squash commit? That would solve the problem here, because anytime bisection locates a squash commit you could just follow the pointer to the unsquashed history and repeat.
- SassyGrapefruit 6y ago>I've always disliked Git's squash feature, and here's one of the many reasons to dislike it. Had this happened in a code base with a 'nice history' (as the squash proponents like to present it), that small commit would have been bundled with various other commits. This question always boils down to this metric. Which happen more often... 1. A circumstance arises where the dynamics of fine grained commits make the problem more obvious. 2. I have to interact with the git history Now(and this is just me) I interact with the git history for one reason or another multiple times a day every day. I have been using git for 10+ years and I haven't yet encountered the first circumstance yet. Its not to say I won't encounter that circumstance and when I do I'll probably pine for the fine grained commits that would make it stand out. For me the clean, easily browsable history benefits me every day. Furthermore because of how often I have to use it I should optimize it heavily at the expense of just about anything else unless the benefits of that item can be realized with similar frequency. The access to fine-grained commits hasn't helped me once in the 10 years I've been using git. With that calculus in mind I must conclude that I should squash the commits and pay the piper on the other thing when/if that bill comes due. EDIT: Imagine this. If clean browsable history saves me 5 minutes a day then it has saved me ~10,000 minutes since I started using git. That equates to 1 full time working month. I really can't think of a circumstance where having access to fine grained commits would deliver a similar net savings.
- jkubicek 6y agoHaving fine-grained commits and having a clean git history aren't mutually exclusive. If you create and submit small PRs and git-squash those PRs into the main branch, you'll get the best of both worlds. FWIW, the way I use git results in me committing a lot. I treat it as, essentially a save point that I can undo to if I need to. The history for my personal branches is a mess of broken tests and false-starts on code paths that just aren't going to work. Useless for anyone but me (but fantastic for giving me an hour-by-hour breakdown of what I've been working on and what attempts I've made).
- SassyGrapefruit 6y ago>If you create and submit small PRs and git-squash those PRs into the main branch, you'll get the best of both worlds. I definitely advocate for this. But I've worked some at some spots with "never squashers" that deliver every PR as 119 commits and tell me this yarn about "the time having all those commits totally saved their bacon". I've just never regretted squashing my commits into something manageable.
- carapace 6y agoForgive me if this is a dumb question, git squash deletes information?
- cameronfraser 6y agoThis reads like the articles of people who are afraid of rebasing
- ed25519FUUU 6y ago> Clearly, you're not familiar with this code base, but even so, you might be able to intuit the problem from the above diff. You don't need domain-specific knowledge for it, or knowledge of the rest of the code base. A question I ask, is that if I reverted/rolled back to this commit, would the system still work? Commits that require other commits to work should never be in isolation in case a rollback is required. You should be able to check out any commit in the entire system and have a (hopefully) working program.
- fchu 6y agoYou can always "bisect" through space (execution flow) OR time (commits) to find a bug. I usually prefer through space, because 1. it actively helps narrow the buggy line in the code you're working on, instead of reflecting over different snapshots of code 2. sometimes you have nothing to bisect on, if it's code you're actively writing. The author seems to have started bisecting through space using the debugger. Unfortunately they had the wrong understanding of the execution flow and quickly stopped after one step at the start of the route. Had they realized the route wasn't triggered they could have checked what happened in the authentication code.
- ptx 6y agoAre you saying he should have stepped through the framework's authentication implementation in the debugger? The bug was in the configuration, if I understand correctly, so there wasn't much to step through there.
- SPBS 6y agoThe main focus shouldn't be on keeping the commit small, it should be on ensuring that the commit comprehensively encapsulates a unit of change in the codebase. Dogmatically pursuing such small commits will lead to a noisy commit history, and all for the meagre payoff for when such niche situations occur. BTW if such a thing really does happen in production and you already squashed and merged, you can always fall back on on git reflog on the developer's machine. Commits are never really destroyed in git, the branch simply shifts focus to another set of commits. The old commits are still there, dangling and not referenced by any other branch but they are still reachable.
- waffletower 6y agoSome software engineers can be fanatical and rule-bound; commit etiquette is merely one milieu where rigid thinking can be applied, and with the usual detriment. I applaud the author for presenting a clear case how reality is more complex than many beginning to intermediate developers would like it to be when it comes to commit history.
- sys_64738 6y agoWhen you commit your fully tested and bug free code to the master branch, developer WIP commits should be squashed. It's one thing to commit your WIP commits to a toy git master branch you control but as soon as others look at it, it doesn't scale. I've asked this question in job interviews. It's amazing what people say.
- waffletower 6y agoSad that you use your strict etiquette as a leading question in job interviews. How do you know that code is bug free when you commit to master? What qualifies as fully tested? If you are convinced that there are absolute and known answers to these questions I would suggest analyzing your logic again.
- sys_64738 6y agoNot squashing commits doesn't scale. That's the key to discern who has experience on large projects. Intermediate commits are of no interest to a codebase with dozens of developers.
- jcorrington 6y agoI've really never understood this obsession with "preserving" history. When I work on a feature or bug I tend to commit often, because it's cheap, and it can save my ass if I go in a bad direction and break something that used to work in one of my previous commits. I subscribe to a code fast philosophy, and work quickly to explore the space, find unknowns, get a working solution, write tests, and do final cleanup, docs, and optimize if needed. As such this would be a common commit history for me that would end up in a small PR. A large PR for me might have 50+ commits. - initial prototype of feature x, mostly working - feature x working as intended in requirements - fixed edge cases not originally identified when thinking about feature x - rearchitect feature x a bit now that it's better understood - write tests for feature x, most pass - get all feature x unit tests pass - a few more test cases for feature x - and a couple of integration tests for feature x - code cleanup for feature x - fix typo from last code cleanup resulting in bug - better docs for feature x - fix typos in feature x docs - update main readme to include notes about feature x Can anyone explain to me the value in preserving this history, because I'm just not seeing it, and it would completely muddy up the overall git history. There's tons of intermediate work we're not concerned about preserving, like the notes I scribble on paper, so why are we so concerned about preserving all these commits? I admit there's a small chance it's useful in rare circumstances, but I'd rather optimize for the common scenario.
- blntechie 6y agoUnrelated to git but as soon as the authorization was failing with even a correct JWT token, the first thing I would have looked at is the AuthorizeHandler and registration stuff in Startup.cs. It got crystal clear when the breakpoint was not hit or when removing the Authorize attribute worked. The repeated saying of my tests were passing and it must be the framework kind of annoyed me as the Authorize attribute is not magic and there need to be wiring stuff which need to be written.