10 ms·
My new Git utility `what-changed-twice` needs a new name
- chris_wot 1y agoYou know, I find myself partially agreeing that a number of utilities for git could be done quite nicely in perl.
- eru 1y agoGit's repository includes quite a bit of Perl, but they want to get rid of it.
- chris_wot 1y agoIs there any reason for doing so?
- kruador 1y agoIt's a pain in the backside to run on Windows, for two reasons. Firstly, Windows doesn't have (by default) a lot of the tools that are preinstalled in most nix environments. Git for Windows ships half a Cygwin distribution (MSYS2) including Bash, Perl, and Tcl. Second, Windows doesn't really have a 'fork' API. Creating a new process on Windows is a heavyweight operation compared to nix. As such, scripts that repeatedly invoke other commands are sluggish. Converting them to C and calling plumbing commands in-process has a radical effect on performance. Git for Windows is more of a maintained fork than a real first-class platform. Also, I believe it's a goal to make it possible to use Git as a library rather than as an executable. That's hard to do if half the logic is in a random scripting language. Library implementations exist - notably libgit2 - but it can never be fully up to date with the original. Search for 'git libification'. Many IDEs started their Git integration with libgit2, but subsequently fell foul of things that libgit2 can't do or does inconsistently. Therefore they fall back on executing `git` with some fixed-format output.
- 1718627440 1y agoI don't get why everything needs to be a library? Using the OS to invoke things gets you parallelism and isolation for free. When you need to deal with complicated combination of parameters to an API, it doesn't become too different from argument parsing, so you might as well do that instead. You can still wrap the interface to the executable in a library.
- codewritero 1y agoJujutsu has a command which is helpful for this sort of workflow called absorb which pushes all changes from the current commit into the most recent commit which modified that file. (Each file may be merged into a different commit).
- metadat 1y agoYes, totally useful compared to default git base commands. And also - melding the "changed twice" (or thrice...) mutations into a single commit is a brilliant isolation of a subtle common pattern.
- goku12 1y agogit-absorb does exist [1]. It seems to be inspired by a mercurial subcommand of the same name. It's also available in most distro repos. [1] https://github.com/tummychow/git-absorb https://github.com/tummychow/git-absorb
- koterpillar 1y agogit-absorb (https://github.com/tummychow/git-absorb https://github.com/tummychow/git-absorb) does a bit more, figuring out the exact changes that should be fixed up.
- CorrectHorseBat 1y agoJJ absorb does the same as far as I understand
- globular-toast 1y agogit-autofixup is better and easier to install: https://github.com/torbiak/git-autofixup https://github.com/torbiak/git-autofixup
- operator-name 1y agoCould you elaborate how it is better?
- MontagFTB 1y agoI am familiar with an algorithm that stably brings a disjoint selection of items together around a specified point. Sounds similar to this case, where the disjoint selection are changes that happened to a given file. The name of the algorithm is “gather”, by Sean Parent and Marshall Clow.
- JKCalhoun 1y ago"muster" comes to mind and is different than "gather".
- quuxplusone 1y agohttps://github.com/stlab/adobe_source_libraries/blob/765924419421c1d76c4c4abb9761141962ef9222/adobe/algorithm/gather.hpp https://github.com/stlab/adobe_source_libraries/blob/7659244... https://listarchives.boost.org/Archives/boost/2013/01/200366.php https://listarchives.boost.org/Archives/boost/2013/01/200366... I gotta say, I don't see the greatness any more than most of the repliers in that Boost thread — it's just two stable_partitions in a row. "[...] Or is there some optimization that gather provides over (stable_)partition? —— Nope. [...]"
- MontagFTB 1y agoThe Boost thread starts with an example of how Bjarne replaced a bunch of complicated code with it. It may be just two stable partitions, but “just” is doing a lot of work there. The algorithm becomes obvious once someone has identified it.
- quuxplusone 1y agoThe talk: https://www.youtube.com/watch?v=OB-bdWKwXsU&t=52m49s https://www.youtube.com/watch?v=OB-bdWKwXsU&t=52m49s Sadly the 25-line original code isn't presented; the code that is presented is the 5-line replacement using the STL's `find_if` and `rotate`. Bjarne sketches the idea that those five lines can be further condensed into two lines with the non-STL `gather` algorithm: auto dest = std::find_if(v.begin(), v.end(), contains(p)); stdx::gather(v.begin(), dest, v.end(), [](const auto& elt) { return &elt == &*source; }); But this is overkill — replacing an O(distance(source,dest)) non-allocating rotate with an O(v.size()) potentially-allocating stable_partition — and more importantly it re-complicates the code. Now, I think part of his point is that `stable_partition` is "simpler" than `gather` only because it's in the STL. If we add `gather` to the STL too and everyone learns what it means, then there's no objection to using `gather` for "simplification" like this: it would be a straightforward simplification in almost the same way that `std::equal_range(first, last, x)` is a straightforward simplification of `std::make_pair(std::lower_bound(first, last, x), std::upper_bound(first, last, x))`. The "almost" is that actually there is an algorithmic advantage to `std::equal_range`: when you're looking for the upper bound, you don't have to consider any of the elements to the left of the lower bound you already found. You get a (very slight) performance boost by using the combined `equal_range` algorithm. `gather`, on the other hand, has no such advantage; and (as we've seen) has a (very slight) performance disadvantage when compared to the `rotate` that Bjarne's correspondent's code actually required. We're not talking about replacing 25 lines of bespoke code with 1 line of Boost `gather`; we're talking about replacing 2 lines of STL `stable_partition` with 1 line of Boost `gather`. The former is probably worth it. The latter is not.
- eru 1y ago> There's bonus information too. If a commit is not mentioned in the report, then it only changed files that didn't change in any other commit. That means that in a rebase, I can move that commit literally anywhere else in the sequence without creating a conflict. Only the commits in the report can cause conflicts if they are reordered. This is only true in the textual level. Semantically, re-shuffling commits like this can still cause conflicts. Ie it can break your tests. Not at the end, but for the intermediate commits.
- _dark_matter_ 1y agoThis is why I no longer do atomic commits. I've just never had it be a benefit to walk through and guarantee that each commits tests and builds successfully. I so rarely back out changes that when I do, I test then that everything is working (and let's be honest, I back out usually at the PR level, not the commit).
- mjd 1y agoI agree. I decided years ago that that was a lot of work for little or no benefit. It's enough for the tests to pass at each merge point.
- baq 1y ago…and that’s why squash merge should be the default setting in PRs.
- eru 1y agoYes, it should be the default, but ideally you have the option of preserving history (for PRs where that makes sense) and then your CI/CD should also check that the individual commits build and pass tests. In general, your CI/CD should make sure that each commit that appears in the 'public' history of main builds and passes tests.
- WorldMaker 1y agoYou can `git bisect --first-parent` just fine without needing to squash.
- nicr_22 1y agoFlipFlopStop? FFS for short, which has suitably disgruntled other exclamatory meanings.
- allseeingimei 1y agogit-delta -n <times> i.e. git-delta -n 2 = 'what changed twice' or if its just what changed twice in every case then just 'git-delta-delta'
- GuB-42 1y agoWhy does it needs a new name? I had a good idea of what it did before reading the article, it is a long name but not Java-long, and none of the suggestions so far are clear to me, even after reading the article. The only somewhat confusing part is the "twice", because it can be more than twice. But if you think about it, if it has been changed more than twice, it had to be changed twice at some point, so it is not totally wrong.
- mjd 1y agoAt the time I started writing the article, the utility was called `analyze-commits`. Hard to think of a worse name than that! By the time I finished writing it I had come up with a less crappy name, but I thought I'd leave the question in the post anyway.
- antonvs 1y agoIf you’re looking for something descriptive and not clever/catchy, I propose ‘find-repeat-changes’.
- deleted 1y ago[deleted]
- 0manrho 1y agoJust gonna +1 this. It's still fairly short, descriptive and to the point, which I generally prefer to something more "trendy" or "clever". I like it.
- alex-moon 1y agoOr indeed "Find Repeat EDits" or fred for short.
- 1718627440 1y agoWhat about git n-changed or even git nchanged. I feel like these commands need to be short and not consist of >3 words.
- squeaky-clean 1y ago"what-changed-twice" tells me exactly what the command does. "squash-what" tells me nothing, why is the program name asking me what to squash, and then why does it not squash? The only inaccuracy I can think of in the name is that it's technically "what-changed-more-than-once." But if something has changed thrice, by definition it's also been changed twice.
- protocolture 1y agoDouble Jeopardy?
- zahlman 1y agoWhen I make Bash aliases or functions for Git functionality, I always name them as `git-something-or-other`. That way they're namespaced in a way that I find pleasant both for tab completion and for easy of memory. I think that should apply to more complex utilities, too. By my usual naming conventions, this one would be `git-repeatedly-changed`.
- nothrabannosir 1y agoLast but decidedly not least: if you have `git-foo` on the PATH, you can do `git foo` and it will automatically pick up your program. If I remember early git days correctly, that's how git was implemented: a bunch of separate utilities working together on the database which is the .git folder.
- gavmor 1y agoThese are called alternative "porcelains:"[0] third-party, user-friendly interfaces built on top of Git's stable, low-level plumbing commands. 0. https://git-scm.com/docs/git.html#_low_level_commands_plumbing https://git-scm.com/docs/git.html#_low_level_commands_plumbi...
- mjd 1y agoI usually do that too, but this seemed to me like it's not really a git utility. It's just a filter. I can see the argument in favor of `git-` also. But I think I'd prefer `git-changed-twice` to be a wrapper that takes a reflist argument, and runs `git-log --stat reflist | what-changed-twice`.
- st3fan 1y agoOidia- Oops I did it again
- quuxplusone 1y agoSuggestion: `git squash-report`. (Or `git rebase-report`, except I wouldn't call it that because it would interfere with my tab-completion of, and/or muscle memory of, `git rebase -i`.)
- atoav 1y agoNo it does not.
- pfannkuchen 1y agochange-cluster?
- gorgoiler 1y agoTools like this are also useful if you need to cherry pick a patch onto a release branch and want to know potential dependencies: ↑ newer D* fixes bug in crypto.py C B* rewrites crypto.sh in Python A 0 last month’s release ↓ older In this example, if the release needs the fix in D you’ll also need to cherry pick the rewrite in B. You get false positives and false negatives: if B fixed a comment typo for example it’s not really a dependency, and if C updated a module imported in the new code in D you’d miss it. (For the latter, in Python at least, you can build an import DAG with ast. It’s a really useful module and is incredibly fast!) So I would say the author’s tool is really multiple tools: 1/ build a dependency graph between commits based on file changes in a range of commits; 2/ automate the reordering and squashing of dependent commits on a private dev branch; 3/ automate cherry-picking commits onto a proposed release branch (which is basically the same as git-rebase -i); and 4/ build a dependency graph based on external analysis (in my example, Python module imports) rather than / as well as file changes. Their use case is (1) and (2), (3) is a similar but slightly different tool to (2), and (4) is a language specific nicety that goes beyond the scope of simple git changes for, arguably, diminished returns.
- handsclean 1y agoI suggest group-commits-by-file , group-commits , or group-by-file, depending on whether you want it to make sense out of context and whether you ever group commits differently. You might then feel compelled to add a final line like “… and 12 files with 1 commit each”, or even to enumerate them, which sounds like it’d be useful anyway. “what” isn’t doing any work, there’s already an implicit “what” in the call-response paradigm. “Changed” implies you’re detecting changes, but you’re not, you’re operating on a data structure that happens to represent changes.
- nferraz 1y agoWhy did you opt for "highly-abbreviated commit IDs"? Instead of: ``` calendar/seasons.blog 196 40 d1 196 196e749 40 40c52f4 d1 d142598 ``` The tool should simply display: ``` calendar/seasons.blog 196e749 40c52f4 d142598 ``` That's it! The second table only complicates the output. PS: `what-changed-twice` is a good name.
- paulddraper 1y agoWebsite is down. https://archive.ph/52C1y https://archive.ph/52C1y
- cozzyd 1y agooops-i-did-it-again
- perfmode 1y agoYou could shrink the prefixes in your report. 40 and 33 could become 4 and 3 without losing correctness.
- mjd 1y agoThere were commits in the original log input for which 4 and 3 would have been ambiguous, and the abbreviations are already short enough.