13 ms·
The Tyranny of the Diff
- TazeTSchnitzel 15y agoThat would be an interesting feature to add to diffs - if there are huge structural changes in a block, just show before and after instead of trying to show differences. No idea how you would define the criteria or implement it though.
- _delirium 15y agoI'd like an option to collapse additions or removals of entire functions into just one line saying "foo() added" or "foo() removed", instead of the 50 lines of the function's body scrolling past. However that does have to be language specific.
- pnathan 15y agoI like using a UI to do large diffs in a side-by-side fashion.
- knowtheory 15y agoThe point is fair enough, but to call this a tyranny seems a bit much. If one believes (as say that Literate Code movement does) that code alone isn't sufficient to convey an author's intent, why should diffs alone, which mark changes in a codebase, be any different? Comprehensible commit messages and comments explaining the purpose of a method itself (and perhaps the mechanism around which a block of code functions if necessary) are no vice and go a long way to mitigate or counteract any such "tyranny".
- dmlorenzetti 15y agoI often stage changes with an eye toward making the diffs clearer -- not for myself at commit-time, but for myself and colleagues in the future.
- bricestacey 15y agoIf you use vim, try the fugitive plugin[1]. You can do :Gdiff to see side-by-side diff between what is staged and the working tree. It provides some context, but if you need more you can unfold it. [1] https://github.com/tpope/vim-fugitive https://github.com/tpope/vim-fugitive
- sophacles 15y agoMore generically, many diff viewers exist, and certainly help understand what is going on with a diff, in a side by side comparison. However, in a multi-file refactoring, sometimes this can still be a bit tricky. Say for instance you split one method into 3 smaller methods that are somewhat inter-dependent (you have x(), but now it is a(), b(), c(), and you sometimes call a(); c() and sometimes b(); c(), and sometimes just c()); in that case you have a complex diff that is hard to make sure everything is correct in. Even if you can turn it into a series of changes keeping x() as a wrapper around a(), b() and c(), you'll frequently end up with several multi-file diffs to look at as you migrate. I somewhat agree with the author, that maybe there is some better change viewing paradigm we aren't seeing for these complex cases, that could benefit everyone.
- Chris_Newton 15y agoIsn't this a fundamental difficulty with any flat text-file representation for code, though? Many programming languages don't lend themselves to easy semantic analysis based on even a single file, never mind a combination of files. We usually don't work with canonical textual representations of even quite simple programming ideas such as the example in the blog post of replacing a nested condition with a guard clause. At our current level of expressiveness, the refactoring tools that make these semantic changes mostly aren't a huge step up from brute force text editing, and the mechanical changes we can automate are usually trivial compared to the way a developer might conceive a required change in the behaviour of his code. Trying to reverse the process to identify the semantic significance of changes after the fact is far beyond us today. In fact, it's hard to see how we could ever move beyond that level without developing some new and much more semantically rich way to represent our programs. At that point, the idea of diffing raw text code side-by-side might seem like it comes from the dark ages anyway...
- marshray 15y agoThere have been a thousand graduate school and commercial project to try to develop program representations and languages that were superior to the flat text file. Obviously, none of them have received wide usage. What has succeeded for many is scriptable text editors, cscope-like tools, autocomplete-like features, and refactoring support in editors. But as you point out, these are better tools to work with program text. I think text is safe for the forseeable future. Partly because we have thousands of years experience with it and partly because it's a winning combination of a extremely powerful representation format and a KISS solution.
- Chris_Newton 15y agoI agree that, for better or worse, we'll be using text-based formats for a while. I don't think the hypothetical alternative representation I mentioned is going to happen off the back of a single grad school research project, or anything even close to that scale. It's more like something that's going to take the R&D lab of an industry heavweight a decade to develop and refine and then launch into widespread industrial usage with the backing of at least one major platform developer when it's ready for prime time. Sadly, as long as text-based (or, perhaps more accurately, line-based) programming languages are good enough to produce acceptable software, there isn't a truly compelling motivation to develop a completely new model. I wonder whether increasing pressure on the software industry to provide quality and security as a backlash against the current cheap-and-nasty trend will drive us toward more radical programming models first, and perhaps get us close enough to make the quantum leap to an entirely new kind of representation from there.
- dochtman 15y agoSplit up your changes. Seriously: in most cases it's possible to split up a 30-line patch into 5 separate patches that, applied in order, monotonically improve your code base and are easier to review, in total, than one larger patch. Changesets are cheap, and we should be optimizing them for easy review, so any eyeball we get can see what's going on.
- morsch 15y agoWhen I'm ready to commit, I usually have solved the problem in code. In order to create five separate patches that illustrate the line of thinking, wouldn't I have to go back in time and re-create the intermediary steps? I suppose I could create patches as I am actively solving the problem, but at that point my code may very well be a mess that I'd have to clean up each time. Of course all of that is moot if your problem naturally segments into several patches, if nothing else than simply by the virtue of being larger than a 30-line patch or involving several mostly independent components.
- cpeterso 15y agoIf you've completed a big change and can't stage it in separate commits that build upon each other ("telling a story" of the feature development), I recommend at least splitting non-overlapping chunks that can stand (compile/test) independently into their own commits. git-cola is a good GUI to visually stage chunks into separate commits. I haven't really used git-cola's other functionality, but I really like its visual staging features. http://git-cola.github.com/screenshots.html http://git-cola.github.com/screenshots.html
- drothlis 15y agogit-gui (part of the core git distribution) provides similar mousey-clicky staging -- you can stage hunks or individual lines at a time.
- drothlis 15y agoBasically, yes. It does mean more work for you, for the benefit of the reviewers of the code. Arguably, it is more important to optimise for the readers of the change (many people, spread over time) than for the writer (one person, once). If you use, say, git, you might commit every single, tiny change separately, on a private (local) branch; then use `git rebase --interactive` to reorder/merge the commits as necessary. This is the easiest way I have found, but it still involves more work. If you wanted to take it further, you could have your editor automatically commit on every save, with a post-commit hook that takes your changes into a staging area, compiles/tests, and provides a report for each commit.
- naner 15y agoI realized that I didn't want to look at the diff any more. I just wanted to see the full body of the affected method before and after my changes. Oh, and I also wanted to see whether I added new tests or changed existing ones. Well there are multiple ways to do that. This doesn't really have anything to do with the diff but how you chose to view it.
- sliverstorm 15y agoAbsolutely, I was just thinking that. Many version control systems support an arbitrary diff command, e.g. via environment variables. In that case, one might be able to try tkdiff, for example.
- davvid 15y agogit-difftool lets you plug in your own viewer. tkdiff, xxdiff, meld, kdiff3, and many others are built-in. You can also configure your own. http://schacon.github.com/git/git-difftool.html http://schacon.github.com/git/git-difftool.html
- zdw 15y agoDiff works best when things are single idea per line, and when control structures don't get in the way. One example - the use of the trinary (?:) operator as a replacement for if/else assignment statement can quick and easy when programming. The problem is that diffs with it can look like a total mess because a whole lot of things are happening in that one line. Similarly, certain languages where program flow or control structures make it so that people are inclined to make many things happen in one line (inline regex, lisp or scheme syntax lanagues) can diff in a confusing manner. Diff is great when everyone is using a coding/whitespace standard, and things tend to atomically happen on one line. I'd encourage that when refactoring code that's not to the standard, you do two passes - one to clean up the code to the standard, then another to make the actual changes.
- gbog 15y agoYes, and commit on every smallest atomic refactorings.
- stcredzero 15y agoI've used environments with quick and reliable Undo/Redo stacks. To be completely compatible, one would have to treat code as data in a "nondestructive" editing scheme and save off the actual refactoring steps, which can then be committed automatically. Or maybe, a "quick snapshot" facility could be developed to make it easier to save off intermediate steps and commit the series of them automatically.
- makecheck 15y agoI definitely agree with doing formatting/nonfunctional changes independently from other changes. In general all commits should be as focused as possible: have one purpose, don't just change whatever else you happened to think of while editing the same file. A "diff" is also incredibly useful prior to a commit to make sure that you changed only what you thought you did. Programmers should go to great lengths to adopt a style that is "diff compatible" so they aren't likely to miss anything important.
- memset 15y ago
- _cavalle 15y agoFor the reasons described in this post I tend to separate in different commits those changes that alter functionality from those that are refactorings. Sometimes I refactor before making my changes, and sometimes I do it afterwards. In any case I try not to mix a change in the functionality and some refactoring in the same commit. That way, in retrospect, it's easier to me to understand each commit: the ones related to changes in functionality have simple, easy to understand diffs, and the ones related to refactorings have messy diffs but at least I know that they don't change any functionality.
- zvrba 15y agoI have the same problem when maintaining Latex files in some RCS. After any "big" change (rewritten sentence, etc), I customarily reformat the text in emacs so that it looks nice on screen, which also totally messes up the diff. emacs has a mode which allows one "logical" line to wrap and to be edited as many "physical" lines, but when I tested it few years ago, it was rather broken. Fortunately, for editing Latex and such, I don't really need the diffs, I'm just interested in archival.
- drunkpotato 15y agoEmacs virtual-line-mode has gotten much better; check it out again if you're interested. What I do with my Latex files is have one sentence per (logical) line. I've found diffs at the sentence level much more helpful.
- renata 15y agogit add --patch can also help with this. You don't necessarily need to make all your changes simultaneously.
- pwpwp 15y agoOne of the most interesting things re diffing/VC I've seen in a while is "Towards Structural Version Control" https://www.cs.indiana.edu/~yw21/slides/ydiff-slides.pdf https://www.cs.indiana.edu/~yw21/slides/ydiff-slides.pdf
- ExpiredLink 15y agoSimple suggestion: Don't change the old code. Copy the old method and then rename the old method. Refactor the copied code. Diff will only mark your 'newly added' code which will check in without conflicts. Later remove the old method.
- makecheck 15y agoMost languages and file formats don't enforce an inherently "undiffable" layout, so it's really an issue of programming style. People tend to lazily do what's easiest to write instead of thinking about what will be easiest to read later (in "diff" or otherwise). A common example I see is something like a one-line list of files to build in a makefile. It is certainly possible to put every file on its own line and backslash-escape each line ending, and doing so produces a very readable "diff": if someone adds a file you see "+ xyz.c" (or whatever) and that's it instead of a mangled mess of file lists repeated with one word that's different.
- wnoise 15y agoA lot of people have mentioned ugly diffs when lines are considered the basic unit of granularity, often giving examples that are much cleaner if words are considered fundamental (e.g. latex, or lists of files in makefiles). But there are many tools that can handle word-based diffs, and most have options to change what's considered a word. git diff can use --word-diff(=color) and --word-diff-regex=... There is also the venerable "wdiff" program. I've only found one program that can do "word patches" though: "wiggle" http://freecode.com/projects/wiggle http://freecode.com/projects/wiggle . It works well, though I've found the interface to be slightly confusing. But turning it into a git diff and merge driver isn't that hard.
- michaelfeathers 15y agoI just want diffs at method scope. Show me the methods that have been added, changed, or deleted. Use line diffs for things external to methods.
- keypusher 15y agoUse meld, or any one of the many other excellent visual diff viewers.
- slurgfest 15y agoThe problem isn't that the diff has "tyranny". The problem is that what you are diffing is bytes instead of the program's AST.