17 ms·
What if Git worked with programming languages?
- afavour 5y agoI do kind of love the idea of Git using ASTs instead of source code. It makes a ton of sense. Even just in the immediate term I wish I could make Git(hub) tabs/2 spaces/4 spaces/whatever agnostic. Seems crazy to me that in 2021 we still have to make opinionated choices across orgs about what to use... why can't we pull the code down, view it in whatever setup we want, then commit a normalized version? [whispers] this is actually something tabs allow you to do natively by setting custom tab widths in text editors but I've given up trying to sell people on tabs at this point and just want to be able to do my own thing
- enriquto 5y ago[whispers] don't give up! There's quite a bunch of us. Our day will come! Long live glorious tabs!
- silon42 5y agoI'm fine with using tabs, but my tab width will be set to 8... be sure to obey line length limits with that in mind.
- jcelerier 5y agoanything beyond 2 is heresy, and some days I'm tempted to go down to 1
- enriquto 5y agoHeretic! From the book of Linus [0], chapter one: > Tabs are 8 characters, and thus indentations are also 8 characters. There are heretic movements that try to make indentations 4 (or even 2!) characters deep, and that is akin to trying to define the value of PI to be 3. [0] https://www.kernel.org/doc/html/latest/process/coding-style.html https://www.kernel.org/doc/html/latest/process/coding-style....
- jcelerier 5y agoof course PI isn't 3, it's 1 (from a distance)
- a1369209993 5y agoOnly if you're a cosmologist.[0] 0: http://xkcd.com/2205/ http://xkcd.com/2205/
- klyrs 5y agoCursed April fools update: tabs are now π spaces wide.
- giomasce 5y agoMath trivia: there are cases on which it is sensible, in sufficiently advanced mathematics, to define pi as 3 (or whatever other number). I don't use tabs, but if I'd say that the biggest advantage of using tabs is that everybody can configure their own editor to make them as large as they wish.
- mellavora 5y agoThree shall be the number of the counting and the number of the counting shall be three. Four shalt thou not count, neither shalt thou count two, excepting that thou then proceedeth to three. Five is right out.
- deleted 5y ago[deleted]
- encryptluks2 5y agoThere are no line limits. That is what word wrapping is for.
- convolvatron 5y agohaving presentation by flexible and different than the underlying model is a great idea for code but admit it, tabs are fragile and a pretty weak implementation
- wutbrodo 5y ago> admit it, tabs are fragile and a pretty weak implementation Could you elaborate? I don't have a personal opinion here and have only worked in orgs that require spaces, but I'm not familiar with the criticisms of tabs.
- gregmac 5y agoFor me the problem happens as soon as tabs are used for alignment, instead of just indent. The benefit of tabs is custom tabstop. If anyone does anything that undermines that benefit, you might as well use spaces to avoid all the problems caused. Consider the following code: if (x) { SomeMethod(paramater1, paramater2, parameter3); } If done "properly", it is: if (x) { <tab>SomeMethod(paramater1, <tab><spaces...>paramater2, <tab><spaces...>parameter3); } What I often see, that totally breaks the entire point of tabs: if (x) { <tab>SomeMethod(paramater1, <tab><tab><tab><space><space>paramater2, <tab><tab><tab><space><space>parameter3); } The same thing happens if you are trying to align table-style code: var badMixedTypeArrayExample = [ [ "some", true, 128, x ], [ "long strings", true, 8, someLongVariable ], [ "and", false, 16384, x ], [ "short", true, 12345678, anotherVariable ], ]; If tabs are used between fields, it will look like a hot mess to anyone with a different tabstop than the author.
- cool_scatter 5y agoWhich is the reason for the very common stance "tabs for indentation, spaces for alignment".
- 5y ago
- williamdclt 5y agoIt's not that you're going too far, it's that you're not going far enough! It's not a Git question, it's a programming language question. There's no reason source code need to be stored as plain text[1]! Editors show it as text, we edit it as text, but why wouldn't it be _stored_ as an AST? Not only does formatting becomes an editor concern, but code could even be edited as a tree, as a graph, as whatever you want [1] - well, actually there's plenty of reasons: chiefly because plaintext is very interoperable
- jerf 5y ago"but why wouldn't it be _stored_ as an AST?" It profoundly is. You can't store "an AST". You can only store a serialization of it. The official language grammar is a serialization of the AST custom crafted for that language. It is as much an "AST" as any other serialization would be; all such alternative representations would all produce isomorphic memory representations if parsed from a proper library. At a high level it may sound useful to try to then provide a cross-language AST representation, but it's one of those things that sounds great at a high level but as soon as you actually tried to implement it for, say, Python and C++, you'd rapidly discover that in practice there's not as much opportunity for "generic AST operations" as you may think. The problem isn't that it isn't "stored as an AST" but that $YOUR_LANGUAGE apparently doesn't have good libraries or mechanisms for getting at it. Go, for instance, ships with the relevant bits of the compiler exposed, and as a result there are tons of tools that operate on Go code as ASTs and not textually, because it's readily available and supported by the core language team. I use this only as an example I know personally, there are other languages that have similar sorts of support as well.
- vlovich123 5y agoI feel like you're picking a strawman here. The AST serialization everyone is implying is one where you don't need to token/lex but can just load it directly & manipulate it (i.e. implying the on-disk version is a valid AST or one who's validity can be trivially validated without needing to have the entire language syntax & grammar). First, that makes the compiler much faster because tokenization/lexing is moved to the "save" phase which happens infrequently at human scale vs the compile/processing phase which happens in an automated fashion where the overhead can be notable. Additionally, if you mmap the AST from disk into memory, you can use finer-grained caching to memoize expensive analysis that happens for faster compiles of code that's only changed slightly (e.g. changing whitespace/comments wouldn't recompile anything). More importantly for advocates, it avoids needing to ship the deserialization library and makes tooling simpler. That's really why the idea of a simple AST format is so attractive. Typically compiler frontends are typically very tightly coupled to the underlying middle & back end. There's some work in some languages to decouple this (e.g. LSPs & Idea's failable parsing approach), but the efforts are still very immature & it's still not clear to me that it's worth it (see the last paragraph). The main underlying challenge with making sure the on-disk contents is well-formed according to the syntax rules is that frequently you want to pause work at an intermediate stage. This means you either have to make sure that whatever state the user saves is a valid AST via editor tricks (although I think this also typically means you have to design the language around it), you reject saves, every tooling library has to be capable of parsing malformed ASTs, or you save a dirty transformation to apply to the last known saved version so you can have the user resume editing but otherwise tooling uses the "last known good" version. That's the real challenge with having a serialized version that's amenable to 3p tooling for interop. Finally, all the "serialize the AST" solutions ignore the problem of wanting to grep the codebase. This means you need to change out several decades of line-oriented manipulation tools in favor of new ones that are AST-based & likely more complicated to write/maintain as compared with one-line regular expressions. At least I've yet to see any AST manipulation libraries that aren't drastically different from existing text manipulation tools if clang-tidy and Rust macros are any indication about what good solutions to the problem look like today. I think eventually we'll get AST serialization, but I think it will be packaged into an entirely new language (like Rust did with ownership) that also considers the tooling aspect end-to-end rather than as a retrofit into existing languages. Once that's successful, then I think we'll see retrofits because the space will have been better explored & other languages will benefit from the R&D into what a successful path would look like.
- BiteCode_dev 5y agoYes, but only if it falls back to text diff as soon as there is the smallest doubt it can't provide a good AST diff.
- thrwyoilarticle 5y agoYou can also write git hooks to turn their spaces into your tabs & vice versa.
- OJFord 5y agoThat's not a good solution - every commit with an author (well technically committer I suppose) whose opinion differs to the last will be horrendous.
- thrwyoilarticle 5y agoThere won't be any difference. OP will run the script when they checkout, work with their tabs, then run the script when they commit. Spaces in, spaces out.
- OJFord 5y agoOh ok, sure. But then that's just a weak version of what's being requested - an entirely neutral more agnostic, abstract format that stores the meaning without any formatting at all.
- fstrthnscnd 5y agoTabs do work as long as they aren't fixed width (I don't know what you mean by "custom"). For instance, in many languages, one will sometimes have to split a function call to many lines, and in most languages function names aren't of fixed length, thus in order to get a correct alignment for parameters, the tab width at that point will have to match the function name length. #include<stdio.h> int main(int argc, char* argv[]) { printf("%s %s %s %s\n", __FILE__, __LINE__, __DATE__, __TIME__); return 0; } I agree with your idea of storing a normalized version of the code in the repo: it wouldn't then matter whether that version contains characters to align the code properly, it would just be inserted by the editor/linter as needed. The difficulty is that sometimes linting isn't enough, and some manual formatting is needed. Or perhaps the formatting rules are under specified? Another issue with AST diffing is when languages allow some form of syntactic sugar as preprocessing: the compiler might just see the simplified tree, not the one with the "sugary" forms. A tool capable of parsing such languages should also be able to handle these extensions.
- Asraelite 5y ago> the tab width at that point will have to match the function name length. This is a non-issue. Use tabs for indentation and spaces for alignment.
- njharman 5y agoThat is the kind of problem solution that ends up with you now having 2 problems. New problem(s); having tabs and spaces, having to think when to use them, having to train/document everyone in usage, having to debate that usage, having to correct code and chastise people who get usage wrong. Use a automatic code formatter with minimal options. Automate either running code formatter on commit or denying commits that change when code formatter is run on them.
- Asraelite 5y agoAbsolutely, I wouldn't dream of doing any kind of fancy alignment by hand, only with an auto-formatter. If I had to break arguments onto multiple lines without an auto-formatter I would just keep it simple and use another level of indentation instead of aligning them with the function name.
- Anon_troll 5y agoThe whitespace and formatting are not significant to the compiler, but they can provide a lot of information to the reader of the code. You can often see where the writer put the most effort and thought by just seeing how they wrote it. This can help analyzing a codebase considerably. If everything is normalized, you lose those valuable cues.
- geofft 5y agoOne of the practical issues here is, if your code fails to compile in CI with an error like /home/ci/src/foo.c:123:45: error: use of undeclared identifier 'a' or /home/ci/src/bar.py:50: syntax error in type comment or crashes in production with an error like java.lang.NullPointerException at com.example.Baz.doThings(Baz.java:1337) you really want to be able to find line 123 column 45, line 50, or line 1337 in your editor, and have that be the same line as what your CI compiled and deployed. On its own, tabs vs. spaces only affects columns, and you can probably figure things out without columns (although it's a shame to lose it). But different tab sizes affect how long your lines are, and line wrapping is a thing that people care about at least as much as tabs vs. spaces (people with different size monitor or fonts will easily see too-long or too-short lines on their display; if your spaces are equivalent to the tab stop, the distinction is literally invisible). And once you start rewrapping lines, everyone's line numbers are different. I think it's possible to solve this by using some sort of AST-based index into the file and teaching IDEs to let you seek based on that, but it's suddenly a more complex problem.
- ratww 5y agoThis is already a very common problem with a solution: transpiled JS already needs source maps to display errors correctly.
- geofft 5y agoNo, I don't think that's the same problem / the same solution. A source map translates between a layout checked into the code and a format generated at build time. I'm talking about translating between a layout in a developer's local workspace and the layout checked into the code. Since the developer can choose whatever formatting options they want, there isn't a single source map that can be referenced in the compiled version of the code, so backtraces etc. So the transformation cannot be done at the point the error is displayed (compiler output or backtrace output), it has to be done in the context of the developer's local workspace. I think source maps could probably be inspiration for solving this problem, but I don't think they would work directly - and even if they did, the real problem here is not designing a solution, it's getting everyone's IDEs to work properly with it. Source maps work largely because the major browsers know how to deal with source maps in JS. You'd have to extend this to all the other ecosystems, at the very least.
- thefreeman 5y agoIf you append `?w=1` to the diff view URL on a pull request it makes it whitespace agnostic just FYI
- pbiggar 5y agofwiw, this is what we do in Dark [1]. We store (serialized) ASTs, then then we pretty print them in the editor. This converts the AST into tokens that you see on your screen, complete with configurable* indentation, line-length, etc. Code would be displayed according to your config* and the same code displayed differently to a different developer looking at the same code. [1] https://darklang.com https://darklang.com * I haven't actually enabled users to configure this, but it's just some variables called 'indent' and `lineLength` in the code
- iso8859-1 5y agoThis is on the http://lamdu.org http://lamdu.org roadmap.
- ghoward 5y agoI'm actually working on a VCS based on this idea and on tracking changes to binary files based on their structure as well. (It turns out that the same techniques work for both.) AMA and please give me feedback!
- tombert 5y agoI would definitely support a Lisp-centric Git. Whenever I do Clojure, something that can get difficult when working with multiple people is how the parentheses/brackets/braces stack up, especially when everyone seems to have different opinions on how that works. As a result, if you're not careful, when there's a merge conflict you can have a ton of extra parentheses, which can be irritating to debug. Obviously this is at some level an issue inherent to Lisps (and to be clear, I love Lisps, and these small headaches are worth it), but I think problems like that could be reduced if our source controls were aware of the ASTs.
- timgilbert 5y agoYeah, I've long thought a diff tool that works on s-exprs instead of lines would be invaluable for Lisp programming. It doesn't seem like it would be too hard to write, either, although getting GitHub etc to use it seems like it would be its own challenge...
- fulafel 5y agoGit can use an external tool for merging, so there could be eg a Clojure merge plugin even now. Apparently there are some commercial merge tools that Java programmers use for this like SemanticMerge. After all you hit the similar curly braces merge problems with other languages.
- aardvark179 5y agoI’ve done quite a lot of work on version management on structured data (in my case this was for a version managed GIS database) and it’s not an easy problem, and is likely even harder with something like an AST that is generated from a text file and so does not preserve the identity of nodes. I’m not saying that it’s impossible, but it is more work and requires more tooling around it than people think, and it keeps coming up here and other places as a, “really good idea.”
- cormacrelf 5y agoCounterpoint: a quick google reveals diffsitter: https://github.com/afnanenayet/diffsitter https://github.com/afnanenayet/diffsitter The output could be a lot more compact, it could do better at adding context (in the same way https://github.com/romgrk/nvim-treesitter-context https://github.com/romgrk/nvim-treesitter-context does, etc), but if you're interested in this it's really within reach, go help out. I wonder if you can use it for automerge yet.
- deleted 5y ago[deleted]
- nerdponx 5y agoStoring AST instead of source code is one of the goals of the very interesting Unison programming language: https://www.unisonweb.org/ https://www.unisonweb.org/ Part of what's nice about Git (and plain text in general) is that it's the lowest common denominator for a lot of things. This is why traditional Unix tools are built oriented around streams of bytes. Text is a low level carrier protocol; you can encode almost anything in it, but you need to agree on some kind of format. The good part is that you can use very very generic tools on almost arbitrary pieces of data. The bad part is that you might have to do a lot of parsing and re-parsing of the same data, and you have to contend with the dangers of underspecified formats. Git follows the Unix tradition in this regard. As a result, it is nearly universal in what it can store. You can use it to store pretty much anything, but you are now at the lowest common denominator of support for any particular data format. Git-for-ASTs will no longer have this universality property, but will gain a lot more power in the covered domain. This is a design tradeoff. One thing that's nice about Git is that you can specify arbitrary diff drivers with the "attributes" system. So even if the Git database is storing plain text, your diff driver can parse your source code into ASTs and present AST diffs to you when you run `git diff`. Perhaps more impressive, you can configure custom merge drivers, so you can (theoretically) implement semantic merging of ASTs right inside Git. There are probably some fundamental limitations of this system, because the underlying data is still stored as blobs of bytes. But you can get pretty far as long as you don't mind parsing and re-parsing the same text over and over.
- ssivark 5y agoHas this approach been tried? (Unison or otherwise…)
- nerdponx 5y agoI believe Unison is the only attempt to do this at a programming language/environment level. For Git diffs, there is Diffsitter, which uses Tree Sitter to generate semantic diffs of code files: https://github.com/afnanenayet/diffsitter https://github.com/afnanenayet/diffsitter I have not used it, but it is high on my todo list.
- Karellen 5y ago`git` generally doesn't work with lines of text. Mostly it works with opaque file blobs and directory trees. `git diff` and `git merge` work with lines of text by default - but they don't have to. You can supply your own `diff` and `merge` tools with the `difftool.*` and `mergetool.*` config options, try them out with `git-difftool` and `git-mergetool` commands, and set the default with the `git.diff` and `git.merge` config options. If someone wanted to create AST-based diff and merge tools for a given language, they could be plugged right into the existing `git` infrastructure and it would work with them absolutely fine.
- bspammer 5y agoThis feature is useful in so many different places. I use it to diff small encrypted files in my repo - just add `gpg -d` as a diff configuration and now I can use git log, diff etc in a meaningful way with binary files. I've heard of people using it with pdfs as well - a pdf to html converter lets you get a good idea of what changed in the document.
- colonwqbang 5y agoYes, I think this article is coming at it from the wrong end. Git is hardly the problem here, nor is it going to provide the solution. The problem seems to be that we are lacking the format and the toolchain to manipulate it, and that is not the fault of git. What is the state of the art in this area? Does somebody know of a viable format and toolchain, or any interesting projects looking to build them?
- tyleo 5y agoI believe that semantic merge does something like this: https://www.semanticmerge.com/ https://www.semanticmerge.com/
- paul_h 5y agoCame to link to that :)
- kapep 5y ago> If someone wanted to create AST-based diff and merge tools for a given language, they could be plugged right into the existing `git` infrastructure and it would work with them absolutely fine. There's a lot tooling in the Eclipse modelling ecosystem which could be easily used for this. Storing XML-based models in git is no problem and there's tooling for diffing and merging models via a GUI or programmatically. Combined with the fact that xtext DSLs use EMF models to represent ASTs, it wouldn't be too hard to glue together an AST-based a diff/merge tool for an xtext DSL.
- Smaug123 5y agoI'm surprised they didn't mention Unison (https://www.unisonweb.org/ https://www.unisonweb.org/), whose big idea is an immutable content-addressable store of ASTs. I really hope it changes everything.
- renox 5y agoExcept that Unison created its own language which makes pretty sure that they are doomed to fail.. I don't know if there is a technical reason for the new language or if it's NIH syndrome.
- mangecoeur 5y agoInteresting they mentioned Jupyter Notebooks but not NBDime https://github.com/jupyter/nbdime https://github.com/jupyter/nbdime which is a Jupyter plugin specifically to address this problem. Without it, diffing notebooks is not feasible.
- auscompgeek 5y agoNote that you can specify a custom merge driver for different file types using a combination of gitattributes and git-config: https://git-scm.com/book/en/v2/Customizing-Git-Git-Attributes#_merge_strategies https://git-scm.com/book/en/v2/Customizing-Git-Git-Attribute...
- atonalfreerider 5y agoSelf-promote: Primitive does AST diffing and represents the changes graphically primitive.io
- jrm4 5y agoWhat if Programming Languages worked with Lines of Text?
- vxNsr 5y agoIs it just me or is he describing an IDE with source control?
- cies 5y ago> Structure editors haven't really taken off yet despite several historical and contemporary attempts. This is a nice contemporary one: https://github.com/projectional-haskell/structured-haskell-mode https://github.com/projectional-haskell/structured-haskell-m... Lisps also have all kinds of options available in Emacs, but it is more special to see this outside of the land of s-expressions.
- raxxorrax 5y agoTheoretically it might work, but I don't think I am too fond of the idea. I used git pull to completely waste my source and it would have been nice for git to have more intelligence here, but in the end I think some of its success lies in its simplicity. SVN isn't too bad and not too much of a difference to git if you use a central repository anyway. The main neat thing was to just have one hidden folder, not in every subdirectory. Git would also need the ability to transform from AST to source for every language. A bit unrealistic and there is no benefit to it. Could also do that with Assemblies and some meta info for the decompiler.
- ufo 5y agoI'm trying to remember the citation, but I remember seeing a presentation once from someone who studied this and they said that the thing that worked best was a hybrid approach: use structured diff at the top level of the program (modules / methods) but use line-based for statements and expressions. According to them, the structured diff can give unintuitive results if applied at the lowest syntactic levels.
- jakeinspace 5y agoI think the only useful way to implement AST-level diff/merge for non-trivial codebases would require the compiler to provide the parsed AST, since per-file ASTs would lack a lot of context. You could also ask the user to provide a separate file or files that describe the code topology, but why bother when the compiler can spit out an AST itself? A diff tool which targets a few of the bigger build systems (CMake, Maven, Gradle) and compilers might work, and could worry about small build environments after gaining momentum.
- kitplummer 5y agoI'm just a bit more "generally" curious. Is `git` being the _only_ DVCS a good thing? Not to say that `hg` or `darcs` don't exist, just that the hub on top of git has pushed us in a singular direction. I would like to see, at least academically, something more.
- tomphoolery 5y agoThe choice of DVCS tooling is ancillary to the success of GitHub. People learned Git so they could use GitHub, not the other way around. At least, this is the way I remember it. If someone comes along and builds the best forge software ever, but uses Mercurial instead of Git, I'll bet a lot of people would switch technologies at some point. Until then, I'd say most people use GitHub because GitHub works for them, and they use Git because that's how you interact with GitHub. They don't care about the ivory tower benefits of their particular DVCS tooling, all they care about is easily collaborating with their teammates. It would definitely be great if you could have a GitHub-like experience using Mercurial or Darcs, but so far I haven't seen anything close to that.
- aayjaychan 5y agoDoes GitLab count as a GitHub-like experience? There is a fork of GitLab called Heptapod that supports Mercurial. https://heptapod.net/ https://heptapod.net/
- gumby 5y agoShared (concurrent) code editors might work better if their CRDT/OT model worked at that level. Not that I really want to edit code in a shared environment (editing documents that way is bad enough), but just musing…
- mabbo 5y agoReading this article, I feel as though the author doesn't deeply understand git. git works on blobs of data, not files, and not lines of text. It doesn't just happen to also work on binary files- that's all it works on. Now, if the author is suggesting that git-diff ought to have a language specific mode that parses changed files as ASTs to compare, now I'm interested. Let's do that. I'll help! But git does not need to change how it works for that to happen. Git does not even need git-diff to exist to serve it's main purpose.
- dboreham 5y agoPretty sure OP does understand, and is proposing what you deduced. Incorporate some semantic understanding of the version controlled data into the VCS. Currently this work is subcontracted to humans.
- mabbo 5y agoMaybe I'm misunderstanding. It's just lines like this: > The text-orientated design of git reflects... > The current version of git is also able to find differences in binary files. > if we were storing information as ASTs, rather than lines of text These all, to me, show a gap in the authors understanding of how git works. And that's okay- git is often easier to use than is to understand. But if they had a better understanding, they could make their point far better. And without understanding, they won't be able to implement this idea.
- mbauman 5y agoYou can already choose different `diff` programs to use for particular filetypes. E.g., nbdime for Jupiter notebooks: https://nbdime.readthedocs.io/en/latest/vcs.html#git-integration https://nbdime.readthedocs.io/en/latest/vcs.html#git-integra...
- hardwaregeek 5y agoI dunno I feel like you're focusing on a detail that's not particularly relevant. The author's main thrust is precisely what you described about parsing changed files as ASTs.
- ClassAndBurn 5y agoGit is designed to require human oversight. This is usually a feature, but in recent years has become a bug with things like GitOps. It's important to remember that Git is a terrible database because of its lack of semantic structure. All conflicts require a human who does have to context. This is why almost no one builds a system that uses Git as a two way interface. And when they do, its via Github Pull Requests (which go to humans) and not Git itself. In all, this makes it a wonderful general purpose shared filesystem. And that's about it.
- Jensson 5y agoI don't see how this could ever work on evolving languages, different GIT versions would produce different commits and read commits differently based on the latest C++ standard. This would potentially lead to version control bugs where different GIT versions creates different results from the same commit, that is horrible, version control needs to be 100% bug free in that regard. The only reasonable application would be to use a language AST parser to better identify relevant text diffs, but the commits still needs to be stored as text.
- dboreham 5y agoThis doesn't really make sense, because in order to have those code changes compile correctly, there must be a corresponding commit to CI config that changes the complier version or compiler switches for the new language version. The "semantic-diff-er" can also be driven by that commit such that it uses the correct language version.
- verdverm 5y agoIt's non trivial to support multiple versions of the same language on a host system. You have to account for dev machines and workflows as well. Docker can help with this, but often devs don't want to run a container to build their code. It's a hard habit to change. Now, consider how difficult it would be to get the differ to understand where to find compiler versions.
- shepherdjerred 5y agoCommits could be stored as is, the difference would be that diffs are clearer when presented to a human.
- pkghost 5y agoHow is this different from any other problem that is already solved by version pegging?
- jcrites 5y agoI’ve had loosely similar ideas before. The basic idea is to make the compiler tool chain aware of diffs, and help scrutinize and implement them. Refactoring suggestions could be included with the diff. For example, say you’re dependending on a module and it renamed a class/method/trait/macro/constant/whatever. A synchronous method has become a sync or vice versa. The diff could include programmatic instructions for consumers to apply to their code bases switching them over to the new method. This could be as simple as semantically changing the name used, or in the case of changing sync to a sync it could add `await` in the appropriate spot. There’s no limit to how complex the rewrite rules could be. You could totally reorganize the parameters to a function and ship that refactoring, or even add a parameter along with the default code necessary to provide it. Unlike the author I don’t think code in Git will likely ever move beyond source, plus the refactoring instructions needed to update a change from a dependency — perhaps a macro-like syntax. Too much text manipulation is required from source control for me to conceive of it being anything but human programming text in a future I can imagine. Machines can already parse it; there doesn’t seem to be a compelling reason to store some other kind of structure. Refactoring wouldn’t need to be any special Git extension, just a file accompanying the commit with instructions for the language tool chain. Your IDE or CLI could walk you through interactively everywhere it’s getting applied, or you could apply all and review the result in your app that consumes the module. This would also open up security risks from accepting diffs from dependencies and applying their refractors, but unless modules are sandboxed quite well that’s a risk you take with updating dependencies anyway. And you can always scrutinize the refactoring-diff manually after it runs before accepting it. The industry would probably standardize onto the notion that a change that requires running automated room factoring from a dependency across your codebase is a major version change; in other words a breaking or backwards incompatible change, just one that’s much easier to upgrade to. Languages with macros or other programmatic transformers would be well suited to this concept I think. Maybe Rust macros could be enhanced for the purpose to pattern match over an existing codebase somehow: not just the macro invocation point, but anywhere, e.g., a given trait or function is used; and then the output of the macro would not feed into the next stage of the compiler step but would instead result in rewriting code on disk to produce a diff that you examine and apply to your code. A capability like this would make it much easier to manage large aggregate code bases consisting of many dependencies. OSS package maintainers or infrastructure providers at companies could ship nominally backwards -incompatible changes that are still actually compatible when you run the macro transformer that updates the code that uses them. For a simple example, imagine that `foo()` was previously a function and the implementation chooses to add some optional parameters or those with default values and change it into a macro `foo!()`. The accompanying transformer would semantically identify references to `foo()` and make the necessary updates. You could rename global constants or traits or other code elements this way. Consider the way in which Google Guava has had to evolve over time. A number of its features have become part of the Java language, and thus the classes deprecated and removed gradually. With a compiler facility like what I am describing, users of Guava could run the transformer to migrate older could bases that use methods like Guava’s `Preconditions.checkNotNull(Object, message))` to use Java’s now-standard `Objects. requireNonNull (T obj, String message)`. Because the maintainers of Guava wanted to keep it modern, current, avoid redundancy, and designed in the best way that they knew how, they made a number of breaking changes for which the project lead later apologized [1]. Most changes wouldn’t have been be painful if accompanied by automated refactoring. You could allow the transformer to produce code that still needs work from humans to finalize and compile. In that case it could change as much as it’s able and leave instructions at each call site. At my company we have bots that submit proposed code changes to our codebase that need to be taken over as author by a human, reviewed, sometimes lightly edited, and shipped, and they work quite well. One bit finds unused code and submits diffs to remove it if it’s been in the code repository long enough. Another detects when launch experiments have been at 100% all on one treatment for a long period of time (meaning the feature has launched) and submits code removing the experiment check. The latter sometimes require removes surrounding code that would subsequently become unused from removing the experiment check. These have provided meaningful value in helping keep the codebase tidy and I look forward to more automation like this in the future, including diff-aware compilers and refactoring tools. [1] https://www.reddit.com/r/java/comments/mr03mi/comment/guk8482/ https://www.reddit.com/r/java/comments/mr03mi/comment/guk848...
- CodeIsTheEnd 5y agoI don't understand why GitHub hasn't solved the issue of diffs starting with a '}' (or ')' or 'end'). Just slide the diff over while it starts with a closing token! I suppose it's an artifact of the diffing algorithm, but aren't there better diffing algorithms, even built-in within git? This is by far the most obvious example of "git doesn't understand programming languages", but it also seems like the most straightforward to fix.
- nemetroid 5y agoGit supports a few different diff algorithms. GitHub only seems to support the standard Myers algorithm, though: https://github.com/isaacs/github/issues/455 https://github.com/isaacs/github/issues/455
- mynegation 5y agoIt is because diff is syntax agnostic. You might be able to get away with this hack in some cases but that complicates algorithm and will break in some other cases (how about nested brackets? Multiple brackets on one line?). Once you want to handle this properly you need syntax aware diff algorithm and some resources are linked in this discussion.
- hardwaregeek 5y agoI've wanted this for a while, but I will say there's some caveats. Sometimes I want to commit just as a "it's the end of the day, I want to leave, here's a code dump". I suppose you could have multiple tiers of code saving. I've also wondered about whether you could do code analysis with time as a dimension. If you can analyze the evolution of the code and pull old implementations, what can you do? Autocomplete is a good example, as it can pull previous patterns you've used. Maybe some way to tell the programmer "hold up, you've made this mistake before, don't do it again"? I'm not sure.
- inetknght 5y ago> I want to commit just as a "it's the end of the day, I want to leave, here's a code dump" 1. git checkout -b eod-$(date -Id) 2. commit 3. leave 4. return 5. git checkout - 6. git merge - --no-commit
- hardwaregeek 5y agoYes, but if we're talking about some hypothetical tool that requires a valid AST, there might be a situation where I don't have a valid AST and want to save the code. Similarly, I had a job where we used pre-commit hooks that ran a linting script. I had to override the hook to commit which was slightly annoying at times.
- lsaferite 5y agoPerhaps you store the invalid syntax portion as an in-tree comment until it's valid syntax.
- kazinator 5y agoThe blogger does not understand Git, fundamentally. Git does does not work with text. It stores snapshots of artifacts. The diffs that you see when you use the various commands like git log -p are recovered from the snapshots, when those artifacts happen to be text files. Git absolutely works with texts when you connect it with external representations and tooling, such as when you "git format-patch" and then "git am" to import that; and the rebasing workflows obviously have textual merging with conflict resolution. Still, that seems like something that could be externalized. A language-specific three-way-diff tool can handle a merge by parsing all three pieces and working with ASTs. It's something that could be developed later, yet still work with your old commits. There is this: https://git-scm.com/docs/git-mergetool https://git-scm.com/docs/git-mergetool No idea how well it works.
- skybrian 5y agoIf you’re interested in this sort of thing you might want to look at Dolt (for sharing databases in a git-like way) and Pijul, which records diffs explicitly, rather than calculating them on the fly. I wonder if there might be a clever way to encode source code in a Dolt database? Maybe each function should be a record?
- LukeEF 5y agoAuthor is CTO of TerminusDB (https://terminusdb.com/ https://terminusdb.com/), which is a more graphy version of Dolt! Check it out.
- jpitz 5y agoDidn't the VisualAge IDEs do this with their built-in version control? This was 20 years ago, and I seem to remember that the version control was at the method level, not file level.
- mumblemumble 5y agoI would maybe be interested in Git allowing you to plug in your own diff generators for different file types. But I would not want Git itself trying to understand the contents of files. That seems to me to be an idea that lives on a misconception of the "things programmers believe about names" variety. Not every file in source control is source code. Not every programming language's grammar maps to an abstract syntax tree. In some files, such as makefiles, the difference between tabs and spaces is semantically significant. Some languages (such as Fortran and Racket) have variable syntax. And so on and so forth. So I think that we really don't want the source control system itself trying to get too smart about the contents of files. That will inevitably make the source control system less compatible with the various kinds of things you might want to put into source control. And it will also make the source control system a lot more complicated than it would otherwise be, in return for a largely theoretical payoff. But if we want to delegate the work of generating diffs off to other people, so that Git can allow for syntax or semantics-aware diffing without having to personally wade into that quagmire (and perhaps also allowing language communities to support multiple source control systems, a bit like how it works with LSP), that might be an interesting thing to experiment with.
- saurik 5y ago> I would maybe be interested in Git allowing you to plug in your own diff generators for different file types. This is already supported.
- lux 5y agoA common example is UnityYAMLMerge for merging the Unity game engine's generated files. https://docs.unity3d.com/Manual/SmartMerge.html https://docs.unity3d.com/Manual/SmartMerge.html Configuring it to work with Git and others is a little ways down the page, but would apply the same for other diff tools.
- franga2000 5y agoI looked this up and for anyone wondering, it's called "diff/merge drivers", but there are only a handful of them out there. Some highlights from a few minutes of searching: - MS Office: https://github.com/lcnittl/DMFO https://github.com/lcnittl/DMFO - SQLite: https://github.com/cannadayr/git-sqlite https://github.com/cannadayr/git-sqlite - Jupyter notebooks: https://nbdime.readthedocs.io/en/latest/vcs.html#git-integration https://nbdime.readthedocs.io/en/latest/vcs.html#git-integra... One big caveat of this is that since git doesn't really store just a stack of diffs, despite the fact it presents itself as such to the user, a custom merge driver will not make your .git grow any less than it would normally.
- olodus 5y agoEver since I learned about Git merge strategies and wrote a very basic one myself, I've been wanting to write one that syntaxticly understands a bit of the test framework code we use at work. It is super annoying when you copy a test because you want to vary a very specific case and gig gets all confused about what code is and isn't the same. (yeah I know I should break out the copied part but who always has time for that)
- deleted 5y ago[deleted]
- tomxor 5y ago> The fact that git works on lines of text [...] we could be looking at the alterations to the abstract syntax tree. Fundamentally git does not operate on text, it operates on files (content addressed SCM not a ledger of text diffs); diffs are generated upon request between arbitrary merkel trees. So there is no need to implicate git in such a tool, it can be independent: GIT_EXTERNAL_DIFF When the environment variable GIT_EXTERNAL_DIFF is set, the program named by it is called to generate diffs, and Git does not use its builtin diff machinery. For a path that is added, removed, or modified, GIT_EXTERNAL_DIFF is called with 7 parameters: path old-file old-hex old-mode new-file new-hex new-mode
- lamontcg 5y agoThis has been posted before
- bialpio 5y agoThis made me think of Unison: https://www.unisonweb.org/ https://www.unisonweb.org/ Discussion: https://news.ycombinator.com/item?id=27652677 https://news.ycombinator.com/item?id=27652677
- maweki 5y agoWorking on the AST is quite an interesting idea, until your comments aren't in the AST and you want to commit a syntax error of work in progress. Not to mention changing ASTs (while maintaining concrete syntax) in different versions of the language.
- deleted 5y ago[deleted]
- ozim 5y agoI feel this is just an example of "worse is better" and whole proposition as interesting but totally not practical and I would not like for GIT to go anywhere near that idea.
- alkonaut 5y agoI’d give anything just to get a few basic merge modes. For example “this file can treat two one line additions as unordered”. So any shared append-only file (a change log, an enumeration,…) doesn’t automatically conflict. Syntax aware diffing would be great too, but I’d take something much simpler. For syntax aware stuff I’d love something that could tell semantic changes from noise.
- aidenn0 5y agoThe problem with a tool that depends on structured data is that it only works with structured data. Of course the problem with a tool built for unstructured data is that it's dumber than it need be, and when you do treat the data as structured, it's ad-hoc and often buggy. When talking about how "powerful" a tool is, there's always this tension between structured and unstructured.
- g051051 5y agoBack in the early 2000's Visual Age for Java allowed you to version individual methods. Since Visual Age for Java was derived from Visual Age for Smalltalk (and was actually written in Smalltalk) I suppose it inherited the capability from there.
- igouy 5y agoMore specifically from ENVY/Developer https://www.google.com/books/edition/Mastering_ENVY_Developer/ld6E19QIMo4C https://www.google.com/books/edition/Mastering_ENVY_Develope... which was available for the three main commercial Smalltalk platforms: VisualWorks, IBM Smalltalk, Digitalk.
- g051051 5y agoI never knew that! Thanks.
- elif 5y agoIt sounds inspirational and revolutionary but the longer I think about it, the less utopian it feels. The idea of forgetting about pieces of code or how things are linked. Files also provide an opportunity to purposefully communicate organizational intent.
- shoo 5y agoThere's a good blog post about auto-merging JSON/XML structured data files (for game content) on the bitsquid blog from 2010: > having content conflicts is no fun either. A level designer wants to work in the level editor, not manage strange content conflicts in barely understandable XML-files. The level designer should never have to mess with WinMerging the engine's file formats. > And conflicts shouldn't be necessary. Most content conflicts are not actual conflicts. It is not that often that two people have moved the exact same object or changed the exact same settings parameter. Rather, the conflicts occur because a line-based merge tool tries to merge hierarchical data (XML or JSON) and messes up the structure. > In those rare cases when there is an actual conflict, the content people don't want to resolve it in WinMerge. If two level designers have moved the same object, we don't really help them address the issue by bringing up a dialog box with a ton of XML mumbo-jumbo. Instead, it is much better to just pick one of the two locations and go ahead with merging the file. Then, the level designers can fix any problems that might have occurred in the level editor -- the right tool for the job. -- http://bitsquid.blogspot.com/2010/06/avoiding-content-locks-and-conflicts-3.html http://bitsquid.blogspot.com/2010/06/avoiding-content-locks-...
- WalterBright 5y agoI'm fine with the line oriented nature of git. I wouldn't want to be confronted with problems due to the constantly shifting nature of languages vs whatever plugin exists for them on git. Even C is constantly changing.
- zomglings 5y agoI maintain a free/open source project that does exactly what the author asks for: https://github.com/bugout-dev/locust https://github.com/bugout-dev/locust. Our tool uses git as the foundation of its functionality. It superimposes git diffs on top of ASTs. It is insanely powerful. For example, we use it to power semantic code search and current support Python, Javascript, and Java. We generate a JSON object describing the AST differences between initial and terminal commits on GitHub PRs. A full text search on the JSON objects performs surprisingly well when we want to answer questions like, "When did we add dateutils as a dependency?" or "When did we last change the /journals handler on the API?" The Python integration currently sees the most use but if you are interested in other languages, we would be happy to support it. Do drop me a DM if you want help getting started with Locust.