6 ms·
I think agree (but I think I think about this maybe a one level higher). I wrote about this a while ago in https://yoyo-code.com/programming-breakthroughs-we-ne
by panstromek 8mo ago
I think agree (but I think I think about this maybe a one level higher). I wrote about this a while ago in https://yoyo-code.com/programming-breakthroughs-we-need/#editing-code-in-general-doesn-t-work-well https://yoyo-code.com/programming-breakthroughs-we-need/#edi... .
One interesting thing I got in replies is Unison language (content adressed functions, function is defined by AST). Also, I recommend checking Dion language demo (experimental project which stores program as AST).
In general I think there's a missing piece between text and storage. Structural editing is likely a dead end, writing text seems superior, but storage format as text is just fundamentally problematic.
I think we need a good bridge that allows editing via text, but storage like structured database (I'd go as far as say relational database, maybe). This would unlock a lot of IDE-like features for simple programmatic usage, or manipulating langauge semantics in some interesting ways, but challenge is of course how to keep the mapping between textual input in shape.
- conartist6 8mo agoCome the BABLR side. We have cookies! In all seriousness this is being done. By me. I would say structural editing is not a dead end, because as you mention projects like Unison and Smalltalk show us that storing structures is compatible with having syntax. The real problem is that we need a common way of storing parse tree structures so that we can build a semantic editor that works on the syntax of many programming languages
- panstromek 8mo agoI think neither Unison nor Smalltalk use structural editing, though. [edit] on the level of a code in a function at least.
- conartist6 8mo agoNo, I know that. But we do have an example of something that does: the web browser.
- fuhsnn 8mo agoStructural diff tools like difftastic[1] is a good middle ground and still underexplored IMO. [1] https://github.com/Wilfred/difftastic https://github.com/Wilfred/difftastic
- panstromek 8mo agoIntelliJ diffs are also really good, they are somewhat semi-structural I'd say. Not going as far as difftastic it seems (but I haven't use that one).
- zelphirkalt 8mo agoWhy would structural editing be a dead end? It has nothing to do with storage format. At least the meaning of the term I am familiar with, is about how you navigate and manipulate semantic units of code, instead of manipulating characters of the code, for example pressing some shortcut keys to invert nesting of AST nodes, or wrap an expression inside another, or change the order of expressions, all at the pressing of a button or key combo. I think you might be referring to something else or a different definition of the term.
- panstromek 8mo agoI'm referring to UI interfaces that allow you to do structural editing only and usually only store the structural shape of the program (e.g. no whitespace or indentation). I think at this point nobody uses them for programming, it's pretty frustrating to use because it doesn't allow you to do edits that break the semantic text structure too much. I guess the most used one is styles editor in chrome dev tools and that one is only really useful for small tweaks, even just adding new properties is already pretty frustrating experience. [edit] otherwise I agree that structural editing a-la IDE shortcuts is useful, I use that a lot.
- pegasus 8mo agoSome very bright Jetbrains folks were able to solve most of those issues. Check out their MPS IDE [1], its structured/projectional editing experience is in a class of its own. [1] https://www.youtube.com/watch?v=uvCc0DFxG1s https://www.youtube.com/watch?v=uvCc0DFxG1s
- flowerbreeze 8mo agoI'm quite sure I've read your article before and I've thought about this one a lot. Not so much from GIT perspective, but about textual representation still being the "golden source" for what the program is when interpreted or compiled. Of course text is so universal and allows for so many ways of editing that it's hard to give up. On the other hand, while text is great for input, it comes with overhead and core issues for (most are already in the article, but I'm writing them down anyway): 1. Substitutions such as renaming a symbol where ensuring the correctness of the operation pretty much requires having parsed the text to a graph representation first, or letting go of the guarantee of correctness in the first place and performing plain text search/replace. 2. Alternative representations requiring full and correct re-parsing such as: - overview of flow across functions - viewing graph based data structures, of which there tend to be many in a larger application - imports graph and so on... 3. Querying structurally equivalent patterns when they have multiple equivalent textual representations and search in general being somewhat limited. 4. Merging changes and diffs have fewer guarantees than compared to when merging graphs or trees. 5. Correctness checks, such as cyclic imports, ensuring the validity of the program itself are all build-time unless the IDE has effectively a duplicate program graph being continuously parsed from the changes that is not equivalent to the eventual execution model. 6. Execution and build speed is also a permanent overhead as applications grow when using text as the source. Yes, parsing methods are quite fast these days and the hardware is far better, but having a correct program graph is always faster than parsing, creating & verifying a new one. I think input as text is a must-have to start with no matter what, but what if the parsing step was performed immediately on stop symbols rather than later and merged with the program graph immediately rather than during a separate build step? Or what if it was like "staging" step? Eg, write a separate function that gets parsed into program model immediately, then try executing it and then merge to main program graph later that can perform all necessary checks to ensure the main program graph remains valid? I think it'd be more difficult to learn, but I think having these operations and a program graph as a database, would give so much when it comes to editing, verifying and maintaining more complex programs.
- panstromek 8mo ago> what if the parsing step was performed immediately on stop symbols rather than later and merged with the program graph immediately rather than during a separate build step? I think this is the way to go, kinda like on Github, where you write markdown in the comments, but that is only used for input, after that it's merged into the system, all code-like constructs (links, references, images) are resolveed and from then you interact with the higher level concept (rendered comment with links and images). For programinng langauge, Unison does this - you write one function at a time in something like a REPL and functions are saved in content addressed database. > Or what if it was like "staging" step? Yes, and I guess it'd have to go even deeper. The system should be able to represent broken program (in edited state), so conceptually it has to be something like a structured database for code which separates the user input from stored semantic representation and the final program. IDE's like IntelliJ already build a program model like this and incrementally update it as you edit, they just have to work very hard to do it and that model is imperfect. There's million issues to solve with this, though. It's a hard problem.
- zokier 8mo ago> but storage format as text is just fundamentally problematic. Why? The ast needs to be stored as bytes on disk anyways, what is problematic in having having those bytes be human-readable text?
- thesz 8mo ago> Dion language demo (experimental project which stores program as AST). Michael Franz [1] invented slim binaries [2] for the Oberon System. Slim binaries were program (or module) ASTs compressed with the some kind of LZ-family algorithm. At the time they were much more smaller than Java's JAR files, despite JAR being a ZIP archive. [1] https://en.wikipedia.org/wiki/Michael_Franz#Research https://en.wikipedia.org/wiki/Michael_Franz#Research [2] https://en.wikipedia.org/wiki/Oberon_(operating_system)#Plugin_Oberon_and_slim_binaries https://en.wikipedia.org/wiki/Oberon_(operating_system)#Plug... I believe that this storage format is still in use in Oberon circles. Yes, I am that old, I even correctly remembered Franz's last name. I thought then he was and still think he is a genius. ;)
- panstromek 8mo agoInteresting. It looks to me this was more about the portability of the resulting binary, IIUC. Dion project was more about user interface to the programming language and unifying tools to use AST (or Typed AST?) as a source of truth instead of text and what that unlocks. Dion demo is here: https://vimeo.com/485177664 https://vimeo.com/485177664
- thesz 8mo agoI took a look. Their system allow for intermediate state with errors. If that erroneous state can be stored to disk, they using a storage representation that is equivalent to text. If erroneous state cannot be stored, this makes Dion system much less usable, at least for me. They also deliberately avoided pitfalls of languages like C. While they can do that because they can, I'd like to see how they will extend their concepts of user interface to the programming language and unifying tools to use (Typed) AST to C or, forgive me, C++, and what it'll unlock. Also, there is an interesting approach of error correcting parsers: https://www.cs.tufts.edu/comp/150FP/archive/doaitse-swierstra/error-correcting.pdf https://www.cs.tufts.edu/comp/150FP/archive/doaitse-swierstr... Much extended version is in Haskell at Hackage: https://hackage.haskell.org/package/uu-parsinglib https://hackage.haskell.org/package/uu-parsinglib As it allows monadic parsing combinators, it can parse context-sensitive grammars such as C. It's interesting to see whether their demonstration around 08:30 of Visual Studio unable to recover from an error properly can be improved with error correction.