3 ms·
First, I specifically said that textmate takes less effort for "the most basic highlighting", that would be keywords, literals, comments, etc. Trying to encode
by debugnik 3y ago
First, I specifically said that textmate takes less effort for "the most basic highlighting", that would be keywords, literals, comments, etc. Trying to encode a CFG in it would obviously be painful, but why would you? Use an actual parser at that point, but you might as well use a correct one.
Second, there's this idea that compiler parsers are inherently slow and require a full build to use or something. As if tree-sitter parsers were using some magic tech unavailable to compilers, and most new ones weren't modular or didn't offer APIs. It's not even hypothetical, increasingly more modern languages have semantic highlighting in real-time as part of the toolchain.
And in fact, I actually wanted to use that tech in a compiler, I said I tried to use tree-sitter for one, but that's where my big issue with tree-sitter came in: It would return parse trees that blatantly didn't match the grammar I wrote, and not mark them as bad syntax in any way. That's not error recovery, that's swallowing errors. If it does that, I'd rather work on making toolchain parsers better than fighting tree-sitter.
Even the "run the parser on the editor" part is solvable, since tree-sitter parsers are being shipped as WASM. A WASM interface for running possibly incremental parsers (that tree-sitter would also use) sounds more interesting, really.
- DannyBee 3y agoThere's a lot here, and i'mn not sure where to start. You are complaining about a lot of things simultaneously, and it's hard to tell exactly which point you want to tackle. For correctness - gonna leave this alone, i don't know of issues here and it seems fairly orthogonal to speed which is where we started. I've done parsing and lexing for decades now, including multiple compilers. So i'll try again - for magic tech, tree-sitter does actually have "magic tech" unavailable to most compilers, and i explained exactly what it was - but you simply want to deride it for some reason - not obvious to me. the magic tech is that 1. tree-sitter ast/lex tokens (for purposes of incremental lexing/parsing) are compiler/language agnostic. it does not care. This is not true of most compilers 2. it allows incremental lexing/parsing when you keep those ast/tokens under the text of your editor. It only needs to be informed of what edit was made (add/change/remove of a range of characters is fine) and have random access to the text. This is not true of most compilers. 3. The incremental lexing/parsing is extremely fast. It is optimal in the amount that has to be re-lexed/parsed, which for both is bounded by the amount of max used lookahead in the grammar, which for most languages, is really small (again, not the theoretical lookahead, but the actual-used-during-lex/parse lookahead). Not true of most compilers. Can you design compilers that can do this? Sure, i was at IBM when we did visualage, which could do this in other ways, though not optimally. As mentioned, i've even implemented the optimal incremental lexing/parsing algorithms in other parser/lexer generators, and even some compilers. Is it common? Not a chance. Your claim that increasingly more modern languages have semantic highlighting in real time "built in" is simply false - most cannot real time semantic highlight even a 100k file on every keystroke. There are a very small number which can, and it's mostly by luck - they fall down on larger files or other complex cases because they have no incrementality. clangd is about the closest you get. rust-analyzer definitely can't. etc Lexing has always been really fast, and doable in a single pass, that has never been the problem. Meanwhile, what i just described (updated highlighting on ever keystroke) is trivial with tree-sitter for all languages because of its optimality and agnostism. If you want to see your claim in action - turn on semantic highlighting for vscode and type fast - most of the time you will lose syntax highlighting because things can't keep up. see, e.g., https://github.com/microsoft/vscode-extension-samples/issues/851 https://github.com/microsoft/vscode-extension-samples/issues... Again, yes, you could make a compiler for every language which supports a mode that does what tree-sitter does. and people could carefully implement parsers/lexers in each of their favorite compilers and languages that support optimal incremental parsing in them. and then pay the cost of integrating 50 language specific ways of transforming these to work with the editor, basically reinventing what tree-sitter already did right. (As per above, the current semantic token and highlighting support does not resolve this, and the interface it provides can't support it, last i looked). You seem to really otherwise be complaining that you believe it does always generate correct parsers. Like I said, that seems totally orthogonal to anything about the speed issue, and if that's your real concern, have at it.
- debugnik 3y ago> You are complaining about a lot of things simultaneously, and it's hard to tell exactly which point you want to tackle. Thanks, that's exactly how I felt about your first reply. > tree-sitter does actually have "magic tech" unavailable to most compilers No, not using it doesn't mean it's unavailable, at best it means the toolchain is up for an upgrade. > the magic tech is that I too have written compilers and language servers, even tested novel parsing techniques. I don't need tree-sitter explained to me, I understand the theory behind it perfectly well and, as I already said, understand what they're going for in terms of ecosystem. I just don't think it's the sweet-spot you think it is. In any case, making incremental edits is not novel, it's easy even with good old recursive descent by simply traversing the ast, skipping the prefix of the change, reparsing the inner node and checking that the edit was well-bounded (hopefully your language doesn't have multiline strings, or you'll have to accept that some incremental changes consume the whole file in the worst case anyway). I've done this, and again, if some toolchain doesn't then it's up for an upgrade. The actually interesting parts of tree-sitter, to me, are the error recovery and the shared parse tree interface. The latter is orthogonal to the parsing technique, yet with tree-sitter it's all or nothing. The former would be great if it actually marked the error in the parse tree, which it didn't when I checked and it wasn't deemed important. > most cannot real time semantic highlight even a 100k file on every keystroke Even you concede that the bottleneck here is not usually parsing, and even then I just said it can be made incremental easily, but all the other work that the compiler is choosing to do before answering back to the editor. When this work is unnecessary (at this step anyway), then sure, you can speed past them by doing the work yourself on the tree-sitter parse tree. > you could make a compiler for every language which supports a mode that does what tree-sitter does Yes, we could. Maintaining and testing tree-sitter grammars also has a cost, maybe the difference is that compiler maintainers are externalising this cost right now. > and then pay the cost of integrating 50 language specific ways of transforming these to work with the editor No! We went through this with LSP. We can have any number of parsers with a shared driver interface, maybe like tree-sitter's, but we don't need to push every grammar through a single parser toolkit that isn't good enough to replace the existing parser in the toolchain. In fact, tree-sitter grammars are shipped precompiled to editors, so that's kinda the situation already, but their current interface AFAIK expects a particular implementation on the parser (theirs). > You seem to really otherwise be complaining that you believe it does always generate correct parsers I guess you mean doesn't. Then yes, but that's not an "otherwise", that's a big deal, because if tree-sitter parse trees were actually sound I could just maintain a single tree-sitter grammar for the entire toolchain, compiler and editor, and this debate would be redundant.