4 ms·
As someone actually working on some symbol indexing stuff at the moment, I can mention that this model is ignoring some pretty important details, and I think th
by ninepoints 3y ago
As someone actually working on some symbol indexing stuff at the moment, I can mention that this model is ignoring some pretty important details, and I think there's some classic overengineering happening here (easy trap to fall into when you're still in the design space).
Generally, when you need a code completion action, it's expected that you have an AST (post semantic analysis). This AST is still very useful even after a user continues to edit the code! Outside of where the cursor is, most of the symbols, definitions, and declarations are relevant, and source locations can be easily translated from a previous version of the document, provided that deltas are tracked. Cancelling an existing translation unit parse on edit is wasteful, because the majority of the time, that parse will produce meaningful results. The better approach (IMO) is to let the parse finish, but immediately enqueue a subsequent parse (with some debouncing timer to avoid overly consuming user resources). If you wanted to level this up further, you could perform incremental parsing/analysis, provided your language supports it. In the presence of a preprocessor, this can be very difficult, but it's the next "upgrade" from the previous approach mentioned in my opinion.
- matklad 3y agoThis depends on the compilation model in use. If that’s a traditional pipeline of phases, where the result is an AST data structure for the whole CU which gets annotated with types, then, yes, “enqueue new analysis” makes sense. If the compilation model lazy & query based, then theres’s just no “enqueue a subsequent parse”, but rather “give me the type on this thing doing as little analysis as possible”. You got to re-use symbols and declarations not because you heuristicsly adjust offsets and assume they are otherwise valid, but because underlying analysis reasons, precisely, they they are valid and re-usable. The two models are quite different at the core, and aren’t really an evolution of one into another. Which of the two approaches to use, and, consequently, whether to think about cancellation at all, depends heavily on the language in question. If the language allows for lazy analysis, it’s probably a better bet, as that gives you correct results faster.
- ninepoints 3y agoI only know a little bit about zig, but I would have assumed that the latter mechanism you are describing isn't really possible. Most languages with "advanced" features and type systems really need a full semantic tree to do anything meaningful due to the complexities of compile time code and type instantiation. Even if we are using this latter model, the user is only editing one file at a time. What if the completion source is in a different file altogether? Generally, you're going to be indexing the entire codebase anyways, so we're splitting hairs over a potential optimization in just one facet of the indexer.
- 59nadir 3y ago> Even if we are using this latter model, the user is only editing one file at a time. I'm not necessarily arguing against your bigger point but this is very frequently a wrong assumption and strikes me as designing around an idealized view of the problem that you would like to be the case, not actual reality. Code generation scripts, formatters, auto-fixers and the like can modify many files and many parts of those files "at a time", unless you have a pedantic view of what "at a time" means. Almost all LSPs fail in these modes of operation and disallow external tools from participating in the code base, which is very annoying and makes them less useful. Having to actually (re-)open a file so the LSP can "see" changes made to it by an external program is not something that should ever be needed but happens a lot with some of the most worked-on language servers (TypeScript comes to mind).
- ninepoints 3y agoEither way, all those files need to be marked dirty and reindexed, so I suppose I'm just not sure how the topic at hand is relevant. Is the proposal that we attempt to treat each dirtied file as though incremental edits were imminent? Because this is precisely the wrong assumption you're raising. Ultimately, if you need to reopen a file in order to see changes on disk reflected, that's a bug with your lsp server, nothing more.
- 59nadir 3y ago
- levodelellis 3y agoI pretty much entirely agree with you. Below is a copy/paste of what I said when I saw this article elsewhere. What language is your LSP for? > I written a LSP for my prototype compiler. I don't like any of the options you listed. My LSP didn't do any typechecking, it didn't build an AST, it didn't need immutable data structures etc. > Typically when a person is typing into the editor the code is in a broken state (incomplete variable name, missing semi colon, maybe an open but no close parenthesis etc). What I did was look around what part is being edited and using the previous 'build' (when a user saves or ask the compiler to build), I would look up vars and type names. There's no need to rebuild everything on every keystroke. Maybe you can do it on a newline if you really wanted to but midsentence sounds like a bad place to try and you're not really gaining anything from compiling/parsing a single line change