6 ms·
Why building a Rust LSP is hard
- mitxela 16d agoUsing a JSON TCP connection for what on Windows would be direct function calls (COM) or in Eclipse would be direct function calls between Java modules always felt a bit gross.
- throwaway17_17 15d agoI agree, the only reason LSP exists as it does is for a world running on electron based applications. Plug-ins and application extensions are not new technology and they are nearly universally more efficient in the forms designed and used prior to 2005-ish. I understand why VSCode exists and why it is used so often by developers, but it is certainly a downgrade from more language specific options that could exist. There are some arguments that do carry water in favor of using a client/server protocol transmitting JSON, in particular, the ability for a nearly complete decoupling of analysis of code and the displaying and editing of that code. Also, LSP was (to my knowledge) the first language/platform/usecase agnostic protocol intended for use in code editors. I get why this is where a big chunk of developers have ended up, but I do bemoan the lost potential for a language to grab mindshare and popularity on the usability and performance of its tooling and developer experience via-a-vis a custom designed and hyper specific code editor. I mean, Rust and Elm received endless praise for their error messages as a massive boon to developer experience, so it is a facet of language design and implementation that can act as great advertisement. I just hate that the prevalence of LSP at this point precludes custom editors as a first choice in the current zeitgeist.
- jiehong 15d agoLSP helps with not having to develop the same tooling in each editor for each language. This isn’t for Electron app only, but is also helpful for Emacs, vim or helix (especially when they lack said language plugins). Now, some IDEs provide capabilities for a language far exceeding what can even be implemented in a LSP.
- stevenhuang 15d agoThis comment makes no sense. You realize how much lsp is used in other editors like vim/neovim? The existence of lsp is exactly what lets custom editors to flourish. Any language/lsp client can add their own extensions too. Any new editor can come up with it's own way of doing things and implement it on top of lsp, as long as the lsp client supports those extra features then there's no problem. I don't know how you can get it so wrong.
- mitxela 15d agoThey are not custom editors, they are all the same editor because they have to implement the same LSP protocol.
- imtringued 14d agoYou don't know what an editor or LSP is, if you genuinely believe that.
- pjmlp 15d agoIt only makes sense in the context of host security and stability, as proven by plugin issues in those IDEs. However, to come back to your point, there are much better high performance OS IPC mechanisms for inter process communication than sending JSON down through a TCP wire. As others point out, Electron.
- bheadmaster 15d agoIn order to have direct function calls, you need to load the plugin's code in your own memory. Which means you expose your own memory to the plugin. Any malicious/buggy plugin could wreak havoc on your program - even managed code doesn't solve that problem in general. IPC is the natural solution to that - run a process in its own address space and communicate with it via I/O. TCP is not most optimal, but it is uniquous. Same thing with JSON. In LSP, most of the processing happens in the LSP server anyway - communication overhead between client and server is negligible in comparison to it.
- mitxela 15d agoFault isolation is good, but are the fault isolation benefits worth everything else though? Windows COM does have a way to run the server out-of-process with a performance cost, and this fact is transparent to the client (if it only uses COM interfaces and not global variables or something). I assume you can even switch between the modes at runtime (with restarting the plugin).
- braiamp 15d agoWhy not implement unix sockets too, since we want to avoid the ip/tcp stack? Now you have two platforms to support for no real gain. Also, LSP applications are shared between projects/users, how you centralize that if each system needs its own installed program and dependencies? Using networking is the path of least resistance and it's drawbacks are well understood between the ones that need to communicate, there's no point to implement IPC based communication.
- simonask 15d agoI think this is a very reasonable question to ask. Language servers are certainly a significant source of slowdowns and editor unresponsiveness, and any time you're doing IPC, that's another source of latency and fragility. My current Zed editor instance has 5 different language servers running, implemented in at least 3 different languages, some of which require a full VM runtime. Could Zed (written in Rust) host a full .NET runtime to run the Roslyn language server for C#? Could an Electron-based editor? It feels a bit daunting. Plugins could be native shared libraries, but what happens when multiple language plugins all try to initialize huge process-wide runtimes like .NET or the JVM? Even multiple instances of the Python runtime are going to potentially be stepping on each others' toes. IPC and separate processes is probably the pragmatic solution for now, even though I would love to live in a world where in-process was more of an option. WASM could be the solution, but would lock language server authors into the subset of programming languages that can be compiled to WASM, and WASM still looks a lot like IPC in practice. (Zed already uses WASM components for its plugins.)
- Panzerschrek 15d agoTCP isn't really necessary. Usually a language server just uses stdin/stdout, which are just pipes. And overall overhead for running a language server in a separate process isn't that big. LSP is designed in such a way that only minimal amount of information is needed to be passed, like edits or short responses. Packing/unpacking JSONs isn't a bottleneck, the heaviest job like program analysis is done in the language server itself without interprocess communication overhead involved.
- octoberfranklin 16d agoEnter hell: LSP assumes that it's the source of truth, but you still need to access the filesystem yourself, and do it in a synchronized way LSP is an example of utterly horrid technical design. Stop letting Microsoft design protocols and APIs. They are so. bad. at. it.
- duttish 16d agoI've never looked into LSP, how would you design it?
- packetlost 16d agoWell for one I wouldn't design it so both the LSP and the editor need a synchronized view of the underlying file.
- mitxela 16d agoStart with synchronous function calls instead of JSON. Microsoft knows how to do that - they invented COM and OLE. Function calls enable whatever data sharing is necessary to maintain a coherent view. Imagine trying to do OLE with JSON - just wouldn't work. (Does OLE still exist?)
- cyberax 15d agoStupid idea. So the LSP crashes and/or goes into a runaway memory consumption loop. And your main application dies with it. Or what if you want, you know, to be able to use the same LSP from TWO different applications at the same time? Never mind issues with other managed runtimes not expecting to deal with something else in their address space.
- cyberax 15d agoMoreover, it's not just stupid, it also does not actually solve _anything_. A synchronous function can also have obsolete indexing information if it races with the code updates.
- klodolph 15d agoThe particulars of Rust make this a little more difficult, I think. There’s a certain tension between making your language more concise and adding useful redundancies, and Rust has generally gone to the “concise” side, with some redundancies that can make the tooling a little more painful. Like with imports. impl std::fmt::Display for Blah { } If your language makes you qualify your imports (like above) then your LSP can, delightfully, still reliably do certain ops like renaming, even when chunks of your project aren’t parsing. But if you glob import std::fmt, and glob import something else, you are fucked. Display could come from anywhere (maybe from a module that has a parse error at the moment). I really appreciate languages where glob imports (or their equivalent) are either disallowed entirely or where typical code doesn’t use it. Meanwhile, if you add a new file, there’s this little dance where you say: mod mycoolmod; And then you create mycoolmod.rs. Or you do it the other way around. A little redundancy (the file exists and it is declared), that seems to just create a little friction in the LSP because mycoolmod doesn’t get a working LSP until it’s declared in the parent (you have to create both, and then you get a transient diagnostic that your module is unused for a while yet). A small issue, just another little bit of friction in the tooling of Rust that has nothing to do with the type system.
- verandaguy 15d agoHaving worked with rust for nearly two years now (granted, on one team with agreed-upon standards): - glob imports are rare in my experience, less for the LSP’s sake and more for code self-documentation - the `mod foo` line exists because omitting it cannot fall back to a reasonable default visibility level (`pub`/`pub(crate)`/<none> (private))
- yencabulator 15d agoOf course it could just default to private, and if you wanted something you'd have to specify that.
- yencabulator 15d agoWildcard imports are a horrible idea. (And anyone pushing to have a "prelude" for their library is making it worse. Please don't.) Clippy `wildcard_imports = "warn"` is your friend.
- Panzerschrek 15d agoA lot of things described in this article applicable not only for Rust, but for almost any language. Like it's obvious that requests should be handled asynchronously and that conversions from/to UTF-16 are needed. But it's actually not so hard. I have written a language server for my language too. The hardest thing was to find a way allowing providing useful autocompletion for a document in edited state, when it's not syntactically-correct. This is the trickiest part how to deal with such incorrectness without missing all the context necessary.
- rhdunn 15d agoI've not written a language server but have written a language plugin for IntelliJ. I started with writing a correct recursive descent parser. I then extended it to detect, report, and recover from common syntax errors as I encountered them so that the parser is robust. And adding a parser test case for each of these (e.g. one test for each branch through an EBNF construction). Some examples are: 1. missing keywords when the keyword can be detected from the current context (e.g. missing semicolon at the end of a statement); 2. using the wrong token (e.g. `:` instead of `::` in a C++ namespace qualified name); 3. detecting and ignoring whitespace in a whitespace-sensitive qualification (e.g. in XML QNames); 4. keeping in the prolog state (where functions are defined) when there are errors so that functions after the error don't get lost; 5. lexing incomplete literals like `10e` so they can be handled as integers in the parser and emitting an error for them.
- Panzerschrek 15d ago> recover from common syntax errors It's a dead-end. Sure, it can work in simple cases, but there will be always a case where such syntax recovery isn't possible. That's why relying only on syntax recovery isn't an option. Because of that I use a different approach. I do parse on each document editing, but such parsing is guaranteed to produce valid results only up to the point with broken syntax, where editing usually takes place. Such parsing is enough to reconstruct location of the point where editing takes place (namespace/class/function) and to reconstruct local context (local variables declared prior to editing place). This allows to perform almost perfect autocompletion by suggesting global and local names available at the editing point. In order to provide proper suggestion of non-local names declared after the editing point, I do keep a structure for the most recent document state with valid syntax. With features like "go to definition" I do the same. I store a hash-table with location to definition point mapping, but it's updated only from time to time and only if document syntax is valid. In order to be usable for cases with edits made after building such hash-table I just perform text-based position mapping using accumulated edit events.
- deleted 15d ago[deleted]
- chrisjj 15d ago> Why building a Rust LSP is hard Not as hard as understanding what LSP means, apparently. I was interested until I saw this project is simply an LSP server.
- mohd_rafay 15d ago[flagged]
- weinzierl 15d agoReading this made me realize that I want two different things from an LSP that are sometimes at odds with each other: 1. Editing help 2. Reliable and comprehensive analysis The editing help needs to deal with incomplete and inconsistent state and answers on a best effort basis. This is good when I'm writing code. When I'm trying to understand code it usually is in a complete and compiling state but best effort is not enough. I expect complete and exhaustive answers. Independently of that I'd love to read a similar analysis that compares the approaches of rust-anslyzer, rust-glancer and the JetBrains analysis engine in Rust Rover.
- ashkankiani 15d agoThis is a good observation, with the observation that there are two different levels of latency requirements. Most of the time, I would be pretty happy with an untyped, unsemantic, mildly smart heuristic based ident completion while the asynchronous semantic one finishes. The nice thing about this dual setup is that I tend to only want the semantic one if I'm thinking more, so there's naturally a larger time budget for it. One must always think about the experience they want in UX first, rather than the tools they want to build.
- wannabe44 15d agoI feel most LSPs I used are fairly good at both. Sometimes I get no completion suggestions, look at the sidebar and fix compile errors, then start getting suggestions. For whatever reason, it's a fine trade off for parsing errors.
- dominotw 15d ago[flagged]
- qudat 15d agoTangentially, since I now write a lot of Zig and there’s no official LSP, I decided to see what life would be without an LSP for all my projects (Go, TS, Python). Honestly, my LSP requirements are minimal: go to def, find all references, and symbol search. I’ve been wanting to try ctags for awhile and finally made the switch. Another consideration were all the posts about how grep is better than LSPs for LLMs. It really pushed me to just get better at grep. There have been a ton of benefits: I can symbol search outside of my editor, it’s much faster than an LSP, minimal configuration. It feels more like a universal tool I can use instead of a tool for IDEs.
- monocasa 15d agoI do like that vscode understands paths with line numbers in the terminal, so grep -n then ctrl+click takes you to that line.
- wredcoll 15d agoIt's faster than an lsp?? What on earth are you doing? Even just the context switch to typing the right grep command is slower than ctrl-g or whatever.
- qudat 15d agoLSP load times can be brutal (or crash entirely) and I have a shortcut to grep the word under the cursor. Ctags is a symbol indexing system so I run it once for a project and can do symbol searches in my editor or using fzf. I have ctags index the languages stdlib as well which is really handy for quick symbol searching -> open my editor to read more