13 ms·
Asynchronous IO in Rust
- vvanders 11y agoReally awesome stuff, great example of how well composability works with traits and the nice handling that tagged unions bring to state machines.
- jerf 11y agoIf you use threads, green or otherwise, you don't have to "implement" special code for composing things together, you get the full set of tools for composing code together, which includes, in passing, state machines, among all the other things it includes. This basically implements an Inner Platform Effect of an internal data-based language for concurrency that the language interprets, which will A: forever be weaker than the exterior language (such is the nature of the Inner Platform Effect) and B: require a lot of work that is essentially duplicating program control flow and all sorts of other things the exterior language already has. There are some programming languages that have sufficiently impoverished semantics that this is their best effort that they can make towards concurrency. But this is Rust. It's the language that fixes all the reasons to be afraid of threading in the first place. What's actually wrong with threads here? This isn't Java. And having written network servers with pretty much every abstraction so much as mentioned in the article, green threads are a dream for writing network servers. You can hardly believe how much accidental complexity you're fighting with every day in non-threaded solutions until you try something like Erlang or Go. Rust could be something that I mention in the same breath for network servers. But not with this approach. There's plenty to debate in this post and I don't expect to go unchallenged. But I would remind repliers that we are explicitly in a Rust context. We must talk about Rust here, not 1998-C++. What's so wrong with Rust threads, excepting perhaps them being heavyweight? (Far better to solve that problem directly.)
- pcwalton 11y ago> What's so wrong with Rust threads, excepting perhaps them being heavyweight? (Far better to solve that problem directly.) Nothing. If threads work fine, use them! That's what most Rust network apps do, and they work fine and run fast. A modern Linux kernel is very good at making 1:1 threading fast these days. > And having written network servers with pretty much every abstraction so much as mentioned in the article, green threads are a dream for writing network servers. Green threads didn't provide performance benefits over native threads in Rust. That's why they were removed. Based on my experience, goroutine-style green threads don't really provide an performance benefit over native threads. The overhead that really matters for threads is in stack management, not the syscalls, and M:N threading has a ton of practical disadvantages (which is why Java removed it, for example). It's worth going back to the debate around the time of NPTL and look at what the conclusions of the Linux community were around this time (which were, essentially, that 1:1 is the way to go and M:N is not worth it). There are benefits to be had with goroutine-style threads in stack management. For example, if you have a GC that can relocate pointers, then you can start with small stacks. But note that this is a property of the stack, not the scheduler, and you could theoretically do the same with 1:1 threading. It also doesn't get you to the same level of performance as something like nginx, which forgoes the stack entirely. If you really want to eliminate the overhead and go as fast as possible, you need to get rid of not only the syscall overhead but also the stack overhead. Eliminating the stack means that there is no "simple" solution at the runtime level: your compiler has to be heavily involved. (Look at async/await in C# for an example of one way to do this.) This is the approach that I'm most keen on for Rust, given Rust's focus on achieving maximum performance. To that end, I really like the "bottom-up" approach that the Rust community is taking: let's get the infrastructure working first as libraries, and then we'll figure out how to make it as ergonomic as possible, possibly (probably?) with language extensions. My overarching point is this: it's very tempting to just say "I/O is a solved problem, just use green threads". But it's not that simple. M:N is obviously a viable approach, but it leaves a lot of performance on the table by keeping around stacks and comes with a lot of disadvantages (slow FFI, for example, and fairness).
- jerf 11y agoI'm down with all that. I'm really arguing in favor of threading here rather than any particular model of it. Plus, if you write threaded code, you can change the runtime out. Rust guarantees all the hard parts, anyhow. I wouldn't even be surprised once Rust settles down further it turns out there's some sort of hybrid solution that is better than either 1:1 or M:N on its own because it has superior insight into what its functions are doing. However, if you are sitting down in front of an editor, faced with the task of writing a network server, and you immediately reach for the overcomplicated solutions like this rather than starting with threads, you are paying an awfully stiff development price for a performance improvement that probably won't even manifest as any useful effect, because the window where this will actually save you is rather small. If you're sitting at 90% utilization, and this sort of tweak can get you to 80% utilization, that's essentially a no-op in the network server space, because it means you still better deploy a second server either way. If you're close to a capacity problem (and the code has been decently optimized to the point where that's not an option anymore), the solution is not to rewrite your code to squeeze out the 10% inefficiency in stack handling, the solution is to deploy another server, and if a rewrite is needed, rewrite to make that possible/effective/practical. In the Servo design space, it makes perfect sense to be upset that you lost 10% of your performance to green threads, because that's time a real human user is directly waiting, and what does a web browser need with 10,000 threads anyhow? (I mean, sure, set one the deliberate task of using that many threads and they can, but in practice you'd rather be doing real work than any sort of scheduling of that mess.) In the network space it's way less clear... if your infrastructure is vulnerable to 10% variances in performance, your infrastructure is vulnerable to 10% variances in incoming request rate, too. (10% is a bit of a made up number. I believe, if anything, it is an overestimate. Point holds even if it's a larger number anyhow.)
- nostrademons 11y agoAre you talking about threads-the-programming-model (vs. events, callbacks, channels, futures/promises, dependency graphs), or are you talking about threads-the-implementation-technique (vs. processes or various select()-like mechanisms)? And if you are talking about threads-the-programming-model, which synchronization mechanism: locks, monitors, channels/queues, transactional memory? If the various threads in your application don't share data, then you don't need to worry about all the pitfalls of threading. But if they don't share data and you don't care about small (~10%) performance differences, why not use processes? Then you can use a really straightforward blocking I/O model, and you get full memory isolation & security provided by the OS. The interesting questions happen when you a.) want to squeeze as much performance out of the machine as possible or b.) need to share data between concurrent activities. Then all of the different programming models have pros and cons, and I'm not sure you can define a "best" approach without knowing your particular problem. (Which, IMHO, validates Rust's "provide the barest primitives you can for the problem, and let libraries provide the abstractions until it becomes clear that one library is a clear winner" approach.) FWIW, my experience in distributed systems is that threads+locks is a terrible model for writing robust systems, and that once you're operating at scale, you really want some sort of dependency-graph dataflow system where you specify what inputs are required for each bit of computation and then the system walks the graph as RPCs come back and input becomes available. This lets you attach all sorts of other information to nodes - timeouts, latency statistics, error statistics, whether or not this node is required and what defaults to substitute if it fails, tracing, logging, load-balancing, etc. It also adds a huge amount of cognitive overhead for someone who just wants to make a couple database queries. I wouldn't use this for prototyping a webapp, but it's also invaluable when you have a production system and ops people who need to be able to shut down a misbehaving service at a moment's notice while still keeping your overall product up.
- tailhook 11y agoWell, I believe that it's almost impossible to make rust threads lightweight because every green thread needs a stack anyway. This may be fixed with some stuff like `async/await`. But let's talk about what's wrong with threads: 1. Timeout handling is ugly: you need to account a timeout in each read and write operation. At least timeout handling makes coroutine/threaded code no better than state machine code. But actually in state machine approach I can set a deadline once and update it only when changed (note to myself: should add such an example to documentation). In Python, it's usually fixed by making another coroutine with sleep and throw an exception to the one doing I/O. It works well, but Rust will never get exceptions (I hope) 2. When you make a server that receive a request, looks in DB then responds, there is an incentive to own (create or acquire from the pool) a DB connection by each thread. This is a sad trick. The better thing when the DB connection is handled by a coroutine on its own. Because in the latter case you may pipeline multiple requests over a connection, monitor if the connection is still alive, reconnect to DB while no requests are active, or the contrary, shut down idle connections. By pipelining you may keep less number of connections to DB so make the load to the database a little bit lower. When I'm talking about DB in this paragraph, of course, I mean everything for which this application is a client. Sure you can do that in threaded code too, but it's much harder to get right. You need two threads per connection (because one reads network and the other looks at the queue and does write), you need to synchronize both sides, connection cleanup code is complex, there is more than one level of timeouts now, so on. 3. You need to avoid deadlocks. Rust takes care to avoid data races, but deadlocks are possible. And they are not always simple or reproducible, so you will have a hard time debugging them. In the single-threaded async code, you are the only user. Even if you have an async thread per processor, you are more likely to own resources instead of locking on them. You may duplicate many things for every thread. You can have more coarse-grained locks, so never hold two of them. But it's almost impossible to write lock-free threaded server. All of the issues above are neither fixed with async/await nor with any M:N or 1:1 threading approaches.
- jerf 11y agoI've been writing in this model for nearly ten years now. In practice, what you cite as problems aren't. 1: In either approach, somewhere in your event loop you're setting yourself a timeout to fire. Haskell & Erlang do use exceptions, but Go does not, it simply makes this a first-class concern of the core event loop. This is only a problem in languages where the threading was bolted on after-the-fact. Which is a lot of languages, which matter because they have a lot of code. I don't mean to dismiss those real problems. But it's not a fundamental problem, only accidental. 2. In practice, this is not a problem I ever worry about. You get a DB library, it provides pools, unless you're talking to a very, very fast DB (like, memcached on localhost fast) this is one of those cases where IO really does dominate any minor price of thread scheduling. 3. This has been solved for a long time. Go has the nicest little catch phrase with "share memory by communicating instead of communicating by sharing memory", but each of Haskell, Erlang, and Go have their own quite distinct solutions to their problems, and in practice, all of them work. There's other solutions I merely haven't used, but I hear Clojure works, too. (Perhaps arguably a subset of the several approaches Haskell can use. Haskell kind of supports darned near everything, and you can use it all at once.) This is part of why I write this sort of thing... at its usual glacial pace (despite how much we like to flatter ourselves that we move quickly), the programming community is finally getting around to being really seriously pissed off about how bad threading was in the 1990s. Good. We should be. It sucked. Let us never forget that. But what has not been so well noticed is that the problems with threading have basically been fixed, and in production for a long time now (i.e., not just in theory, but shipping systems; go ask Erlang how long it's been around). You just have to go use the solutions. Don't mistake debates about the minutia of 1:1 OS threading vs. M:N threading and which is single-digit percent points faster than the other for thinking that threading doesn't work. Lest I sound too pollyannaish about what is still a hard domain, the way I like to put this is that threading has moved from an exponentially complex problem to a polynomially complex problem. (And Rust is leading the way on making even the polynomial have a small number in the exponent.) There's still a certain amount of complexity in making a threaded program go zoom, and it does require some adjustments to how you program, it's not "free", but rather than requiring wizards, it merely requires competent programmers who take a bit of care and use good tools and best practices now.
- mzs 11y agoOne nice thing about Rust is how well suited it can be for embedded use. It is something like a smaller safer easier C++. For embedded you need to bound resources. Small network appliances are a good candidate for async IO. Many threads and stacks with a lot of dynamic allocation are not a good fit for a resource constrained device.
- acconsta 11y agoThere are still other blockers on embedded though. LLVM doesn't target every architecture, and still no allocator API or OOM handling.
- kibwen 11y agoThe dynamically-allocating portions of the stdlib may panic on OOM, but in an embedded context you're not even linking that code into your program. And you're free to provide an alternative stdlib that bubbles up OOM for those rare occasions where you want to dynamically allocate and you want to be able to do something sane in the face of OOM and you're on a platform that doesn't overcommit.
- Gankro 11y agoMinorish point: a true OOM won't panic, it'll abort the program completely. Some APIs may identify that the requested amount of memory is impossible (would overflow) and panic instead, though. Aborting once the allocator has actually been invoked is important for perf and exception safety, though we have briefly mused allowing it to panic: https://github.com/rust-lang/rust/issues/26951 https://github.com/rust-lang/rust/issues/26951
- acconsta 11y agoOOM is an abort, not a panic (contrary to the official docs, interestingly) > you're free to provide an alternative stdlib But like... should you have to? > for those rare occasions In kernel programming, allocation failures are common and handling them is essential. To quote my (very talented) classmate, who actually wrote part of a kernel in Rust: "The only really major issue is with how allocation failure works, which makes rust a (very) poor choice for real kernel development" https://www.reddit.com/r/rust/comments/341v3n/cs_honors_thesis_reenix_implementing_a_unixlike/ https://www.reddit.com/r/rust/comments/341v3n/cs_honors_thes...
- halayli 11y agoYou might want to take a look at: lthread & lthread cpp https://github.com/halayli/lthread https://github.com/halayli/lthread https://github.com/halayli/lthread_cpp https://github.com/halayli/lthread_cpp
- cwp 11y agoIsn't this a rehash of the C10K problem[1] from a decade ago? That was pretty much resolved in favour of single-threading and asynchronous IO, with Nginx and Node.js replacing Apache and Ruby as the platforms that the cool kids use. So, if threads are the way to go today, what has changed in the last 10 years to turn the conventional wisdom on it's head? 64-bit processors and servers with more memory? Hypervisors/containers? Are threads in Rust more practical than they are in C? (BTW, these aren't rhetorical questions, genuinely interested.) [1] http://www.kegel.com/c10k.html http://www.kegel.com/c10k.html
- girvo 11y agoMore cores, perhaps? Though I don't know if that's actually true in a server context.
- dboreham 11y agoNo, backwards: the C10k techniques are the solution to the problem "how do I achieve my performance and scaling goals given the characteristics of the OS, languages, tools available to me today?". The current thread (sic) is asking the question "if we can change the language, tools (perhaps the OS too), what's the best approach?". And: I have to counterbalance the assertion that Node.js is in any way good for anything besides "I need to code in JS, but not in the browser". I'd rather use VAX assembler and the $QIO syscall, typing on a VT100.
- cwp 11y agoThat's exactly the point. So my question remains: if threads are the right answer today, either the C10K folks were wrong or something has changed since then. What?
- krenoten 11y agoModern operating systems are better at handling many threads.
- dboreham 11y agoThe question is a different question, hence the answer is different. C10k wasn't considering the option of new languages, this thread is. The same question could have been asked back then (although most folks wouldn't have considered using anything other than C/C++ for high-scale servers then). Perhaps the thing that's changed is more memory and CPU power has made the option of new (and less efficient than C) languages a practical choice for new projects. I know that if, back then, I'd proposed the solution to my product's scaling issues with network I/O to be : "re-write all 500k loc in a different language", I'd have been laughed at.
- amelius 11y agoIt would be much nicer to use coroutines instead of state machines. This means that you could write code as if it were doing i/o in a "blocking" fashion, but behind the scenes it is actually asynchronous.
- mtanski 11y agoI think a lot of folks think of Network IO when they say Asynchronous IO. That's only half the story, unless you're just building proxies and caches you have to deal with Disk IO at some point in time. And, async disk IO is horrible in every OS / language.
- saaadhu 11y agoCare to explain why async disk IO is horrible? AFAIK, the kernel APIs only deal with abstract handles and descriptors - they aren't written specifically for disk IO or network IO. And C#'s async/await and BeginXXX also work fine with all kinds of IO.
- kentonv 11y agoOn Unix, the APIs for async disk I/O are completely different from the APIs for async network I/O. This is because, on Unix, the APIs for network I/O are "readiness"-based (they tell you when data is available in the read buffer or space is available in the write buffer), but a disk file handle is always ready (except at EOF). So, for disk I/O you need a "completion" model, where you queue an operation and then get a callback when it is done. This is a very different kind of API, and as a result it does not fit well with the network I/O APIs, but of course if you are doing async you probably need to do both disk and network at the same time, so you need the APIs to play nice with each other (you need a single event loop that waits for both kinds of events). (On Windows, as I understand it, you can do completion-based I/O on all handle types.) And then, even if you manage to use async disk I/O, you can only really use it for reads and writes. Filesystem calls that manipulate the directory tree usually don't have async versions. At some point you have to give up and do things in threads.
- stuaxo 11y agoThe new async keyword in python looks a lot nicer than what we had before.
- justincormack 11y agoWell, apparently not on Windows. But this is a kernel interface issue not a language issue, and clearly needs fixing now we have SSDs where having thousands of outstanding requests is useful, if not required for performance, unless the hardware APIs are going to change (maybe if they get memory interfaces this does change).
- geertj 11y agoThis is identical to the Protocol abstraction in Python's asyncio, right?