3 ms·
TLDR summary: 1. Straw man 2. Talk about an unrelated paper 3. Conclusion: I am smart and those people are dumb On a somewhat related note: I'd say 95% of the
by exfalso 3y ago
TLDR summary:
1. Straw man
2. Talk about an unrelated paper
3. Conclusion: I am smart and those people are dumb
On a somewhat related note: I'd say 95% of the software I've written was IO bound(and if it wasn't at the beginning then the goal was to make it IO-bound), and that remaining 5% does not benefit at all from async/coroutines(used them in Haskell, Kotlin(Quasar) and Rust). I'm very curious about what real-world CPU-bound use cases people have that can benefit from async performance-wise. If you're optimizing on that microsecond-nanosecond scale then you shouldn't even have yields/blocks on your hot paths, and synchronization will at most be done using memfences, so what are we talking about?
I'd also conjecture that if you place your problem domain on the IO vs CPU bound axis, the problems where async would provide performance benefits can invariably be solved by other designs that perform even better (GPU/FPGA).
Yes I'm one of those people who thinks coroutines have extremely marginal actual value, and most of that value is the "feeling of how cool this is" that people experience when they first learn about the concept. If there was a way to bin the concept altogether I would do it in a heartbeat. As-is, the entire Rust ecosystem is suffering heavily because of it.
- slaymaker1907 3y agoA lot of the time, the benefit can be from much more controlled scheduling without having to introduce actual threads. A lot of software doesn't want to take the overhead of creating platform threads (memory and taking additional CPU time) but does want to do scheduling internally for various CPU bound workloads. I think web pages are often like this. You don't really want to hog tons of CPU, but you also want to have guarantees like not blocking rendering when running a user provided regex since said regex could take a long time to run, but in practice that regex will be fast. Even if we want to eventually offload that to a separate thread, async gives us the option to first try evaluating it in the main thread for some number of iterations before offloading it. GPU/FPGA programming is kind of irrelevant because as expensive as developing high performance code is for CPUs, costs go up by an order of magnitude for those platforms unless there is some existing library you can utilize (mainly applicable for ML/AI). These platforms are also very expensive. It's like saying just rent a helicopter if you need to get somewhere quickly instead of asking what "quickly" actually means and optimizing accordingly.
- exfalso 3y agoSuspendable computations like regex statemachines are actually a good example use case. However I'd still categorize this as extremely marginal. I have seen literally zero Rust crates that handle suspendable computations (ofc outside of the scheduler runtime). Also note that by their nature regex statemachines do not actually necessitate the use of async, it's effectively syntax sugar on top. My point with GPU/FPGA is that if you're at the point where this level of optimization matters then you're actually dealing with situations where it's worth to invest in the big boy tools. Examples are HPC in fintech (low latency trading) and scientific computations, game development, video codecs etc. You know, "actual" computations. Webservices are not in this category. Generally speaking with web services your goal is to "hide in the shadow of IO". If you max out your network and/or database and/or filesystem capacity, further CPU optimizations will have literally no effect. I have yet to encounter a web service where this wasn't the case. What I do see sometimes with webservices is simply unnecessary compute, bad internal structuring, dynamic dispatch, parallelism overcommit, lack of batching, fragmented apis, fragmented data accesses etc etc all of which appear as CPU capacity saturation and also sometimes as kernelspace overhead in profiles. Async does not help solving any of these issues. And once you do solve them, you reached IO boundness and it doesn't matter anymore. Again this is just my experience, and I'm happy to learn about what kind of web service can utilize coroutines with measurable performance benefits over a managed threadpool.
- TexasMick 3y agoI write software for FPGA soft cores. Even with all this acceleration around the soft core, we need some sort of scheduling and kernel context switches are really a big killer. We have the same issue as web developers, dealing with 50,000 things per second means we need to avoid kernel context switches.
- exfalso 3y agoAvoiding context switches is not a problem async solves. A threadpool popping work items from a queue(or something like lmax disruptors) has the same effect on context switches. The only thing one could argue is that the async runtime's threadpool "homogenizes" work. Again, yet to encounter a case with web services where this was the issue. I'm not familiar with FPGA scheduling, are there any resources that explore the issue? I was under the impression it's akin to GPU compute where the main bottleneck is the bus.