3 ms·
I'm not talking about whether the scheduler can block the task on IO. This is trivial. Let's see a couple of examples. Database connection contention: as
by exfalso 3y ago
I'm not talking about whether the scheduler can block the task on IO. This is trivial.
Let's see a couple of examples.
Database connection contention:
async function() {
let connection = getConnection().await; // Allocate resource
connection.execute("SELECT 'Hello World!'").await; // <- yield to the scheduler, connection is held onto until continuation is scheduled
connection.close().await; // Release resource
}
This is an extremely trivial example, and this already leaks resources. If you launch 100 of these then there is a chance that there will be 100 open connections at the same time. Had you used a simple thread pool, the size of that pool limits the number of open connections.
Memory leak:
async function() {
let object = stream.readLargeObject().await; <- Memory allocated
process1(object).await; <- Memory leaked to scheduler until continuation is scheduled
return process2(object).await;
}
Again, very trivial example, already leaking. process1 and process2 are not tied together with direct control flow edges but rather the flow dispatches through the coroutine scheduler, and they are holding onto the resources while sitting in the scheduler's bookkeeping. Again, yields disrupt the allocation's scope, and if you launch a lot of these coroutines you'll run out of memory.
This issue is made worse if you need to use combinations of resource allocations. For example DB connection + memory allocation or network + disk or even network+network are common combos. Depending on the complexity of the allocation nesting you can run into very nasty contention issues or even deadlocks, where - again - a threadpool would have worked well. N threads, N number of resources. Done.
I want to stress that there are ways to mitigate the above (queues/resource pools/semaphores), however they are not nearly as intuitive as using a threadpool, and you need to be constantly aware of these leaks when you're writing async code.
- clavigne 3y ago> I'm not talking about whether the scheduler can block the task on IO. This is trivial. The whole point of coroutines is that they are trivial to block and resume, right? So the way to deal with limited resources... is to block on them. So when you need a connection, you await until one from the pool becomes available. From the point of view of the consumer it's very natural because grabbing a connection is just a standard async call, and returning it to the pool can be done the same way you would usually close it. If anything, it's a lot easier to manage them like this because unlike in a non async program you don't have to worry too much about blocking progress.
- exfalso 3y agoThat's not my point. The issue is that when the blocking happens, with coroutines control is yielded to the scheduler which will now schedule other tasks. Those tasks may again request and block on resources. This is where the leak is coming from. A resource pool is one way to get around this, however this stops working if you have several kinds of resources. On the other hand, with threads the IO block is a "proper" block. No new task will be scheduled, the thread will only continue when the IO operation finishes, providing a very natural backpressure mechanism that prevents overallocation/contention.