8 ms·
Not mentioned in TFA, but I'm utterly convinced that the only compelling reason for async is to avoid the per-thread stack memory allocation of 8MB per thread o
by doubleunplussed 6y ago
Not mentioned in TFA, but I'm utterly convinced that the only compelling reason for async is to avoid the per-thread stack memory allocation of 8MB per thread or whatever it is, in order to be able to scale to an extremely large number of concurrent threads/coroutines. You can't do this with threads.
Async lets you do this whilst still storing the state of unfinished things in a stack, i.e. not having to make a million callbacks. Making it look like you're using threads even though you're not.
The better async code and interfaces become, the more it looks like regular old multithreaded code, except there aren't actual threads underlying it. You still need to make sure you serialise access to shared resources, don't share data that shouldn't be shared, etc. All the same considerations as multithreaded code.
Almost ends up looking like the underlying mechanism: threads with a GIL vs async ought to be an implementation detail that doesn't require you to modify your entire programming model.
Extension code not holding the GIL can still run in true parallel with real threads, so that's a meaningful difference but is usually not relevant for IO where async is usually used.
- nurettin 6y agoThis explanation works great as long as the underlying implementation of threads in your particular python implementation is ultimately concurrent, but not parallel.
- matheusmoreira 6y ago> I'm utterly convinced that the only compelling reason for async is to avoid the per-thread stack memory allocation of 8MB per thread Yes. The point is to use a single thread with one stack to process several tasks. This consumes less memory. For example, Go currently has a minimum stack size of 2 KiB so a machine with 4 GiB of memory will be able to process less than 2 million goroutines. An event loop uses a single thread with a single stack, reducing memory usage at the cost of complexity. Asynchronous functions are just like coroutines. The difference is they return to the awaiting caller instead of yielding to another function. The order of execution is determined by the underlying loop.
- dorfsmay 6y agoAlso there is a point where managing many threads become a load on the CPU, which my inderstanding is why Rust is not provinding green threads anymore.
- calpaterson 6y agoWhy are you "utterly convinced" of that though? If I run 20 threads in Python, unix top reports resident memory usage to me as 14mb. Why do I see that number instead of 160mb? Code sample: https://gist.github.com/calpaterson/ab35377da9275ca3af7072dbd03ec3a6 https://gist.github.com/calpaterson/ab35377da9275ca3af7072db...
- Scarbutt 6y agoYou mean 14MB?
- calpaterson 6y agoHtop shows me "14016" as resident memory and I don't think it is using kilobytes as the unit.
- anonymoushn 6y agoIt is using kilobytes as the unit.
- calpaterson 6y agoYes you're right. What a silly mistake. Normal Python interpreter = 9mb, with 20 threads = 14mb. Edited the original post.
- mlyle 6y agoAnswer to your question was already discussed right here: https://news.ycombinator.com/item?id=24429221 https://news.ycombinator.com/item?id=24429221
- mlyle 6y agoThat's embarrassing.
- 6y ago
- harikb 6y agoThere are also other language implementation styles that get roughly the same benefit without writing async and await all over the codebase. If the implied yield at defined synchronization points coupled with a decent scheduler like in Go would make this all a non-issue [1]. [1] https://journal.stuffwithstuff.com/2015/02/01/what-color-is-your-function/ https://journal.stuffwithstuff.com/2015/02/01/what-color-is-...
- lmm 6y ago> avoid the per-thread stack memory allocation of 8MB per thread or whatever it is, in order to be able to scale to an extremely large number of concurrent threads/coroutines. You can't do this with threads. It's 8 K B per thread, so you can scale a thousand times further than you thought. One dark secret of the async movement is that if your goal is C10K (10,000 concurrent clients) then actually bog standard threading will handle that fine these days. > The better async code and interfaces become, the more it looks like regular old multithreaded code, except there aren't actual threads underlying it. You still need to make sure you serialise access to shared resources, don't share data that shouldn't be shared, etc. All the same considerations as multithreaded code. Depends what approach you're using. I prefer making an explicit distinction between sync and async functions ( https://glyph.twistedmatrix.com/2014/02/unyielding.html https://glyph.twistedmatrix.com/2014/02/unyielding.html ), so you effectively invert the notion of a "critical section" - instead of marking which sections can't yield, you mark which sections can yield, so your code is safe by default and you can introduce concurrency explicitly as and when you need it for performance, rather than your code being fast-but-unsafe by default and you're expected to fix a bunch of rare nondeterministic bugs with minimal support from your tools, which is how it works in a multithreading world.
- mlyle 6y ago> It's 8 K B per thread, so you can scale a thousand times further than you thought. One dark secret of the async movement is that if your goal is C10K (10,000 concurrent clients) then actually bog standard threading will handle that fine these days. Default virtual memory allocation for threads on Linux distributions tends to be 8 megabytes. Actual memory used is the peak stack depth used, rounded up a bit. It'd be pretty unusual to only use as little as 8 kilobytes per thread; just the standard per-thread libc context information for concurrency is a few kilobytes, plus at least one page of stack, plus the kernel's information about the thread (which isn't counted against the process)... Yes, you can spawn thousands of threads on relatively modest hardware; I was spawning thousands of threads a decade ago. Spawning 5000 bare-minimal python threads that do nil seems to use about 300 megs of ram on my system; real threads that do anything substantial will use a whole lot more, even if their use of the stack depth is intermittent. Not to mention allocators that cache part of freed heap per-thread, etc.
- otabdeveloper4 6y agoThe "8 megabytes" is virtual memory. You're only incrementing a counter in a table, nothing is actually allocated until you actually start using that stack.
- jnwatson 6y agoIn general, it is far easier to reason about locking with async code, as the number of preemption points is far far lower. As a result, you get very low overhead inter-task communication.
- laurencerowe 6y ago> I'm utterly convinced that the only compelling reason for async is to avoid the per-thread stack memory allocation of 8MB per thread or whatever it is, in order to be able to scale to an extremely large number of concurrent threads/coroutines. If it's possible to avoid shared state then I tend to prefer threads, but in the presence of shared state there are good reasons to think that explicit coroutines are easier to reason about than threads or green-threads: https://glyph.twistedmatrix.com/2014/02/unyielding.html https://glyph.twistedmatrix.com/2014/02/unyielding.html
- gpderetta 6y agobut then you need to forfeit parallelism. If you add parallel execution to (shared memory) async, the advantage disappear.
- laurencerowe 6y agoWith Python you forfeit parallelism either way due to the GIL.
- justsomeuser 6y agoIt’s not the only reason. For me, async code is easier to debug and understand because it is composed of regular functions (tagged with async). This means if you have a function that is 10 layers deep, you can pause the debugger and see the stack context. Same with your IDE, you can jump to each function. With threads that 10 function stack would be 10 threads, each requiring some sort of tooling at both software write time and runtime to get the context. In summary composing a system of just “functions (sync + async)” is easier than vs “functions + threads”
- gpderetta 6y agohum, I'm missing something. The same call stack you get with async would map to exactly one thread call stack. Sure each thread will get its own call stack, but all related continuations that participate in an async call stack would be owned by the same thread (Assuming the same application design).
- justsomeuser 6y agoMy point is that with async both the runtime and the dev tools try to make async call stacks look exactly like sync ones. This keeps the context of how your code gets into certain states. If you have 10 threads, it’s like you have 10 different processes (without the context of how they are related - which you get with composing functions) which is harder to understand. Also assuming that each thread does not have an event loop.
- arethuza 6y ago"With threads that 10 function stack would be 10 threads" Presumably only if you have one-thread per function - which I guess you could do but not sure if anyone actually writes multi-threaded code that way.
- justsomeuser 6y agoIf you wanted each function to wait for IO or an event without blocking a thread you need an event loop, or one thread per function that needs to block to wait on incoming events.
- ynik 6y agoIn a multi-threaded context yes, reduced memory usage is the main benefit. But async/await can also be used for other things! In a C# Windows GUI application, it's normal to use async/await on the UI thread. Your UI can await multiple tasks at the same time, yet you don't need any locks when accessing the UI state; because all your code runs on the UI thread. This is a really useful programming model made possible by cooperative task-switching via `await` on a single thread. Here the "await" being explicit is a crucial feature, it allows the programmer to reason about when the shared state might be mutated by other tasks (or maybe by the user clicking cancel while the current task is waiting). Any pre-emptive task switching adds a lot of additional complexity and isn't really suitable for UI code.
- spaetzleesser 6y ago“Your UI can await multiple tasks at the same time, yet you don't need any locks when accessing the UI state; because all your code runs on the UI thread.” Is that true? I thought the code after each await runs on a different thread from the code before the await. At least that’s what I have observed when debugging things.
- heeen2 6y ago> I'm utterly convinced that the only compelling reason for async is to avoid the per-thread stack memory allocation of 8MB per thread I needed to write an interface to a SSE API (tldr: long running socket connection with occasional message arriving as content) There simply was no way to poll a http connection (e.g. requests) if it had new data. It would block no matter what. So I would have to start writing threaded code, or just use asyncio which seemed much more ergonomic.