20 ms·
A Heisenbug lurking in async Python
- deschutes 4y agoFun stuff. Why aren't unfinished tasks gc roots?
- sgt 4y agoIs this something go developers also have to be careful with when using goroutines?
- gerad 4y agoNo. But sometimes goroutines have the opposite problem, where they don’t terminate and get cleaned up. https://betterprogramming.pub/common-goroutine-leaks-that-you-should-avoid-fe12d12d6ee https://betterprogramming.pub/common-goroutine-leaks-that-yo...
- deleted 4y ago[deleted]
- candiddevmike 4y agoIs there an (easy?) test for checking goroutine leaks?
- Snawoot 4y agoYes, it's visible on goroutine profile, provided by built-in profiler pprof. E.g.: https://github.com/mysteriumnetwork/node/issues/5311#issuecomment-1193172089 https://github.com/mysteriumnetwork/node/issues/5311#issueco...
- Jtsummers 4y agoNo. Goroutines don’t generate a reference to hold onto, either. They just run until they or the program terminate.
- deleted 4y ago[deleted]
- acjohnson55 4y agoSomething linters can help with would think?
- ryanianian 4y agoC++ has nodiscard which is super useful for scenarios like this where ownership can be tricky.
- deleted 4y ago[deleted]
- aldenpage 4y agoThat's extremely insidious. I suppose I never encountered this issue because I almost always call asyncio.gather(*), which makes having a collection of tasks natural.
- kortex 4y agoThis is good form. It makes top-level control flow easier to follow, and keeps the concurrency scoped.
- rlpb 4y agoThis issue doesn't exist with Trio's structured concurrency model. In other words, the problem is already solved.
- nbadg 4y agoI'll +1 the Trio shoutout [1], but it's worth emphasizing that the core concept of Trio (nurseries) now exists in the stdlib in the form of task groups [2]. The article mentions this very briefly, but it's easy to miss, and I wouldn't describe it as a solution to this bug, anyways. Rather, it's more of a different way of writing multitasking code, which happens to make this class of bug impossible. [1] https://github.com/python-trio/trio https://github.com/python-trio/trio [2] https://docs.python.org/3/library/asyncio-task.html#task-groups https://docs.python.org/3/library/asyncio-task.html#task-gro...
- Tanjreeve 4y agoOh good so now we can all move to this years Async flavour in Python.
- edfletcher_t137 4y agoThis is a great blog post. Concise, lacking fluff or extraneous prose, it gets right to the point, presents the primary-source reference and then gets right to the solution. A bit of editorializing in the middle but that's completely allowed when writing this tightly. Well damn done, OP. And also it's great information that I - like I'm sure many of you - also never noticed. THANK YOU!
- isoprophlex 4y agoWell, I don't know, I kinda miss the human angle. I'd have loved to first read six paragraphs about how the author's grandmother raised them on home grown threads and greenlets :^)
- nickjj 4y ago> I'd have loved to first read six paragraphs about how the author's grandmother raised them on home grown threads and greenlets. With recipes, often times your problem is you want to learn how to make something where having the steps listed out is the most important thing. The story behind the recipe isn't important to solve your problem but for tech the story around the choice is important. Often times the "why" is really important and I really like hearing about what led someone to use something first. Often times that's more important or equally as important as the implementation details. It wouldn't make sense for this post given its title but if someone were making a post about why they chose to use async in Python I'd expect and hope that half of the post goes into the gory details of how they tried alternatives and what their shortcomings were for their specific use cases. That would help me as the reader generalize their post to my specific use cases and see if it applies.
- bialpio 4y agoOff-topic but the life story is there to make them eligible to be protected by copyright. IANAL. Source: https://copyrightalliance.org/are-recipes-cookbooks-protected-by-copyright/ https://copyrightalliance.org/are-recipes-cookbooks-protecte...
- samwillis 4y agoThis is one of many reasons I'm sceptical of the current trend in Python to "async all the things". The nuance to how it operates is often opaque to the developer, particularly those less experienced. GUI toolkits (like Textual) however are a really good use case for Asyncio. Human interaction with a program is inherently asynchronous, using async/await so that you can more cleanly specify your control flow is so much better than complicated callbacks. Using async/await in front end JS code for example is a delight. Where I'm particularly unconvinced of their use is in server side view and api end point processing. The majority of the time you have maybe a couple of IO opps that depend on each other. There is often little than can be parallelised (within a request) and so there are few performance gains to be a made. Traditional synchronous imperative code run with a multithreaded server is proven, scalable and much easier to debug. There are always places where it's useful though, things such as long running requests (websockets, long polling), or those very rare occurrences where you do have many easily parallelizable IO opps within one short request.
- michael_j_x 4y agoI am not sure I agree that the GUI is a good use case for async. A human interaction with the program must almost always pre-empt whatever the program was running, so I can not see how a cooperative multi-threading runtime like async Python can work in such a scenario.
- harpiaharpyja 4y agoWell it does work so "it can't work" isn't a very substantial criticism. Is there some nuance or detail that you meant to include?
- heavyset_go 4y ago> Where I'm particularly unconvinced of their use is in server side view and api end point processing. The majority of the time you have maybe a couple of IO opps that depend on each other. There is often little than can be parallelised (within a request) and so there are few performance gains to be a made. Traditional synchronous imperative code run with a multithreaded server is proven, scalable and much easier to debug. Traditional synchronous imperative code run with a multithreaded server is proven, scalable and much easier to debug. Python doesn't have multithreading that scales or supports real parallelism. asyncio has very measurable performance benefits for exactly that use case you've mentioned versus threaded servers.
- NelsonMinar 4y agoDoes anyone understand why the event loop only keeps weak references to tasks? It'd seem wise to do something to stop it from being garbage collected while running, maybe also while waiting to run.
- masklinn 4y agoOnly guess I’d have is to protect the system against infinite-loop tasks, but I don’t remember any other runtime caring and an a task which never terminates seems easier to diagnose than one which disappears on you.
- kortex 4y agoBecause it's almost always the case that the consumer is going to keep a reference to the task in some way, so that is the logical choice for the "primary owner" of the task. Python doesn't have ownership per se like rust, but if you keep more than one hard reference to an object around, it'll prevent collection, so in cases such as this it makes sense to designate one primary owner and have all other references be weakref.
- skitter 4y ago> if you keep more than one hard reference to an object around, it'll prevent collection Which is the behavior the parent comment asks for.
- kevin_thibedeau 4y agoThat creates a new problem that you have to remember to kill unwanted threads.
- deleted 4y ago[deleted]
- JonChesterfield 4y agoPython's reference counted - if the event loop holds a reference until the task has run, then drops it, then everything behaves sanely. That's not a cycle. It just means the task that was scheduled will execute, which seems like the right default.
- anthomtb 4y agoWell, looks like I know what I am doing first thing on Monday. I converted a bunch of code to asyncio a while back. I have yet to run into any heisenbug in that code and want to keep it that way.
- smetj 4y agoStart a thread/greenthread/fiber/process/task without holding a reference to at least tie all loose ends at exit? Hmm dunno.
- samsquire 4y agoThank you for this. This is really useful information. I recently adapted some garbage collection code to add register scanning. I can imagine all sorts of subtle bugs where things go away randomly. One problem I have with my multithreaded code is that sometimes a thread crashes and the logs are so long I don't notice. From my perspective the thread is just not doing anything. Sometimes the absence of behaviour can be really tricky to debug!
- kodablah 4y agoIt is for this reason in Temporal Python[0], where we wrote a custom durable asyncio event loop, that we maintain strong references to tasks that are created in workflows. This wouldn't be hard for other event loop implementations to do too. 0 - https://github.com/temporalio/sdk-python https://github.com/temporalio/sdk-python
- jarboot 4y agoIf I want to create a task that runs even after the function returns, ie "async def f(): asyncio.create_task(coro=10_second_coro.run()); return;" is there any way to mitigate this? Function-scoped set of tasks?
- jmholla 4y agoYour task is implicitly not function-scoped as you want it to survive exiting the function. What your doing here would be better architecturally done with threads. async is not a direct replacement for threading. But, you could also return the task object to the caller and have them manage it. There's also nothing async about your function, so you don't need the async or to await it.
- nhumrich 4y agoYes, read the last part of the included documentation and hold onto background tasks.
- nixpulvis 4y agoHey, at least it's documented... good developers actually RTFM. I can't comment on the design of this API, because I don't feel like learning the library, but in some performance critical applications these sorts of contracts aren't all that uncommon. Granted, this is python, I guess it's a bit more suspicious, IDK.
- vbernat 4y agoThe documentation update is quite recent (Python 3.11). It was added after this ticket: https://bugs.python.org/issue44665 https://bugs.python.org/issue44665 (not the first ticket around this problem).
- notatoad 4y agowow. yeah, this absolutely explains a heisenbug that i've been chasing for a while. and i can't count the number of times i've had that exact doc page open on my screen in the last few months, and never bothered to read that block of text that starts with "important"... thanks
- zzzeek 4y agoI think asyncio is kind of neat for what it's good at, but beginner programmers who have never wrote code before are going directly to using Python asyncio (i know this because they are telling me so when they post sqlalchemy discussions). This is just wrong.
- qxmat 4y agoPython has a few weird issues like this. The last one I encountered was with a class inheriting Thread, join and the SQL Server ODBC driver on Linux. Fairly sure I hit page faults thanks to a shallow copy on driver allocated string data but didn't have the time to investigate like the hero of this blog post.
- Lammy 4y agoI experienced a heisenbug exactly like this in Ruby when trying to `while case Ractor::receive`: https://github.com/okeeblow/DistorteD/blob/dd2a99285072982d32e69a2e63d204b0d4aff262/CHECKING%20YOU%20OUT/lib/checking-you-out/ghost_revival.rb#L380-L385 https://github.com/okeeblow/DistorteD/blob/dd2a99285072982d3...
- makomk 4y agoWell, this explains that one really annoying intermittent bug that I was having in some asyncio-based code.
- cmstodd 4y agoThanks for posting.
- dataflow 4y agoThe notion of fire-and-forget is itself the problem. Even with threads, you should have them join the main thread before the program exits. Which implies you should hold strong references to them until then. Most people don't go out of their way to do this even when they're able to, but that's what you're supposed to do.
- SamReidHughes 4y agoI came here to write this comment. Also, you usually need to have some means of canceling the task -- otherwise you have to wait for them to finish, or you leak these stray lost tasks that are doing stuff, like manipulating the state of things.
- sidlls 4y agoConveniences like this library and other threading libraries make it easy for people to trivialize something (concurrent programming) that ought not be.
- Too 4y agoThis. Even if you hold a reference to the task, your program very likely has a bug. At some point you should always await it to see if it failed or not. It's easy to miss this if you observe completion via a side-channel, for example item removed from a queue. But this is also a bad way to write tasks in the first place, let them return meaningful data rather than mutate shared objects. That way you are forced to await them and your code becomes much more straightforward. It's counter-intuitive at first if you think in threads, because there you are more used to worker pools and such, whereas asyncio tasks can be written in a more linear way and don't background workers to the same extent. After having done this mistake several times, I've concluded one should almost never use create_task. It's much better to place them into a top-level list of background tasks, that is always awaited, using this method they are both started automatically and always awaited for errors appropriately.
- dataflow 4y ago> At some point you should always await it to see if it failed or not. What you're sayin is correct, but doesn't quite imply what I'm saying. I'm saying everything that you spawn asynchronously (be they threads, tasks, whatever) needs to be joined - even if they're no-ops whose success or failure is irrelevant. This is similar to how you should always make sure to deallocate memory that you dynamically allocate whenever you can, as a matter of good practice and good hygiene. Sometimes you can get away with not doing so, but you shouldn't really skip it unless you don't have a choice, as it makes the program logic clearer and can make the program more robust too. (e.g., imagine running your main() in a loop where threads are spawned each time but never guaranteed to join.)
- bornfreddy 4y agoWow. What a strange design decision, as evidenced by sheer number of developers who don't / didn't know about this (myself included). I hope this gets fixed instead of just documented.
- jcheng 4y agoAgreed, I’m really surprised at all the comments defending this behavior. I suspect there is a non-obvious reason why it’s this way, but “you should’ve read the docs” and “but why wouldn’t you hold your own strong reference” are weird takes IMHO.
- hitekker 4y agoBugs stemming from the architecture of a poorly specified system become insecurities for the people who rely on that system. One of the major reasons why Python’s leadership refused to optimize python’s performance, besides Guido’s intransigence, was because they treated the CPython implementation as the specifications. Legions of script kiddies built their programming identities around the belief that python must be slow, because to admit otherwise would require changing the system.
- int_19h 4y agoIt was the Python community that consistently treated CPython behavior as the spec, in many cases contrary to explicit statements in the official docs to the effect that "this is a CPython-specific thing" etc, the most notable example being reliance on deterministic refcounting and on the GIL. But that's not what makes performance hard to fix. It's rather the fact that most Python code out in the wild depends on packages written in native code, and the CPython ABI for said packages exposes way too many implementation details. If you ditch ABI compatibility, you can ditch GIL, for starters (see e.g. IronPython and Jython) - but few people are willing to make do without all the affected packages.
- hitekker 4y agoBlaming the followers for not reading the docs is an easy excuse, when the real problem is that the leaders have failed to define Python. The PSF needs to decide what precisely Python is and what it isn't among themselves. Because they did not do so, they can't tell if something is a bug or intentional. See this entire thread as proof. In lieu of a formal definition, the maintainers resort to a hodgepodge of user docs, PEPs, mailing lists, and the "reference implementation". Worries about making that "reference implementation" more complex stymied Python's development. There's been flamewars about this topic with the PSF and their apologists. On that note, Rust encountered a similar problem of bad decision-making by maintainers, who also were opposed to specification. They have ejected those maintainers and replaced them with ones who understand the need for a formal definition of their language. https://blog.m-ou.se/rust-standard/ https://blog.m-ou.se/rust-standard/
- aeturnum 4y agoI really think this writer doth protest too much. Yes, the base async interface is confusing and overly complex. It's a downside! As they note lots of people have stepped in to provide better helpers (like TaskGroups) - but these are the docs for the base library! > But who reads all the docs? And who has perfect recall if they do? Everyone reads the docs? That is why you don't need perfect recall because you can read them whenever you want. Python has lots of confusing corner cases ("" is truthy, you need to remember to call copy [or maybe deepcopy!] sometimes, all the other situations where you confuse weak v.s. strong references). They cause really common bugs. It's just a hazard of the language in general and the choices it makes (much like tasks being objects is a hazard). I do understand why people think they can throw away task references (based on other languages) - but this is Python! The garbage collector exists and you gotta check if you own the object or something else does. Edit: this feels like an experienced Python developer, who has already internalized all the older, non-async Python weirdness, being taken aback by weirdness they didn't expect. Like, I feel you, it does suck - but it's not a bug that values you don't retain may get garbage collected.
- raverbashing 4y ago> "" is truthy Humm, no? Unless you mean ("",) >>> not "" True
- aeturnum 4y agoOh, sorry, you are right - "" is false-y, even though it's a valid empty value. So it's hard to tell the difference between a value not being filled and a value being filled with an empty value. ex: answers = {} answers["I exist"] = "" if answers["I exist"]: print("a") does not print.
- fbdab103 4y agoI guess I am too deeply in the Python ecosystem to see a problem here. Unless you want to check for the existence of "I exist"? In which case, the Python Way would be answers = {} answers["I exist"] = "" if "I exist" in answers: print("a")
- No1 4y agoHis argument hinges on "I can't be bothered to read the docs on the stuff I'm using." So instead of reading the docs on coroutines and tasks before using them, writes a rant about how it's all wrong because he didn't understand how it works. On a more fundamental level, why would anyone assume that a coroutine is guaranteed to complete if it is never awaited? There is no reason a scheduler could not be totally lazy and only execute the coroutine once awaited. At least he bothered to make note of TaskGroups, also clearly shown in his documentation screenshot, immediately above the section marked Important that went ignored, and finishes with "As long as all the tasks you spin up are in TaskGroups, you should be fine." Yep, that's all there was to it.
- zackees 4y ago[dead]
- ptx 4y ago> There is no reason a scheduler could not be totally lazy and only execute the coroutine once awaited. Isn't the point of create_task (which is what the article is about) to launch concurrent tasks without immediately awaiting them? The example in the docs [1] wouldn't work (in the stated manner) if the task didn't start until it was awaited. > At least he bothered to make note of TaskGroups [...] Yep, that's all there was to it. That only works on Python 3.11, which was released just a few months ago. Debian still uses 3.9, for example, so the TaskGroups solution can't be used everywhere yet. [1] https://docs.python.org/3/library/asyncio-task.html#coroutines https://docs.python.org/3/library/asyncio-task.html#coroutin...
- No1 4y agoThe reason I said "on a more fundamental level" is that I'm not talking specifically about Python and asyncio, but coroutines in general. Even for Python, there are multiple event loop libraries available, they do not all work identically, which is why multiple ones exist. Someone here mentioned Temporal Python which works differently from asyncio, and would have avoided the author's problem. If you don't know how the scheduler works, you can't assume that a coroutine is guaranteed to complete just because you yoloed it into the scheduler, no matter how convenient that might be for you. Yes, TaskGroups are a recent addition. If you can't use Python 3.11 for whatever reason, there is also the clearly written code sample at the bottom of the create_task documentation, which the author did not bother to mention. Probably didn't make it that far.
- m3047 4y agoHrmmmm. > But who reads all the docs? asyncio.create_task() doesn't exist in 3.6, and I can't find the string "to avoid a task disappearing" in the doc, so I'll go out on a limb: there is no such doc. However I see the reference to weakref.WeakSet.
- Jtsummers 4y agoThe world didn't end in 2016. Welcome to seven years in the future where this documentation does, in fact, exist: https://docs.python.org/3/library/asyncio-task.html#asyncio.create_task https://docs.python.org/3/library/asyncio-task.html#asyncio....
- m3047 4y agoSome of us have been writing python since 2.x, and quite unsurprisingly wrote asyncio code at 3.6, and still happily support it. Some of us have even asked on HN about maintaining compatibility backwards and forwards between 3.6 and 3.11. The documentation didn't exist at 3.6, when I wrote the code. I went and checked the source code, and the documentation and reported my findings. Good to know that there's a potential problem, don't you agree? What would you do differently?
- m3047 4y agoI've got to say, I've never actually noticed a problem with "fire and forget" although I use it for more or less disposable tasks to begin with. However, I've spent some more time looking through asyncio.base_events and * BaseEventLoop._scheduled is a list() * BaseEventLoop._ready is a deque() There is no change between 3.6 and 3.11 in this regard. So this could be a nothingburger if you don't use asyncgens. OTOH I suppose better safe than sorry; the only question is whether no code addressing it is more mentally taxing for the bystander than having code and trusting its implementation.
- cutler 4y agoMaybe grafting async onto a single threaded dynamic language just isn't such a good idea in the first place.
- murphy214 4y agobingo
- hn_throwaway_99 4y agoWorks absolutely fine in JavaScript, and the language is certainly much better for it than it was before async/await.
- cutler 4y agoJavascript was born with async built-in. The async/await in JS is just syntactic sugar around callbacks/promises.
- whoopdeepoo 4y ago> But who reads all the docs Why is this so common? Do people seriously not read a language/library documentation? That's the absolute first thing I do when evaluating a technology.
- adamckay 4y agoBecause people have deadlines and need to get things working. You read enough to figure out how to do what you need to do and then mostly move on. This function was added in 3.7 with no note on the importance of saving a reference. In 3.9 a note was added "Save a reference to the result of this function, to avoid a task disappearing mid execution." which was then expanded with the explanation of a weak reference in 3.10.
- throwaway81523 4y agoI had a manager who actually told me not to read docs. I was a bad report and read them anyway.
- skitter 4y agoIt absolutely is common. People see there is a len function that takes one argument, they call len(some_collection), see that it indeed returns the number of items in the collection like they expect and move on. They don't expect len to return a negative number instead on Thursdays, and of course it doesn't because that would be a pretty big footgun. People also see that there is a create_task function that takes a coroutine, they call create_task(some_coroutine), see that the coroutine indeed runs like they expect, and move on. Sure, you're supposed to await the result, but maybe they don't need the awaited value anymore, only the side effects, and see that it still works.
- boomskats 4y agoAs someone who happens to be eternally grateful to the author for his contribution to the Python ecosystem [0], I kinda feel like this comment thread is overreacting to his overreaction. When I look at this post all I see is a useful, well explained, byte-size writeup that a search engine might recommend to someone looking for help in writing async Python. Maybe it's because a bunch of my friends are Scottish and I get their sense of humour. [0]: https://rich.readthedocs.io/ https://rich.readthedocs.io/ (yes I'm talking about the fancy new progress bar that pip got recently)
- vore 4y agoTo me, the surprise here is that usually you don’t expect Python finalizers to do something like this: when they dispose something, it’s usually unobservable from the perspective of the program, e.g. an unreachable file descriptor. Here, the runtime is disposing something that is still observably in progress, which is surprising behavior.
- deleted 4y ago[deleted]
- quietbritishjim 4y agoI do wish he dwelt on task groups a bit more at the end. Many comments here seemed too have missed that bit. They're not just a handy way of executing a hack. Instead, they're a revolutionary way (ok maybe that's a but string but not much) to structure your async program to avoid a whole host of bugs. A code snippet would have been nice, or a link to the blog post that introduced them (in trio, another async library): https://vorpus.org/blog/notes-on-structured-concurrency-or-go-statement-considered-harmful/ https://vorpus.org/blog/notes-on-structured-concurrency-or-g...
- quietbritishjim 4y ago*bit strong
- ravloony 4y ago
- bandyaboot 4y agoHe doesn’t really get into what makes this a Heisenbug, only that it’s indeterminate in nature. Would attaching a debugger/stepping through the code make it less likely that your task would get garbage collected out from under you?
- foobarbecue 4y agoYeah, he seems to be re-defining the term to mean "a bug that occurs occasionally depending on system state" as opposed to "a bug that changes behavior when you observe it closely e.g. in a debugger."
- macintux 4y agoThe first is a common way of using the term Heisenbug. I first heard it used that way 10 years ago when discussing Erlang’s error handling model.
- foobarbecue 4y agoTIL. I guess I assumed it would hew more closely to the Uncertainty Principle. Edit: actually, come to think of it, I first heard of it in about 2006 from Jamie Brandon and at the time assumed it was something he'd made up. For a second there I forgot that 2006 is more than 10 years ago! (It was a python bug that went away when run in a debugger.)
- throwaway81523 4y agoCPython does most of its memory management by reference counting, which fails to reclaim circular structure. So to make sure it gets everything, it occasionally runs a conventional tracing GC. If the GC happens to run just after you create that async task, the task itself can get collected, it sounds like. It's good to know about this and is (my own editorializing) yet another reason Python3 should have used Erlang-style concurrency instead of this async stuff.
- Izkata 4y agoYou're probably going to need a reference to the task in order to inspect it in the debugger. Creating that reference prevents the bug.
- BiteCode_dev 4y agoAnd this is why trio got it right, and why I think the task groups (nurseries from trio) can't arrive soon enough in the stdlib. Because not only you must maintain a reference to any task, but you should also explicitly await it somewhere, using something like asyncio.wait() or asyncio.gather(). Most people don't know this, and it makes asyncio very difficult to use for them.
- postultimate 4y ago> task groups (nurseries from trio) can't arrive soon enough in the stdlib Please, no. Asyncio is horrible, and bodging it to make it less horrible just means we will be forced to live with the remaining horror. Far better to replace it with something that works properly (yes, Trio).
- throwaway81523 4y agoThere's a similar thing in tkinter but I guess users discover it faster, since the failure if you don't save the reference shows up fairly quickly.
- dehrmann 4y agoEww. What's especially nasty is this is the opposite behavior of threads.
- dehrmann 4y agoAnother common async footgun I see is unthrottled gathering, and no throttling mechanism in the standard library. Once you gather an unspecified number of awaitables, bad things start to happen, either with CPU starvation, local IO starvation, or hammering an external service. What I like about threads is they make dangerous things like this harder, and you have to put more thought into how much concurrent work you want outstanding. They also handle CPU starvation better for things that are latency-sensitive. I've seen degenerate requests tie up the event loop with 500 ms of processing time.
- rednafi 4y agoHuh! Unless you're using semaphores, you can also recreate similar situation with threads. Spin up a whole bunch of threads and send all of them towards some shared object or make 100s of requests with them. There's not much difference between spinning up threads explicitly and creating async task with asyncio.create_task. In either case, you can throttle them with semaphores.
- dehrmann 4y agoI don't have a source or affected versions, but semaphores can scale poorly. I vaguely remember each blocked acquire getting checked on every event loop iteration, or something silly like that.
- osigurdson 4y agoA little pedantic but HUP concerns the fundamental limits of simultaneously knowing a particle's position and momentum, not about observation impacting outcomes.
- cpburns2009 4y agoI've been working on a PySide6 application recently using asyncio. I read the docs but totally overlooked the requirement to hold references to tasks created with `create_task()`.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- winter_blue 4y agoThis article just makes me feel like Python, while a language with nice-ish syntax, is a language that was poorly hacked and put together with little concern/thought about the real-world implications of poor design decisions like this async design decision (and also dynamic typing – a terrible thing in any language).
- photochemsyn 4y ago'async footguns' returns 20,000+ hits on Google. Top one happens to be: https://news.ycombinator.com/item?id=32086973 https://news.ycombinator.com/item?id=32086973 > "Async seems to be the first big "footgun" of Rust. It's widespread enough that you can't really avoid interacting with it, yet it's bad enough that it makes..."
- crdrost 4y agoMost languages have something like this, usually around async. For instance NodeJS has had a bit of this around promises, and eventually needed to institute the rule “if a promise rejects with an error, anf nobody is around to hear it, we will crash your program on the assumption that you probably needed to clean up some resources but didn't and now they're going to leak. Listen to the error with a handler that does nothing, if we are wrong about that.”
- macintux 4y agoOne of many reasons I like Erlang: everything is async, so you have plenty of tooling/libraries/core language features to support you.
- philwelch 4y agoPython 2.7 was a nice little language. Python 3 is a bit of an overwrought monster at this point.
- deleted 4y ago[deleted]
- deleted 4y ago[deleted]
- crabbone 4y agoIn many years since asyncio has been added, I have never used it willingly, outside of the cases where a third-party library required it. There has never been a practical benefit for any of that stuff when compared to select. It always worked poorly and never justified the effort one has to put into writing code that uses the library. The behavior OP describes is just one of the many bad design decisions that are so characteristic of this library.
- aardvark179 4y agoThe same problem or something similar exists in many languages. Threads are GC roots because the OS knows about them, but this may not be true for lightweight threads or async callbacks. It is hard to fix because you don’t want to introduce references from an old object (such as a list of callbacks) to many new objects as that will introduce GC issues, and many other potential leaks.
- pyuser583 4y agoI don’t find this behavior odd at all. Dereferencing unassigned values is normal Python garbage collector behavior. Threads are an exception (no pun intended), but they’re an exception in lots of ways - just try pickling them.
- mhils 4y agoThe note in the official Python documentation was only added in September 2022 [1], so no wonder this comes as a surprise to many! [1] https://github.com/python/cpython/commit/6281affee6423296893b509cd78dc563ca58b196 https://github.com/python/cpython/commit/6281affee6423296893...
- terom 4y agoThat commit is adding the same disclaimer for the `shield` function, the original one-line mention for `create_task()` is a little older, but also only November 2021, so during 3.11 development...? [1] https://github.com/python/cpython/commit/c750adbe6990ee8239b27d5f4591283a06bfe862 https://github.com/python/cpython/commit/c750adbe6990ee8239b...
- perlgeek 4y agoThank you! I just did a quick `git grep` in a work code base and found one clear instance of this bug, and two more locations where I'm not 100% certain whether references are kept around long enough. Made a note to open a bug on Monday :-) Another surprise in python's base library: I knew that re.search searches for a regex match in a string, so I thought that re.match would match the whole string. I was wrong, re.match only anchors the regex at the start, not the end. re.fullmatch anchors it at both sides. I felt very stupid when I found out; I started at my current work as a Perl developer and learned Python for a new job; but there are two more Python developers (with previous Python experience outside this company) on the same project, and none of them noticed the mistakes I made based on this misunderstanding.
- slewis 4y agoThe response here makes me think most commenters don’t have experience with this particular footgun. To clarify: Python can gc your task before it starts, or during its execution if you don’t hold a ref to it yourself. I can’t think of any scenario in which you’d want that behavior, and it is very surprising as a user. Python should hold a ref to tasks until they are complete to prevent this. This also “feels right”, in that if I submit a task to a loop, I’d think the loop would have a ref to the task! It’d be interesting to dig up the internal dev discussions to see why the “fix this” camp hasn’t won.
- kevincox 4y agoI can see this behaviour being useful if you are no longer interested in the result of a "pure" task. For example imagine fetching some data via HTTP. If you no longer need to response canceling the request could make sense. But I agree that this is unexpected and most code probably isn't ready for being cancelled at random points. (Although I guess in Python your code should be exception safe which is often similar)
- account42 4y agoIf you want to avoid an expensive operation in cases where the result is no longer needed then surely you'd want to cancel the task as soon as you know that and not at some indeterminate time when the GC runs so I don't think this behavior makes sense even for that scenario.
- hgomersall 4y agoThat would make tidying up rogue tasks impossible. Of course we all like to think we do cancellation perfectly, but it's nice to know that the task scheduler has your back. Edit: I don't quite understand why a user would expect a task to remain live _after_ the last reference to it has been dropped...
- int_19h 4y agoBecause they have expressed the intent to run it by scheduling it on the event loop? I don't follow the argument wrt tidying up rogue tasks. What does it mean for the task to be "rogue"? If there was some state change that made the task redundant - because it clearly wasn't when it was submitted! - then the code that makes that change, or some other code that observes it, should cancel it. If it isn't cancelled, the fact that nobody is able to observe the value that the task will yield is not sufficient to auto-cancel, as there may still be a dependency on side effects. And, speaking of tidying up, what if the scheduled task is the one that performs some kind of cleanup?
- sidlls 4y agoA better title would be “Bug lurks in incorrect usage of async Python”. The library documentation clearly calls this out, and incorrect implementations, while buggy, do not mean that async Python is itself buggy.
- deleted 4y ago[deleted]
- kyrofa 4y agoGreat post. Feels like something a linter should catch.
- deleted 4y ago[deleted]
- epakai 4y agoI missed this in a little curses program launcher I wrote. It looks obvious when he puts a big orange box around it, but in the actual docs it's an unassuming paragraph between two border-wrapped blocks with the only distinguishing feature being the bold "Important". It should probably be referenced immediately next to the "Return the Task object" sentence.
- syngrog66 4y agoJava had a similar but inverse problem in early versions. A counter-intuitive behavior that bit people and caused leaks. If you instantiated a Thread, and then start() was never called, that thread object would leak. And thus potentially an entire graph of objects, via the references chains beneath it. Obviously a thread that is never started seems pointless by design. But it could happen easily if, for an example, an error happened or an exception was thrown at some point between the instantiation line and the call of start(). The root cause was because Sun's programmers had made the early implementations of Thread get added to a ThreadGroup by default, under the hood. What would happen is that ThreadGroup stayed alive/reachable and thus it kept your app's thread object reachable too, and thus the GC would never clean it up. It was never eligible. It ended up being the cause of a few weird leaks we saw in production. IIRC in Java 1.4 or 1.5 Sun fixed it by ensuring the thread got cleaned up in those cases.
- fsckboy 4y agoThe oop model (or any "client" model of state) dependent on state outside your code being encapsulated/contained inside an object within your program is always confusing. It's not particular to this library or python. A gui window handle within your program, or simply an open file handle, if the OS does something to your object, it's hard for you to know about it, and your object continually needs to refresh its state if you are concerned about it. I don't know what is referred to as a "task" in this case, but I don't think the lifetime of the actual task is the issue, it's the lifetime or your object. It's always the case that if you instantiate an object with a ctor, you can't count on anything about it continuing if the dtor is invoked. The problem is much more general than this API, this library, this language, this use case. Just as you need to structure your C code so malloc and free will always match up spatially and temporally, you need to structure your oop code so ctors and dtors match up sensibly. Otherwise the confusion in your head will spread to your code. And those who always want the compiler and tools to automatically do as much as possible to free them from the burden, are the ones who are most surprised by Heisenbergs. Computers can (or will try to) do anything, it's up to you to make sure it does what it needs to. Maybe someday AIs will do it better than us, but right now you need to provide the I.
- postultimate 4y agoThis one seems quite easy to fix - just have the scheduler check that task objects have a positive refcount before running them, including the first time.
- remram 4y agoA positive refcount doesn't mean the object is not dead.
- postultimate 4y agoBut the lack of a positive refcount means that it is, so this solves the problem the article was complaining about.
- remram 4y agoIt works unless it doesn't, so it stays a Heisenbug.
- magicalhippo 4y agoDelphi had the opposite bug in its thread pool. The worker threads would dequeue a work item and process it in a loop. The work items were reference counted. Now, Delphi doesn't have scoped variable declarations like say C++ or C#, so the dequeued work item was stored in a local variable. However, it didn't drop (nil/null) the work item reference before it looped. Thus it would hang on to that reference until the next work item got dequeued or the pool was destroyed. The result was that if you in a function started a task which captured a local reference (f.ex. using an anonymous function) and then waited for it, that reference could live after your function returned if the pool didn't have anything else to do. Not what most people would expect.
- collinvandyck76 4y agoIt seems like the library should retain a handle on the task until it completes.
- nerdponx 4y agoThis is partly the goal of structured concurrency, implemented as Nursery in Trio and TaskGroup in Asyncio.
- tbrownaw 4y agoSo a `Task` object is in fact the actual task, rather than a handle or pointer to the actual task.
- jasonlaster11 4y agoUnfortunately we don't support python yet, but if you ever encounter a similar heisenbug in JS, you can try recording it with Replay.io
- phyzome 4y agoThat seems very odd for a default behavior. There might be good reasons to allow GCing of a dropped task reference, but it doesn't seem like that would be the most common case.
- 29athrowaway 4y agoReading this was like the first time I read "for/else". "wtf?!" was my first reaction, then I read the documentation.
- mkarliner 4y agoThank you!
- blacklight 4y agoI've used asyncio.create_task forever, and I've admittedly never read its documentation in depth. However, I've ALWAYS assigned the return value of create_task to some variable. To me this is just a good programming practice. The OP says "tasks are not like threads - that you can just launch and forget" - no! Even a thread should not simply be launched and forgotten! You always need a reference to it, so when your application is terminated you can join all the threads that are still running and things can exit in a clean way! Same goes for the tasks: when your application exits, it's just a good practice to stop the even loop and cancel any pending tasks. And, in general, one should always keep in mind the reference count rule in Python (the author incorrectly calls it "garbage collection" btw): if something you created in your function isn't referenced/assigned to anything, then its reference count will be zero, and it will be removed when the function stack unwinds. This is totally expected behaviour to me.
- qwertox 4y agoI'm using a lot of `asyncio.get_event_loop().create_task(...)` calls without assigning the task to a variable, but the docs [0] don't mention anything regarding this method on the loop object. Can I assume that I'm safe? [0] https://docs.python.org/3/library/asyncio-eventloop.html?highlight=create_task#asyncio.loop.create_task https://docs.python.org/3/library/asyncio-eventloop.html?hig...
- quietbritishjim 4y agoI'm pretty sure that's equivalent to asyncio.create_task(), it just gives you the opportunity to specify a different loop if needed. I think the docs are just less explicit because they're aimed more at power users (since most people don't deal with multiple event loops).
- hhhhhhkkk 4y ago[flagged]
- terom 4y agoThere's a GitHub issue to fix this, arguing that the doc fixes are insufficient, consider it to be a design mistake, and argue that it can be changed without breaking backwards compatibility (current GC behavior is not deterministic): https://github.com/python/cpython/issues/91887 https://github.com/python/cpython/issues/91887 Latest reply from GvR is invoking Chesterton's Fence. Here's to hoping the devs can quickly figure that one out and get this fixed. Per the linked issues [1] even the stdlib asyncio implementation was affected. [1] https://github.com/python/cpython/issues/90467 https://github.com/python/cpython/issues/90467