4 ms·
Is it true that these frameworks only work with async compatible libraries? Will a typical pypi lib work with them? For example, a css or xpath parsing library,
by eduction 3y ago
Is it true that these frameworks only work with async compatible libraries? Will a typical pypi lib work with them? For example, a css or xpath parsing library, or date time library - do I have to worry about those blocking and not working properly with asyncio or the others? Or only network and file reading libs?
I wish the article had spent more time on this. Without more info I would probably just use what the author mentions at the end as the Pareto ideal solution (ThreadPoolExecutor) because usually async frameworks in not historically async languages end up being islands within the larger community that need their own bespoke libraries.
- btreecat 3y agoYou don't use IO/make any external calls when you are parsing, or doing date time, so they don't have an opportunity to "cooperate" by letting another thread execute while they wait for the IO to finish.
- eduction 3y agoThanks. So even chewy compute operations (parsing a big html page then running a complicated xpath against it) are ok? What would be an example of blocking - is it just network io? Presumably file io too? Update: I think I’m conflating a bit what I want to speed up with what is allowed. Presumably heavy compute stuff is perfectly compatible with asyncio but it won’t speed it up - it would speed up a lot of io operations. ThreadPoolExecutor can speed up heavy compute by parallelizing (if it’s heavy enough) but may be overkill for just downloading 20 web pages at once.
- btreecat 3y agoIf you are parsing a big thing of text, there is no opportunity for your thread to pause the CPU and say "ok other threads, I have to wait on some external data" and then hand over the CPU resources. But if you are trying to make a network request, disk or memory access, or any HID stuff, you might have to wait on that thing to do it's job and get back to you. At this point you can tell your thread to raise it's hand and say "hey, I don't know how long, but I know I need to wait for this thing to finish, so while I wait someone else can use the CPU, but when I am done waiting, I will need the CPU back." This is the core idea of "cooperative threads" or green threads or co-routines or any sort of thread that is at the language runtime level and not at the OS level. https://en.wikipedia.org/wiki/Cooperative_multitasking https://en.wikipedia.org/wiki/Cooperative_multitasking So consider if the task you are doing is asking for data from "somewhere external to my thread" then you might be able to make that call non-blocking to enable more throughput. Hope that helps explain it!
- tomnipotent 3y ago> it would speed up a lot of io operations It doesn't speed up the I/O operation, but allows the CPU to continue executing other code rather than blocking and waiting (and the program becoming unresponsive). If your app is querying a database ten times per user request and averaging 5-10ms per query, your CPU is spending 50-100ms doing nothing but waiting on the network to finish so it can resume executing the code that comes next. That's time other code that has CPU instructions that can execute now would benefit from.
- rtpg 3y agoThe basic idea is that file IO and socket IO will allow a task to be suspended while the IO is happening, and let another task do its work. On top of that, you can write a C extension that will let you run compute heavy work in a way that allows your task to be suspended (for example, the C extension spins up its own thread). This extension has to be "very careful", basically by avoiding touching Python-side data during this work. The way this sort of stuff ends up working is you pass data into a C extension, and that extension takes ownership of the data or copies it or whatever, does what it needs, then gives Python back some result. But if you're just pure-Python compute heavy, then your task won't be suspended. So everything will run, but your compute-heavy stuff will hog the CPU, and won't be suspended. (Though if you have compute heavy work that is, like, looping over data, you could add `sleep(0)` between every couple of iterations. This gives other tasks a chance to run! This could be good enough to prevent weird bottlenecks). But the ultimate thing is if you have N compute-heavy tasks, you probably won't get speed advantages. If you have 1 compute-heavy task and N IO-heavy tasks, you can get advantages (even if the IO-heavy stuff is interspersed). But if you have N compute-heavy tasks and not much IO-heavy tasks, multiprocessing can get you where you want (since it's usually IO-heavy stuff that is helped out the most with async/await)
- evil-olive 3y agoCPU-heavy operations are allowed, strictly speaking, they're just "impolite" and may cause subtle problems. for example, say you have an aiohttp-backed web server, and in the handler for some URL route, you do a CPU-bound computation that takes 5 full seconds to execute. a crucial thing to remember is that asyncio programs, by default, are still single-threaded. so for those 5 seconds, your computation is the only thing running. other event loop tasks are unable to run. the loop is "blocked" in the same way it would be as if you called `time.sleep()` (not `await asyncio.sleep()`, but the regular synchronous sleep method) or made a blocking IO call (in a pathological case, you might open and read a file on a remote NFS share mounted over a WAN link, for example) now, suppose your web server exposes an `/alive` healthcheck endpoint. during that 5-second interval where you're hogging both the CPU and the event loop, aiohttp won't be able to dispatch requests to the healthcheck endpoint. if your load balancer has a 3-second timeout for those healthchecks, from its perspective the service will appear to be flapping between online and offline, even though the service itself never actually goes down. the "polite" thing to do, for this sort of big CPU-bound task, is to hand them off to a ThreadPoolExecutor like you said (or ProcessPoolExecutor if you have GIL concerns) with `loop.run_in_executor` [0]. that gives you a future, and when you `await` the future you are politely yielding the event loop to allow other tasks to run, such as those requests to the healthcheck endpoint. 0: https://docs.python.org/3/library/asyncio-eventloop.html#asyncio.loop.run_in_executor https://docs.python.org/3/library/asyncio-eventloop.html#asy...
- whalesalad 3y agoasyncio yes. gevent no. if you patch at an entrypoint you'll be ok 99 times out of 100.