7 ms·
Making 1M requests with Python-aiohttp
- ben_jones 10y agoDoes anyone enjoy doing async work in python? I've done a few hobby projects and honestly I was yearning for javascript + async lib after awhile. As great as python is maybe we should yield async programming to the languages designed for it?
- JustSomeNobody 10y agoI guess I don't know how JS was any more "designed for" async than python was.
- henryw 10y agoFrom https://developer.mozilla.org/en-US/docs/Web/JavaScript/EventLoop https://developer.mozilla.org/en-US/docs/Web/JavaScript/Even...: JavaScript has a concurrency model based on an "event loop". This model is quite different than the model in other languages like C or Java. ... A very interesting property of the event loop model is that JavaScript, unlike a lot of other languages, never blocks. Handling I/O is typically performed via events and callbacks, so when the application is waiting for an IndexedDB query to return or an XHR request to return, it can still process other things like user input.
- tyingq 10y ago>>JavaScript, unlike a lot of other languages, never blocks I think that's a little strong. It's more like "The group controlling Javascript has mostly tried to discourage introduction of things that block". You can, for example, do a blocking XMLHttpRequest. It's deprecated, but possible. https://jsfiddle.net/923d5sda/ https://jsfiddle.net/923d5sda/
- coldtea 10y agoYou can do blocking everything. A 10.000 repetitions for look while block the whole interpreter for its duration. Any JSON parsing does the same. Processing strings. Doing math work. ...
- tyingq 10y agoGuess I should have said "blocking I/O"? I thought it was a given that a single threaded language wouldn't magically inject some kind of concurrency around tight loops or CPU intensive tasks. Node.js people don't really think that sort of thing is "non blocking", do they?
- coldtea 10y ago>Guess I should have said "blocking I/O"? I thought it was a given that a single threaded language wouldn't magically inject some kind of concurrency around tight loops or CPU intensive tasks. Well, Erlang (and Elixir) does just that -- it's preemptive, and implicitly yields under the covers even in loops. >Node.js people don't really think that sort of thing is "non blocking", do they? Judging from forum threads and blog posts, a lot of them do, especially web programmers not familiar with blocking and non-blocking that only know that "Node is webscale".
- Matthias247 10y agoI think that quoted statements are not very good. The other languages don't have a builtin concurrency model. For C, Java and others event loop libraries and applications that are built on top of them can be found (nginx, netty, ...). As well as there are libraries that build on top of synchronous IO. The eventloop was also not tied to Javascript in the older standards. Only the introduction of Promises and other stuff required the existence of an eventloop in order to define when continuations should run. "The event loop model never blocks" is also only true as long as you (and all the libraries that you use) do not block it. There is no automatic "does not block" guarantee.
- coldtea 10y ago>A very interesting property of the event loop model is that JavaScript, unlike a lot of other languages, never blocks. JavaScript actually always blocks. It's only external function calls that somebody took care to write in an evented style that don't block -- but anything written in pure Javascript (from for loops to text manipulation) blocks. JS single-thread async without preemptiveness is not something to write home about...
- darpa_escapee 10y agoThis is a genuine question: in what ways is Python's async implementation lacking? Could it have been baked in a better way? In what ways do languages that were supposedly designed for async programming different than Python? Python is definitely lacking an elegant interface for async programming.
- bdarnell 10y agoI think that Python 3.5 now has a very elegant interface for async programming. I prefer Tornado to the standard library's asyncio, but the new keywords are nice for both packages (disclaimer: I'm the maintainer of Tornado). The downsides have nothing to do with the design of the language. The problem is that introducing a new concurrency model late in a language's life splits the ecosystem. Most existing packages are synchronous, so if you want to build asynchronous systems you must avoid packages like requests, django, or sqlalchemy and find (or develop) asynchronous equivalents for the functionality you need. Javascript has an advantage here not because the design of the language is especially well-suited for asynchronous programming, but because it never went through a synchronous/multi-threading phase. Every javascript package is designed for asynchronous use.
- nbadg 10y agoYeah, my biggest complaint personally is that combining multithreading and async is a massive pain in the ass. Now, realistically, you aren't usually going to want to do that, except if you have multiple event loops, or are bridging between external synchronous code and internal async code. Otherwise, I really enjoy async python -- of course, I'm also the kind of person who has written my own event loops using synchronous code before, so maybe I'm just crazy like that.
- markbnj 10y ago>> The problem is that introducing a new concurrency model late in a language's life splits the ecosystem. Really you could say that about a number of features/changes in python 3, not just the new async syntax. Python 3 itself was an ecosystem-splitting instrument.
- 10y ago
- afarrell 10y agoI found it easier than doing it in javascript because I could insert `import pdb;pdb.set_trace()` into the code and get an interactive debugger. Supposedly you can do this in javascript by running node with a particular flag, then connecting to a port on localhost, and opening the chrome debugger. However, the multiple times I've tried throughout 2014-2016 has show that to be incredibly finnicky. It is especially frustrating when trying to insert a debugger into an automated test.
- rgacote 10y agoI think async and aiohttp are game changes for Python 3.5. After working with Twisted callbacks for over a decade, it's a pleasure to write async code that does not use the callback approach (granted Twisted is a mature environment with lots to offer). I've switched to Python3.5 and aiohttp for all new web service applications. The coding style is clean, enjoyable to write, and easy to debug. Plus, I've never once been stymied for speed. I know there's applications out there where people expect to to be handling zillions of connections -- but the bulk of my use cases think 100 transactions per second is a huge through-put, and aiohttp handles that with ease.
- bjt 10y agoHave you used eventlet or gevent? I thought they were game changers. Gevent has been working very well for me for quite a while, without callback hell.
- coldtea 10y agoJavascript was hardly designed for async work -- it is just that it wasn't designed for anything else.
- ojii 10y agoaiohttp has completely replaced flask for my "small web apps/web apis" needs. For my personal performance needs, just running `python myapp.py` is enough, no need for gunicorn or other "complicated" setups.
- henryw 10y agoLooks pretty interesting to do async on python. I once did something similar in node (async by default) with a few lines of code. I think I scraped 12 or 20 million real URLs in 8 hours on a $5 cloud VM. It was limited by network bandwidth.
- imaginenore 10y ago1,000,000 requests in 52 minutes is just 320 req/sec. Am I missing something? What's so amazing about this? I just deployed some production feed that serves at 1955 requests/second on a cheap VPS in freaking PHP, one of the slowest languages out there.
- aaossa 10y agoWhy you say is not amazing? Honestly curious here :)
- imaginenore 10y agoBecause it's trivial. I would be interested in anything doing 10,000+ req/sec on a cheap VPS. 320 is nothing. People achieve 2 million requests/second with C++ on EC2: https://medium.com/swlh/starting-a-tech-startup-with-c-6b5d5856e6de#.38xs3etwg https://medium.com/swlh/starting-a-tech-startup-with-c-6b5d5...
- aaossa 10y agoOh I see now... This speaks by it self: C++/Proxygen =1,990,130 requests per second Python/Tornado = 41,329 requests per second Thanks for sharing btw
- Cyph0n 10y agoAbsolutely excellent article. I always keep C++ at the back of my mind in case I need it some day, so I think the list of libraries they used will be useful for me in the future. Thanks.
- gst 10y agoIf you want to write fast C++ Web services I recommend a look at Seastar: http://www.seastar-project.org/ http://www.seastar-project.org/
- jorge_leria 10y agoThe article is not about serving, but about consuming. Not the same beast.
- nbadg 10y agoFirst off, awesome to see more benchmarks (even if it's just personal experimentation) for synchronous vs asyncio performance. I think the real argument for asyncio right now is that it makes it very easy for you to write extremely efficient code, even for hobbyist projects. Even though your experiment is only handling 320 req/s, that you were able to do that so quickly and with very, very little optimization is, I think, a testament to the potential for asyncio. Some pointers: The event loop is still a single thread and therefore subject to the GIL. That means that at any given time, only one coroutine is running in the loop. This is important for several reasons, but probably the most relevant are that 1. within any given coroutine, execution flow will always be consistent between yield/await statements. 2. synchronous calls within coroutines will block the entire event loop. 3. most of asyncio was not written with thread safety in mind That second one is really important. When you're doing file access, eg where you're doing "with open('frank.html', 'rb')", that's something you may want to consider moving into a run_in_executor call. That will block the coroutine, but it will return control to the event loop, allowing other connections to proceed. Also, more likely than not, the too many open files error is a result of you opening frank.html, not of sockets. I haven't run your code with asyncio in debug mode[1] to verify that, but that would be my intuition. You would probably handle more requests if you changed that -- I would do the file access in a run_in_executor with a max executor workers of 1000. If you want to surpass that, use a process pool instead of a threadpool, and you should be ready to go, though it's worth mentioning that disk IO is hardly ever cpu-bound, so I wouldn't expect you to get much performance boost otherwise. Also, the placement of your semaphore acquisition doesn't make any sense to me. I would create a dedicated coroutine like this: async def bounded_fetch(sem): async with sem: return (await fetch(url.format(i))) and modify the parent function like this: for i in range(r): task = asyncio.ensure_future(bounded_fetch(sem)) tasks.append(task) That being said, it also doesn't make any sense to me to have the semaphore in the client code, since the error is in the server code. [1] https://docs.python.org/3/library/asyncio-dev.html#debug-mode-of-asyncio https://docs.python.org/3/library/asyncio-dev.html#debug-mod...
- dante9999 10y agoThanks for feedback. > You would probably handle more requests if you changed that -- I would do the file access in a run_in_executor with a max executor workers of 1000. This is really good point. I'm going to check this and edit post adding this information there. > Also, the placement of your semaphore acquisition doesn't make any sense to me. I would create a dedicated coroutine like this: looking into my semaphore code next day after writing it I do wonder if I'm using it correctly. I assumed it works correctly because it fixed my "too many open files" exception, so it seems to mean that I'm no longer exceeding 1024 open files limits. Can you clarify why you think my use of semaphore does not make sense and why your suggestion is better? What is the benefit of dedicated coroutine? > That being said, it also doesn't make any sense to me to have the semaphore in the client code, since the error is in the server code. I admit that I focused more on my client than server. One thing that worries me about my test server is that it does not print any exceptions. Either it does not fail at all, which seems unlikely, or it fails silently, which is more likely and is bad. So I need to check my server code to see what exactly happens there. > it also doesn't make any sense to me to have the semaphore in the client code, since the error is in the server code. main reason for semaphore in client code is that it should stop client from making over 1k connections at a time. My logic here is that if client wont make 1k connections at a time - server wont receive 1k connections at a time and thus there will be no problem of too many open files on server (it won't have to send more than 1k responses). However I see that this logic may not be totally correct, other comment points out that it's possible for sockets to "hang around" after closing: https://news.ycombinator.com/item?id=11557672 https://news.ycombinator.com/item?id=11557672 so I need to review that and edit post. > https://docs.python.org/3/library/asyncio-dev.html#debug-mode-of-asyncio https://docs.python.org/3/library/asyncio-dev.html#debug-mod... this looks really great, will look into this thanks.
- sandGorgon 10y agoI really keep wishing that there would be benchmark comparisons of asyncio/aiohttp with gevent/python2 . Performance would be a killer reason to migrate immediately to Py3. What I suspect though is that asyncio is not all that better than gevent. Can someone correct me on this?
- riyadparvez 10y agoIs there anything inherent to Python3 that is slower than Python2? Or is it just some of the performant packages still have not been ported to Python3?
- sandGorgon 10y agoi keep looking for a reason to switch to python 3 and cant find one. Plus if I want to use the cool stuff in Pypy.. then I better not ! overall - very less reason to consider Py3 at all. Performance would have been one - if there were a comparison between gevent and asyncio.
- mrweasel 10y ago>i keep looking for a reason to switch to python 3 and cant find one. Unicode? Not having to deal with encoding all over the place has been well worth switch to Python 3. If performance is a a huge issue, I honestly don't know why you would stay on Python (regardless of version) I wouldn't want to switch back to Python 2.7 is I can avoid it. There's honestly no reason not to go with 3.4 or 3.5 at this point, unless you happen to have a large Python 2 code base.
- sago 10y agoThis is a superficially trivial bit of syntactic sugar, but an example of the way small tweaks can provide big impact. This: > do_something(*some_args, *some_more_args) is rocking my world right at the moment. That's a massive time saving feature I've been waiting for and worth the price of a 3.5 upgrade.
- tooker 10y agoI have a library for doing coordinated async IO in python that addresses some of the scheduling and resource contention issues hinted out in the later part of this post. It's called cellulario in reference to containing async IO mechanics inside a cell wall.. https://github.com/mayfield/cellulario And an example of using it to manage a multi-tiered scheme where a first layer of IO requests seeds another layer and then you finally reduce all the responses.. https://github.com/mayfield/ecmcli/blob/master/ecmcli/api.py#L456
- simonw 10y agoThis looks really promising. I've often wanted to be able to do exactly this: run a bunch of async code in the middle of an otherwise synchronous block (classic example: writing a Django view which fires off a bunch of parallel HTTP API requests and continues once all of them have either returned or timed out).
- tooker 10y agoThat's almost exactly the use case I began with. It unapologetically requires python 3.5+ but if you're already there I'd be happy to see and support some of your use cases. Hit me up on github if you want to try it and need some guidance (the docs are nonexistent).
- ulyssesv 10y agoCould you please elaborate on why async would be preferred than a task queue solution (would it)?
- deleted 10y ago[deleted]
- azinman2 10y ago"Everyone knows that asynchronous code performs better when applied to network operations" Ummm that seems a bit far reaching.
- 15155 10y agoIt depends on what "network operations" you are trying to do. For high-concurrency purposes, asynchronous programming is far more scalable (see: epoll/kqueue + state machines). For high-throughput, low-concurrency operations, it doesn't matter as much.
- azinman2 10y agoI happen to know of a very major tech company who scale is insane yet their core c++ code is based on highly tuned blocking threads. It's not a given that async is the only way to scale.
- velox_io 10y agoThe 1 million in the title is misleading (1M per hour is nothing to write home about, only 278/sec). There are frameworks that are able hit 1M per minute plus (16,666/sec).
- jorge_leria 10y ago1M per minute it is something. Could you name those frameworks?
- jc4p 10y agoElixir is the name I see thrown around the most when it comes to stuff like this: http://www.phoenixframework.org/blog/the-road-to-2-million-websocket-connections http://www.phoenixframework.org/blog/the-road-to-2-million-w...
- dpc_pw 10y ago1 million per hour is nothing... Here, mioco handling 10M http request per second(1) on my desktop: https://github.com/dpc/mioco/blob/master/BENCHMARKS.md https://github.com/dpc/mioco/blob/master/BENCHMARKS.md 1) with a bit of cheating http server. With actual proper http parsing it goes down to 368K req/s, but that's still a lot.
- pbz 10y agoEven 1M per minute is rather pathetic. The game is around multiple millions per second: https://www.techempower.com/benchmarks/#section=data-r12&hw=peak&test=plaintext https://www.techempower.com/benchmarks/#section=data-r12&hw=...
- terom 10y agoRe the EADDRNOTAVAIL from socket.connect(), If you're connecting to 127.0.0.1:8080, then each connection from 127.0.0.1 is going to be assigned an ephemeral TCP source port. There are only a finite number of such ports available, on the order of ~30-50k, which limits the number of connections from a single address to a specific endpoint. If you're doing 100k TCP connections with 1k concurrent conections, it's feasible that you'll run into those limits, with TCP connections hanging around in some TIME_WAIT state after close(). Not that this would be a documented errno for connect(), but it's the interpretation that makes sense.. http://www.toptip.ca/2010/02/linux-eaddrnotavail-address-not.html http://www.toptip.ca/2010/02/linux-eaddrnotavail-address-not... http://lxr.free-electrons.com/source/net/ipv4/inet_hashtables.c?v=4.4#L572 http://lxr.free-electrons.com/source/net/ipv4/inet_hashtable...
- ahuang 10y agoGenerally its the upper 32k ports that are ephemeral, and if your churn more than that per minute in connections, you'll run into that TIME_WAIT issue. Hacky way to get around that is to enable tcp_tw_reuse which will let you reuse ports, but it can be risky if you get a SYN from the previous connection that happens to lineup with segment number of the current connection (which will close your connection). Shouldn't happen often, and if you can tolerate a small amount of failure is an easy way to get around this limit. [0] http://blog.davidvassallo.me/2010/07/13/time_wait-and-port-reuse/ http://blog.davidvassallo.me/2010/07/13/time_wait-and-port-r...
- e12e 10y agoFor benchmarking loopback connections, addressing really shouldn't be an issue, as you have an entire /8-subnet to split between your client(s) and server(s) (127.0.0.0/8). You would need some logic to set up eg 10.000 listening servers, and 1000.000 clients to get it working, and at some point you'd probably run into memory or other limits. I'm a little surprised some simple googling didn't turn up any examples of this - I'm sure someone have tried it out in order to do some benchmarking of high-performance network servers/services? Apparently ipv6 changes this to a single (loopback) address, but then again, with ipv6 you can use entire subnets per network card.
- takeda 10y agoIMO you should place all requests within a single ClientSession(). This will provide two benefits: 1. You won't need to use a semaphore. To limit connections you will need to create a TCPConnection() object with limit set to the limit you used in the semaphore and pass it to the ClientSession() and aiohttp will not make more connections than the limit set (default behavior is to have unlimited number of connections). 2. With single ClientSession(), aiohttp will make use of keep-alive (i.e. it will reuse same connections for next requests, but it will keep at most the limit of connections you set in TCPConnection() object). This should improve performance further, and (given sane limit) it'll also solve issue with "Cannot assign requested address" error. BTW: Even without limit set aiohttp will try to reduce number of connections open so it might still fix the connection error issue as long as individual requests don't take long. It's still good idea to set limit, just to be nice to the remote server.
- philippb 10y agoI'm the CTO at KeepSafe. We open sourced aiohttp. We wrote aiohttp for our production system. We build everything on aiohttp. In our production systems we constantly run more request then in the benchmark with business logic on each request. The main reason we like aiohttp a lot if that you we can write asynchronous code that reads like synchronous and does not have callbacks.