14 ms·
We have to talk about this Python, Gunicorn, Gevent thing
- nemoniac 7y agoSo..., What's a good alternative? One that's relatively straightforward to implement compared to the Python approach.
- jerf 7y agoIn this very particular sense, almost anything else is better. Dynamic scripting languages that are intrinsically single threaded because they were single-threaded for the first 10-15 years of their lives, and it is virtually impossible to retrofit true threading all the way from their basic runtimes through all their libraries [1], are basically the pessimal case for this particular problem. This is not the whole story of the value of those languages. As the article even alludes to, at small loads or with lots of care this can be made to "work". But it is something that an engineer should know about them before picking them up and using a tool for something it isn't really good at. [1]: I add this caveat because I don't think there's anything about dynamic scripting languages that makes them intrinsically difficult to thread any moreso than any other category of language, it's just that by an accident of history, they all come to us from the 1990s personal computer world, and they all spent at least a decade cooking and setting and building libraries and communities and developer skillsets before a serious need for threading was even on the horizon.
- samatman 7y agoIt's a good caveat, because Lua, in particular, has fully-reentrant functions. You can run a bunch of Lua_states cooperatively or on a threaded basis without problems. Everything the VM does, from the C side, receives a Lua_state as the first argument. It's intrinsically single-threaded, yes. But each instance is quite small and they stay out of each others way. Add coroutines and there's a lot you can safely do with Lua that's a real pain to accomplish with Python.
- mhd 7y agoIn the days of yore, I might've attempted to shoe-horn Tcl in such a situation. Decent event-loop for distributing tasks and (like Lua, but unlike Python/Ruby) more eager to escape to C for performance-sensitive tasks.
- oconnor663 7y ago> I don't think there's anything about dynamic scripting languages that makes them intrinsically difficult to thread any moreso than any other category of language I think there might be some intrinsic factors: 1. Languages like Python don't want to expose the program to undefined behavior. Defining some trivial class, and then manipulating it from two threads at the same time, is not supposed to crash the program or introduce horrible security issues. 2. Languages like Python have "property bag" objects. That means (someone who knows better should check me on this) that most writes in a typical program are hash table operations, rather than primitive stores of an int or whatever. Locking each table separately, or using a fully atomic implementation, can be a significant slowdown in single-threaded programs, compared to using a GIL.
- jerf 7y agoI think if you wrote a new dynamic scripting language from scratch with the intention that it be threaded that you could probably come up with something. Python is blocked from it not because it is impossible, but because they've been unwilling to orphan all their C extensions. The problem is that while it may raise hackles when stated so bluntly, dynamic scripting languages are generally on their way out anyhow (although they have a ways to go before that is even generally recognized, and even longer before they're legacy; we're talking plural decades for the whole process here), and also, fighting really hard to get threading into a language that is also intrinsically slow is just not the sort of compelling end-point that would inspire an author to create it, and a community to back it up. Why would you want a "threaded scripting language" that literally takes a 32-core machine to catch up to what numerous languages would be capable of doing on a single core? (A new language will have a long ways to go before it can match a good JS VM. Look at Perl 6/Raku's performance history.) Especially since such a language would be racing things like Nim, Crystal, a D revitalization, and several other contenders who are getting 90% of the convenience of scripting languages while getting 90% (or more) of the performance of compiled ones.
- pdonis 7y ago> What's a good alternative? You don't necessarily need one; it depends on what kind of application you have and what its bottleneck is. If your application is network I/O bound, event-driven asynchronous I/O, which is what this article is describing, works fine, and a single process/thread even in a dynamic language like Python can handle a large volume of requests. The specific issue this article is describing is due to a particular poor implementation of event-driven asynchronous I/O, not a general problem with the entire concept. If your application is CPU bound, then yes, you need to use threads (or multiple processes), and you shouldn't be trying to mix event-driven asynchronous I/O with that.
- jimbokun 7y agoAnecdotally, seems like a lot of people are using Go, or maybe Elixir to keep the dynamic typed development experience but much more efficient hardware utilization.
- j88439h84 7y agoAsync/await
- orf 7y ago> Go back and look. I said that it forks and then it imports your app. Your app (which is almost certainly the bulk of the code inside the Python interpreter) is not actually in memory when the fork happens. It gets loaded AFTER that point. You can just pass `--preload` to have gunicorn load the application once. If you're using a standard framework like Django or Flask and not doing anything obviously insane then this works really well and without much effort. Yeah I'm sure some dumb libraries do some dumb things, but that's on them, and you for using those libraries. Same as any language. If you want to stick your nose up at Python and state outright "I will not write a service in it" then that's up to you, it just comes across as your loss rather than a damning condemnation of the language and it's ecosystem from Rachel By The Bay, an all-knowing and experienced higher power. I guess everyone else will keep quickly shipping value to customers with it while you worry about five processes waking up from a system call at once or an extra 150mb of memory usage.
- ary 7y ago> You can just pass `--preload` to have gunicorn load the application once. If you're using a standard framework like Django or Flask and not doing anything obviously insane then this works really well and without much effort. Yeah I'm sure some dumb libraries do some dumb things, but that's on them, and you for using those libraries. It's not always trivial to ensure none of your dependencies have import-time side effects. Sometimes the productivity/business benefit provided by the depedendency outweighs the pain introduced by the side effects.
- deleted 7y ago[deleted]
- Skunkleton 7y agoIf spinning up a few more workers will solve a performance problem for you, it’s probably worth the time to throw the preload flag on there and see what it does to your test suite. Since you are already cost optimizing at this point you probably have the time.
- kodablah 7y ago
- pdonis 7y agoThe problem being described here isn't Python. gunicorn, or gevent; it's bad programming. I'd be willing to bet there are systems out there written in C++, Java, and Ruby that do the same dumb things. The solution is to not do dumb things--to understand what your program is doing. It's perfectly possible to do that in Python, gunicorn, and gevent.
- _old_dude_ 7y agoIn the case of Java, the Selector API was introduced in Java 4 (2002) for this exact reason, avoid to have all the threads to all waits/being notified on accept().
- j88439h84 7y agoUsing an ASGI server that supports async/await, such as Uvicorn, instead of green threads, forking, etc, seems like a good idea these days. Also means you can use Starlette which has a much nicer design IMO than some of the old frameworks. - https://www.uvicorn.org/ https://www.uvicorn.org/ - https://www.starlette.io/ https://www.starlette.io/
- rcarmo 7y agoI'm digging https://github.com/RobertoPrevato/BlackSheep https://github.com/RobertoPrevato/BlackSheep myself, since I prefer the terseness of Bottle and stacking decorators to add functionality to handlers. Any of the above (or Sanic) can do ~3K RPS on a single core on a Raspberry Pi (which is where I test things for portability, optimisation and a little fun), and the RAM overhead is generally not that bad small (just did a little "hello world" uvicorn/blacksheep app and I see 22MB resident/10MB shared per worker, and one of my Clojure servers taking up over four times that...)
- seemslegit 7y ago"I will not use web requests when the situation calls for RPCs" I'm surprised how often devs treat this distinction as architecturally meaningful. Web requests are just RPCs with some of the parameters standardized and multiple surfaces for parameters and return values - query string, headers, body. This is completely orthogonal to the strategy used to schedule IO, concurrency, etc.
- jchw 7y agoI mean... you do have to parse text formats. HTTP parsing may be a solved problem, but that doesn't mean the overhead or complexity of doing so disappears. Also, TLS is not really ideally lightweight for RPCs, but you should absolutely encrypt your RPC traffic (imo.) So I really think the whole stack is out. (P.S.: If you are wondering what kinds of 'lightweight' replacements for TLS exist, I think my personal favorite attempt is CurveCP, although it is a bit dated nowadays. I wouldn't often recommend people roll their own, but you could certainly do something simple with NaCl/libsodium directly. Maybe QUIC also fits the bill?)
- seemslegit 7y agoThere is nothing that says that "RPC" can't do multiple request/response cycles over an existing open (and encrypted) connection rather than initiate a new one for every call, just like HTTP. Or even pipeline them like HTTP/2.0
- jchw 7y agoHTTP 2 is more reasonable. But by the time you get to HTTP 3 you're just doing HTTP 2 over QUIC. At which point, why not just send RPC payloads directly over QUIC?
- seemslegit 7y agoBecause what I have in place is a good old HTTP 1.1 ?
- yowlingcat 7y agoI like the author's articles most of the time. While this article contains some truths, I don't think it argues very persuasively for its conclusion. Okay, these parts of the Python ecosystem don't work well together, and it's a bad, unpolished experience. Fair, as with other criticisms of Python. The question, however, is why one would use gevent at this point in Python's evolution. There's async await now, and things like FastAPI. If you want to use, say, the Django ecosystem, use Nginx and uWSGI and be done with it. Maybe you need to spend some more resources to deploy your Python. Okay. Is that a problem? Why are you using Python? Is it because it's quick to use and helps you solve problems faster with its gigantic, mature ecosystem that lets you focus on your business logic? Then this, while admittedly not great, is going to be a rounding error. Is it because you began using it in the aforementioned case and now you're boxed into an expensive corner and you need to figure out how to scale parts of your presumably useful production architecture serving a Very Useful Application? Maybe you need to start splitting up your architecture into separate services, so that you can use Python for the things that it does well and use some other technology for the parts that aren't I/O bound and could benefit from that. But that's not this article is about. This article is about someone making the wrong choices when better choices existed and then making a categorical decision against using Python for a service. I'd say that's what "we have to talk about" if you ask me.
- divbzero 7y agoAgreed on both (a) I usually like the author’s articles and (b) think she’s missing the point on this one. gevent and gunicorn were good attempts to remedy a bad situation. async/await is the solution that the Python community is coalescing around. Even with Django, there are active efforts to support ASGI. [1] [1]: https://docs.djangoproject.com/en/3.0/howto/deployment/asgi/uvicorn/ https://docs.djangoproject.com/en/3.0/howto/deployment/asgi/...
- ghostwriter 7y agoGevent was doing it right and async syntax was a huge mistake that fractioned community-contributed libraries into two incompatible camps with lots of unnecessary cloning happening at present moment. In high-level languages with virtual machines and/or garbage collectors, the runtime system should be solely responsible for scheduling green threads around IO entry points, all without special syntactic markers. GHC has it right (https://www.aosabook.org/en/posa/warp.html https://www.aosabook.org/en/posa/warp.html), Gevent was a right development with on-par async performance metrics (https://gist.github.com/rfyiamcool/41d4004b7fd46516d0b4f34f6065e93b https://gist.github.com/rfyiamcool/41d4004b7fd46516d0b4f34f6...), that had a standard synchronous coding style. It could be adopted into the core language and improved further without splitting the community.
- tus88 7y agoGotta love those people who fail to understand how things are supposed to be used, fail miserably as a result, then throw the baby out with the bathwater in a fit of tantrum. Yes, Python has a GIL. Yes, lightweight threads are mostly good for IO bound tasks. Yes it can still be used effectively if you design your app correctly.
- orisho 7y agoThe problem that's described here - "green" threads being CPU bound for too long and causing other requests to time out is one that is common to anything that uses an event loop and is not unique to gevent. node.js also suffers from this. Rachel says that a thread is only ever doing one thing at a time - it is handling one request, not many. But that's only true when you do CPU bound work. There is no way to write code using blocking IO-style code without using some form of event loop (gevent, async/await). You cannot spin up 100K native threads to handle 100K requests that are IO bound (which is very common in a microservice architecture, since requests will very quickly block on requests to other services). Or well, you can, but the native thread context switch overhead is very quickly going to grind the machine to a halt as you grow. I'm a big fan of gevent, and while it does have these shortcomings - they are there because it's all on top of Python, a language which started out with the classic async model (native threads), rather than this model. Golang, on the other hand, doesn't suffer from them as it was designed from the get-go with this threading model in mind. So it allows you to write blocking style code and get the benefits of an event loop (you never have to think about whether you need to await this operation). And on the other hand, goroutines can be preempted if they spend too long doing CPU work, just like normal threads.
- notyourday 7y ago> The problem that's described here - "green" threads being CPU bound for too long and causing other requests to time out is one that is common to anything that uses an event loop and is not unique to gevent. node.js also suffers from this. Isn't node.js IO native threaded?
- orisho 7y agoI don't know, but I was actually discussing time when the thread is performing pure CPU operations and not doing any IO. In this case the JS code does not yield to the event loop, so any other requests waiting will stall. This is actually an important property of JavaScript that allows it to work without locks (as long as you do not await, your code runs completely synchronously and will not be interrupted). It's not a flaw, but definitely requires awareness. Gevent code that doesn't spawn more than one native thread (the one that runs the event loop) also has this property - you don't need locks as long as you do not perform IO. In Python's case it can be more tricky, as you might end up yielding to the event loop by indirectly performing IO when logging or something similar. In JavaScript, the only case, AFAIK, where this can happen is when you await. Nothing else will cause you to yield to the event loop.
- ary 7y agoThis is spot on. My one and only gripe is with this part: > So how do you keep this kind of monster running? First, you make sure you never allow it to use too much of the CPU, because empirically, it'll mean that you're getting distracted too much and are timing out some requests while chasing down others. You set your system to "elastically scale up" at some pitiful utilization level, like 25-30% of the entire machine. Letting a Python web service, written in your framework of choice, perform CPU-bound work is just bad design. A Python web service should essentially be router for data, controlling authentication/authorization, I/O formatting, and not much else. CPU intensive tasks should be submitted to a worker queue and handled out of process. Since this is Python we don't have the luxury of using threads to perform CPU-bound work (because of the Global Interpreter Lock).
- ghostwriter 7y ago> Since this is Python we don't have the luxury of using threads to perform CPU-bound work (because of the Global Interpreter Lock). You can have it with threads if the CPU-bound work is done inside a C extension - https://docs.python.org/3/c-api/init.html#releasing-the-gil-from-extension-code https://docs.python.org/3/c-api/init.html#releasing-the-gil-...
- MoronInAHurry 7y agoRachel's posts would be so much more useful if she would just say what she meant, instead of twisting everything into knots to find a way to say it backwards so she can be sarcastic and condescending while doing it. I'm sure there's some useful information in here, but it's not worth digging through the patronization to find it.
- gorgoiler 7y agoIt’s an interesting post to read. Have another go if you can, but I very much agree with you on the tone issue. Imagine if these were ones own notes that had to be read through the next time something like this happened. A more succinct operational — dare I say: positive! — way of writing would really be welcome.
- deleted 7y ago[deleted]
- yobert 7y agoI would counter-argue that dry, positive, informational writing is great for Wikipedia but can also be very boring. This blog has a lot of snark and that's what makes digesting the great information so much fun!
- _frkl 7y agoTotally agree, I also enjoy her posts quite a bit!
- spectramax 7y agoSo much fun and so little substance. Fun should be sprinkled here and there with a healthy 95% dose of substance. Everything Rachel writes is a convoluted mess that’s impossible to follow.
- spectramax 7y agoFor the downvoters - I also do not like Paul Graham and Sam Altman - they're the same as Rachel in every way. Little substance, lots of unsubstantiated filler material. To extend this further, I also don't like NewYorker for this reason alone - I don't have time for convoluted novel-like stories that has the important bit buried somewhere in the middle of 6 pages. If I want to read beautiful and creative prose, I need to be in that mindset. Not when discussing Python innards.
- mesozoic 7y agoAs much as I love Python I still tell people don't use it for performance sensitive applications.
- pdonis 7y agoIt depends on what kind of performance you need. For CPU intensive tasks I would agree with you. But for network I/O intensive tasks, even though Python is slow it's still more than fast enough to keep up with a large request volume since network I/O latency is so much longer than CPU/memory latency.
- dmurray 7y agoAnd for data science /ML things it works great these days, even though those applications are 100% performance-bound.
- airstrike 7y agoAt this point, I imagine there's likely a similar law to Betteridge's that states: Any headline that starts with "we have to talk about" can be answered by the words "do we?"
- DevKoala 7y agoIn this crap situation atm, can attest. Currently maintaining a Python app for the delicate snowflakes whose years of math understanding somehow prevents them from being able to learn a language that isn’t Python. We have money, let’s just blow it. /s
- tgbugs 7y agoThis is a great review of what is going on "behind the scenes." As the maintainer of about 5 little services with this structure I have vowed never to write another one. The memory overhead alone is a source of eternal irritation ("Surely there must be a better way...."). Echoing other commenters here, the real cost isn't actually discussed. Namely that there is a solution to some of these problems (re long running tasks?), but it carries with it a major increase in complexity. Its name is Celery and oh boy have fun with the ops overhead that that is going to induce. A while back I did some unscientific benchmarking of the various worker classes for python3.6 and pypy3 (7.0 at the time I think?). Quoting my summary notes: 1. "pypy3 with sync worker has roughly the same performance, gevent is monstrously slow gthread is about 20 rps slower than sync (1s over 1k requests), sync can get up to ~150rps" 2. "pypy3 clearly faster with tornado than anything running 3.6" 3. "pypy3 is also about 4x faster when dumping nt straight from the database, peaking at about 80MBps to disk on the same computer while python3.6 hits ~20MBps" I won't mention the workload because it was the same for both implementations and would only confuse the point, which is that there are better solutions out there in python land if you are stuck with one of these systems. One thing I would love to hear from others is how other runtimes do this in a sane and performant way. What is the better solution left implicit in this post?
- coleifer 7y agoWith python the only thing that matters is the workload.
- viraptor 7y ago> It's around this time that you discover that people have been doing naughty, nasty things, like causing work to occur at "import time". Is this something people actually have problems with in practice? I did lots of python and ran into it once. It was quickly fixed after a raised issue. I feel like non-toy development just doesn't experience it. But maybe that's my environment bubble only. Do people who do serious python development actually have problem with this?
- tgbugs 7y agoPython pretends to be a nice homogeneous "everything is at run time" language, but it is all a big lie and there aren't big flashing letters saying "you really probably shouldn't do this" when you start solving a problem in a certain way. For example, it is almost certainly best practice to _never_ call a function, class method, or static method inside a module that is going to be imported, and certainly never instantiate a class. However, there are certain patterns that almost necessitate it if you don't want to write loads of boiler plate or deal with the performance overhead of metaclasses. There are also a bunch of nice hacks like using `object()` at the top level as an instance distinct from everything else, but I'm sure there is a way that `MYTYPE = object()` will come back to absolutely ruin your day if you have to compare two `MYTYPE` instances in two different dicts derived from a parent process and a subprocess. I have personally made this mistake on two or three occasions where I conflated file/module behavior with class behavior because I wanted a python file to act like it was a bit more declarative. Unfortunately this leads to a world of eternal pain. You can work around it, but you should have made everything a python class and pretended like the files/modules don't exist or at least have staggeringly different semantics hiding behind that innocent little `.` operator. Python simply cannot support the desire to solve a problem in a certain way because of the structure of the problem and forces you into using its happy path patterns if you want it to work in slightly different run time contexts. Two relevant posts from instagram engineering on the subject which suggest that best practices for avoiding these kinds of issues are non-obvious and easy to miss. https://news.ycombinator.com/item?id=20708889 https://news.ycombinator.com/item?id=20708889 https://news.ycombinator.com/item?id=21284669 https://news.ycombinator.com/item?id=21284669
- 7y ago
- nodamage 7y agoYes after reading through the article it's not very clear to me what the actual problem is with using Python/Gunicorn/Gevent. The author seems to be saying something about how if a worker is busy doing CPU intensive work (is decoding JSON really that intensive?) then other requests accepted by that worker have to wait for that work to complete before they can respond, and the client might timeout while waiting? If that's the case: 1. Wouldn't this affect any language/framework that uses a cooperative concurrency model, including node.js and ASP.NET or even Python's async/await based frameworks? How is this problem specific to Python/Gunicorn/Gevent? 2. What would be a better alternative? The author says something about using actual OS-level threads but I thought the whole point of green threads was that they are cheaper than thread switching?
- zzzeek 7y ago> (is decoding JSON really that intensive?) in Python, everything is generally CPU intensive compared to what it would be in compiled languages, even though things like JSON decoding are usually happening in a C library, Python programs that do close to nothing still use way more CPU than you would if you were running in the JVM, or Go, C, whatever. > Wouldn't this affect any language/framework that uses a cooperative concurrency model, including node.js and ASP.NET or even Python's async/await based frameworks? How is this problem specific to Python/Gunicorn/Gevent? CPU bound-ness affects all of these platforms, yes. It affects Python and other intepreted languages the most however because these platforms get the most CPU-bound the most quickly. Also, applications that are written in scripting languages tend to have a lot of business logic going on in the first place; after all, if you just wanted to serve static pages you could use Apache with the Event NPM and if you wanted to proxy HTTP requests you'd use HAProxy; both event-based systems that are very much not CPU bound. But yes, most importantly, Python's asyncio system is completely impacted by these same issues and I would have preferred she address that, as asyncio is part of the standard library now and is way more popular than gevent. > What would be a better alternative? The author says something about using actual OS-level threads but I thought the whole point of green threads was that they are cheaper than thread switching? I will grant she lost me a bit with the "use a real RPC system with <feature> <feature> <feature>" thing, and additionally the "load the application in the child process" thing is pretty typical, a worker process should obviously have either threads or greenthreads in use so that each process can handle multiple concurrent requests, but only as many as you'd want handled effectively by one core since the GIL is going to enforce that (another thing you wouldn't have to deal with in other languages such as the above mentioned compiled languages), but it's typical that child processes are going to have a mostly original copy of things. But as far as the "context switching" thing, I've yet to see benchmarks that show the overhead of OS-level context switching actually being more of a performance burden than the less frequent, but more work intensive context switching that user-space schemes like asyncio have to use. If you are writing a logic-heavy, or even a logic-just-a-bit service that receives Python requests you will also have to worry about CPU-bound issues all the time. Using regular threads with processes, like what you get using something like mod_wsgi, will allow individual processes to attend to web requests more evenly. With mod_wsgi you can configure worker daemons that run multiple OS level threads and you can also have multiple daemon processes. I'm not sure if the multi-process model used by mod_wsgi has solved the accept() problem, however in my experience the bigger problem is when a service configures itself to allow for 1000 greenlets within each process, while each process is realistically capable from a CPU perspective of handling maybe 5 or 10 concurrent requests, there's no mechanism that ensures that each process gets an even balance of requests. That is, you might have all your requests waiting in one process, because you told them it can process 1000 at a time, while other processes are idle. TL;DR I'm in the "event based programming is extremely overrated in Python" camp.
- ris 7y agoI don't disagree with any of this but > "Why in the hell would you fork then load, instead of load then fork?" In python it often seems to make little difference. The continual refcount incrementing and decrementing sooner or later touches most everything and causes the copy to happen whether you're mutating an object or not. I've had some broad thoughts about how one would give cpython the ability to "turn off" gc and refcounting for some "forever" objects which you know you're never going to want to free, but it wouldn't be pretty as it would require segregating these objects into their own arenas to prevent neighbour writes dirtying the whole page anyway...
- wrmsr 7y agoThey took a step towards this with https://docs.python.org/3/library/gc.html#gc.freeze https://docs.python.org/3/library/gc.html#gc.freeze but it doesn't go as far as disabling refcount touching outright. I've experimented with doing that, both per-object and just globally, and the results really were promising if your forkserver can keep up with providing the necessarily much shorter-lived worker processes.
- ris 7y agoThanks for this link - I had completely missed it (I think I was just expecting to disable gc entirely or perform some rudimentary surgery on its linked list)
- jks 7y agoThere's a 2017 writeup from Instagram: https://instagram-engineering.com/dismissing-python-garbage-collection-at-instagram-4dca40b29172 https://instagram-engineering.com/dismissing-python-garbage-...
- ris 7y agoThis isn't quite the same thing, but it is one of the articles that spurred my thoughts on this subject. In cpython, gc != refcounting. Instagram were talking about disabling gc, which would have stopped objects which they weren't using from being falsely copied, but wouldn't have stopped objects that they were using (but not mutating) from being copied.
- diebeforei485 7y agoI find this post to be unintelligible. Given that it's been upvoted to the top of HN though, can someone TL;DR of the intellectual value of this post? It seems to be stepping through the details of what is going on while also being rambling.
- doctoboggan 7y agoI recently started playing around with Google Cloud Run and am running some python/flask/gunicorn code in a docker container on the platform. I noticed in the logs that I am getting a lot of Critical Worker Timeouts and I am wondering if this has anything to do with it.
- countbayes 7y agoWe had that problem and hacked around it with the Dockerfile instructions below, if you find a better solution that would be great :) --Dockerfile snippet-- # Cloud Run concurrency is assumed to be set to 10 but we don't assume that is exact # See 'https://github.com/benoitc/gunicorn/issues/1801' https://github.com/benoitc/gunicorn/issues/1801' so disabling concurrency ENV GUNICORN_CMD_ARGS="-c gunicorn_config.py --workers 1 --threads 1 --timeout 120 --preload" CMD [ "gunicorn", "pkg.http:app" ] # Or just use Flask directly if concurrency is set to 1 #CMD [ "python", "cmd/server/main.py" ]
- doctoboggan 7y agoThanks for the pointer. I was messing around with --preload and --timeout flags and they seemed to work, although I think that isn't fixing the root problem.
- benreesman 7y agoIt seems to me that this submission is getting a lot of blowback in the comments for 1) the style and 2) the implication that wiring up Python services with HTTP is bad engineering. I don’t think this is productive. On the first point, yeah Rachel’s posts are kinda snarky sometimes, but some of us find that entertaining particularly when they are highly detailed and thoroughly researched. I’ve worked with Rachel and she’s among the best “deep-dive” userspace-to-network driver problem solvers around. She knows her shit and we’re lucky she takes the time to put hard-earned lessons on the net for others to benefit from. As for “microservices written in Python trading a bunch of sloppy JSON around via HTTP” is bad engineering: it is bad engineering, sometimes the flavor of the month is rancid (CORBA, multiple implementation inheritance, XSLT, I could go on). Introducing network boundaries where function calls would work is a bad idea, as anyone who’s dealt seriously with distributed systems for a living knows. JSON-over-HTTP for RPC is lazy, inefficient in machine time and engineering effort, and trivially obsolete in a world where Protocol Buffers/gRPC or Thrift and their ilk are so mature. Now none of this is to say you should rewrite your system if it’s built that way, legacy stuff is a thing. But Rachel wrote a detailed piece on why you are asking for trouble if you build new stuff like this and people are, in my humble opinion, shooting the messenger.
- ahuang 7y agoI think the main issue is it seems really one-sided and the intent was to be snarky, vs educational. I posted a comment here detailing some ways to work around some of the pitfalls. I think if she devoted more time in the article to solutions vs. complaining, her points would come across more productively.
- diebeforei485 7y ago> we’re lucky she takes the time to put hard-earned lessons on the net for others to benefit from. I genuinely don't see much of a lesson to learn from this particular blogpost, and it appears neither did many others in HN. If there is one, beyond "don't use x", it's hard to find it. I get the impression that this particular post is being upvoted to the top of HN because of who the author is, not necessarily because this post itself has value. This results in a whole bunch of others reading it, wondering why they're wasting their time with such a rambling post.
- cwp 7y agoSigh. Yes. I have been there and done that (more or less) and it sucks. The root problem is that data scientists really want to use Python for machine learning, but wrapping a Python model in a service that uses CPU and memory efficiently is really difficult. Because of the GIL, you can't make predictions at the same time you're processing network IO, which means that you need multiple processes to respond to clients quickly and keep the CPU busy. But models use a lot of memory and so you can't run all THAT many processes. I actually did get the load-then-fork, copy-on-write thing to work, but Python's garbage collections cause things to get moved around in memory and triggers copying and makes the processes gradually consume more and more memory as the model becomes less and less shared. Ok, so then you can terminate and re-fork the processes periodically, and avoid OOM errors, but there's still a lot of memory overhead and CPU usage is pretty low even when there are lots of clients waiting and... You know I hear Julia is pretty mature these days and hey didn't Google release this nifty C++ library for ML and notebooks aren't THAT much easier. Between the GIL and the complete insanity that is python packaging, I think it's actually the worst possible language to use for ML.
- crimsonalucard 7y agoShe's talking about green threads which is different from regular threading in python. Under nodejs/python style green threads only IO calls are concurrent to a single computation task. There is no parallelism under both styles of threading unless you count concurrent IO as parallel. She is basically complaining about a pattern that was popularized by NodeJS and emulated in python by older libraries like gevent, twisted and tornado. Currently python3 uses keywords async/await as an API around the same concepts implemented in the older libraries. This has nothing to do with GIL.
- cwp 7y agoIn the case of the article, you are correct. I have a slightly different case where I'm wrapping scikit-learn model. We're NOT just calling another service and waiting for a response, we're doing computation, in Python. So the GIL is actually a problem.
- ghostwriter 7y ago
- crimsonalucard 7y agoThere's a huge amount of technical jargon and sarcasm that makes it hard to see her point. Basically she's saying that python async (which the current state of the art implementation uses libuv the same thing driving nodejs and consequently suffers from the same "problems") doesn't have actual concurrency. Computations block and concurrency only happens under a very specific case: IO. One computation can happen at a time with several IO calls in flight and context switching can only happen when an IO call in the computation occurs. She fails to see why this is good: Python async and nodejs do not need concurrency primitives like locks. You cannot have a deadlock happen under this model period. (note I'm not talking about python threading, I'm talking about async/await) This pattern was designed for simple pipeline programming for webapps where the webapp just does some minor translations and authentication then offloads the actual processing to an external computation engine (usually known as a database). This is where the real processing meat happens but most programmers just deal with this stuff through an API (usually called SQL). It's good to not have to deal with locks, mutexes, deadlocks and race conditions in the webapp. This is a huge benefit in terms of managing complexity which she completely discounts.
- eropple 7y ago> Python async and nodejs do not need concurrency primitives like locks. You cannot have a deadlock happen under this model period. This is dangerously wrong and I would suggest that you reconsider the steps that got you to this understanding because something really important has been lost. It is absolutely critical to understand that deadlocks are not why you have locks. Correctness during concurrent operation is why you have locks. Deadlocks are a failure state when you do not have correctness during concurrent operation. So are things like double-increment and double-create. Parallelism does not imply deadlocking, concurrency implies deadlocking, and both NodeJS and Python are concurrent runtime environments. And I can guarantee you that, with a little skull sweat, you can write a deadlock in NodeJS or Python. It is very easy. If you need some help, here's a trivial example (and ordinarily I wouldn't use a link shortener here but this one is hefty, it just goes to the Typescript playground): https://bit.ly/2Tvjyze https://bit.ly/2Tvjyze Also, as a concrete, real-world, yes-it-happens-here example of where locking is important, consider that I've recently built a dependency injection framework in NodeJS--tried to use others' first, but my situation isn't covered by existing ones--and had to resort to a mutex to avoid double creation of objects within a single lifecycle. Creation of objects within this lifecycle happens asynchronously--it has to, as the act of creating the objects might itself rely on asynchronous operations. So, if I were to have a diamond-dependency (A deps B and C, B and C dep D), I will non-deterministically, and based on the creation time of B and C, create either one or two instances of D. I rely upon a mutex, keyed upon the dependency being created, to ensure that this does not happen. . I would also submit that perhaps you should adopt a principle of charity and think real hard about whether your priors are correct before you start talking about what she "fails" to see. Rachel is one of those people who has Been Around and while I also have Been Around, I understand that Rachel has Been Around More and I probably should be listening more than I should be smarming at her. Just a thought.
- ahuang 7y agoI think this conflates a poor implementation of a webserver with python/gunicorn/gevent being bad. There are a few (easy) things to do to avoid some of the pitfalls she encountered: > A connection arrives on the socket. Linux runs a pass down the list of listeners doing the epoll thing -- all of them! -- and tells every single one of them that something's waiting out there. They each wake up, one after another, a few nanoseconds apart. Linux is known to have poor fairness with multiple processes listening to the same socket. For most setups that require forking a process, you run a local loadbalancer on box, whether it's haproxy or something else, and have each process listen on its own port. This not only allows you to ensure fairness by whatever load balance policy you want, but also lets you have healthchecks, queueing, etc. >Meanwhile, that original request is getting old. The request it made has since received a response, but since there's not been an opportunity to flip back to it, the new request is still cooking. Eventually, that new request's computations are done, and it sends back a reply: 200 HTTP/1.1 OK, blah blah blah. This can happen whether it's an os threaded design or a userspace green-thread runtime. If a process is overloaded, clients can and will timeout on the request. The main difference is in a green-thread runtime it's about overloading the process vs. utilizing all threads. Can make this better by using a local load balancer on box and spreading load evenly. It's also best practice to minimize "blocking" in the application that causes these pauses to happen. >That's why they fork-then-load. That's why it takes up so much memory, and that's why you can't just have a bunch of these stupid things hanging around, each handling one request at a time and not pulling a "SHINYTHING!" and ignoring one just because another came in. There's just not enough RAM on the machine to let you do this. So, num_cpus + 1 it is. Delayed imports (because of cyclical dependencies) is bad practice. That being said, forking N processes is standard for languages/runtimes that can only utilize a single core (python, ruby, javascript, etc.). This is not to say that this solution is ideal -- just that with a small bit of work you can improve the scalability/reliability/behavior under load of these systems by quite a bit.
- jasonhansel 7y agoI really think this should be solved at the OS level. Why is it so hard to implement kernel threads in an efficient way? Threading shouldn't need to be done in userspace.
- asveikau 7y agoIt's not hard. The very start of the article says she trusts proper kernel threads more. But a bunch of these languages were designed and written in a time before threads were much of a thing. So they fake threads in user mode rather than fix the assumptions of the runtime. Now, using kernel threads uncarefully will lead to different problems, which are famously tricky... But what I mean is, it is not "omg kernel threads are hard" that is the primary factor preventing direct access to them. It's that the language runtime has a lot of baggage.
- worik 7y agoThis is exactly what I think. Those below who complain about the complaints are missing the point. We (computer programmers as a general class) have not learnt from history. We keep reinventing wheels and each time they are heavier and clunkier. What we used to do in 40K of scripts now takes two gigabytes in python/django/whateverthehellelse. E.g. mail list servers. Mailman3 hang your head in shame!
- dirtydroog 7y agoSomewhat related to the RPC argument, but HTTP is a total joke, adn therefore so is REST. In adtech you send 204 responses a lot. The body is empty, just the headers. Headers like 'Server' and 'Date'. Apache won't let you turn Server off... 'security through obscurity' or some nonsense. Why do I need to tell an upstream server my time 50k times per second? Zip it all up! Nope, that only applies to the body which is already empty. Egressing traffic! A cloud provider's dream. I wonder what percentage of their revenue come from clients sending the Date header.
- CameronNemo 7y agoThere are zero response headers required for a 204 status. See: https://tools.ietf.org/html/rfc2616#section-10.2.5 https://tools.ietf.org/html/rfc2616#section-10.2.5 Seems like what you are having trouble with is Apache, not HTTP 1.1.
- _bxg1 7y agoHow does this compare with NodeJS, given the event loop and the managed way that concurrent tasks happen there?
- crimsonalucard 7y agoIt's the same thing. She and many other people don't know it but she's complaining about the same model that was introduced and popularized by NodeJS. Current python "green threads" use the keywords async/await as an api. Underneath this api, the state of the art implementations use libuv (wrapped in a python library called uvloop) the exact SAME C++ library that powers nodejs.
- fancyfredbot 7y agoI really like Rachel's blog and I think I understand the point she's making here. However I think she sees it from the point of view of very large scale services. In many cases you can have a solution ready more quickly with less developer time if you use these technologies, and at smaller scale this more than pays for the additional hardware you need to cope with the inefficiency. In such cases writing services in python is pragmatic and sensible.
- cakoose 7y agoIt seems to be a complaint against doing process-per-CPU. Let's say your server has 4 CPUs. The conservative option is to limit yourself to 4 requests at a time. But for most web applications, requests use tiny bursts of CPU in between longer spans of I/O, so your CPUs will be mostly idle. Let's say we want to make better use of our CPUs and accept 40 requests at a time. Some environments (Java, Go, etc) allow any of the 40 requests to run on any of the CPUs. A request will have to wait only if 4+ of the 40 requests currently need to do CPU work. Some environments (Node, Python, Ruby) allow a process to only use a single CPU at a time (roughly). You could run 40 processes, but that uses a lot of memory. The standard alternative is to do process-per-CPU; for this example we might run 4 processes and give each process 10 concurrent requests. But now requests will have to wait if more than 1 of the 10 requests in its process needs to do CPU work. This has a higher probability of happening than "4+ out of 40". That's why this setup will result in higher latency. And there's a bunch more to it. For example, it's slightly more expensive (for cache/NUMA reasons) for a request to switch from one CPU to another, so some high-performance frameworks intentionally pin requests to CPUs, e.g. Nginx, Seastar. A "work-stealing" scheduler tries to strike a balance: requests are pinned to CPUs, but if a CPU is idle it can "steal" a request from another CPU. The starvation/timeout problem described in the post is strictly more likely to happen in process-per-CPU, sure. But for a ton of web app workloads, the odds of it happening are low, and there are things you can do to improve the situation. The post also talks about Gunicorn accepting connections inefficiently and that should probably be fixed, but that space has very similar tradeoffs <https://blog.cloudflare.com/the-sad-state-of-linux-socket-balancing/> https://blog.cloudflare.com/the-sad-state-of-linux-socket-ba....
- drenginian 7y agoOk so if there’s a problem, what a solution? If I use uWSGI is problem gone?
- Matthias247 7y agoI’m not sure what the main point of the article is? Telling us that eventloops have problems? Sure, the lack of preemption can cause latency problems in some tasks. But native threads have other issues - that’s why people use eventloops. Is the message that epoll and co are lot efficient enough? That’s also true. Api Problems and thundering here are known. And not only limited to Python applications as users. io completion based models (eg throuh uring) solve some of the issues. Or is this mainly about Python and/or Gevent? If yes, then I don’t understand it, since the described issues can be found in the same way in libuv, node.js, Rust, Netty, etc
- deleted 7y ago[deleted]
- Fazel94 7y agoI can relate to the writer, working with legacy sucks. This was my main take on the blog post, others are just brilliant ways of rationalizing why other people such and why there are other people than me. Definitely, I am smarter than the guy who wrote this because then I wouldn't have these problems(Or He is smarter and I just didn't ask him about his rationale). What I design wouldn't run into these BS problems that I have to fix, It just wouldn't run into problems generally. (Or It would have more problems than this one) I had these conversations with myself at least a thousand times, and then it was just the case in the parentheses.