6 ms·
Here's the deal, most tasks I've seen are both CPU and I/O bound, only in different part of the same task. Of course you can very aggressively optimise the desi
by mdomans 10y ago
Here's the deal, most tasks I've seen are both CPU and I/O bound, only in different part of the same task. Of course you can very aggressively optimise the design, but sometimes it becomes so unreadable, the cost you incur through complication makes it not worth the effort.
What I'm trying to advocate is that the distinction of CPU bound and I/O bound makes sense in CS classroom, but most of the tasks in production code are a mix of both.
Therefore, the distinction is really between blocking and non-blocking, and Celery enables "almost readable" non-blocking coding
- garethrees 10y agoThinking about where the computation is bound is only one half of the space of use cases: you also have to consider the shared/non-shared axis. There's a 2x2 matrix of task types: non-shared shared compute-bound multiprocessing ??? I/O-bound either coroutines I put ??? in the upper right because as far as I know Python doesn't have any general solutions here. Also, if "most tasks you've seen" are compute-bound at some points, then of course it makes sense to use Celery! But you should bear in mind that other people may have seen other kinds of task.
- mdomans 10y agoWell, the point I'm making is that I don't see why we we can't have a simple API in Python to define work to do, e.g. access something over the web, without making 30 decisions about the implementation details. And it's not that I think Celery is the best. It has huge drawbacks. I merely suggesting that the model Celery and GCD enforce is much better for writing CPU, I/O and CPU+I/O bound code.
- sunkencity 10y agowell GCD is an evented model and Celery is not so what's your point? Queue and Thread is available in python and work well up to a point. Twisted (and sugar @inlineCallbacks) or similar makes for much more concise code.
- garethrees 10y agoYou're not addressing the point about shared data. How do I program my Celery tasks to operate safely on shared data? Well, I need a system that is responsible for maintaining the shared data, and I have to program my tasks to send it queries and updates, and receive the results, and so probably I'm going to need serializers and deserializers for all my data structures, and if the tasks have transactions (which they almost certainly do) then I'm going to have to implement locks to maintain the data integrity. Whatever it is, this system looks a lot like a database. Whereas with coroutines each task just updates ordinary data structures in memory. No need for communication, queries, serialization, or locks.
- lqdc13 10y agoFor shared compute-bound you can use MPI4Py http://mpi4py.readthedocs.io/en/stable/tutorial.html http://mpi4py.readthedocs.io/en/stable/tutorial.html
- fovc 10y agoFor your upper right quadrant, I think that's where the biggest problem is today. PyParallel is a (long-term) attempt at solving it. Reading through the presentations, it appears having proper OS support is also necessary, so it's not just a Python GIL problem. [1] http://pyparallel.org/ http://pyparallel.org/
- Symmetry 10y agoAnd just as an elaboration, shared only counts if you're sharing mutable state. We have a fair amount of shared sensor data but only need to return a relatively small plan so thanks to the wonders of copy on write multiprocessing works just fine for us.
- dang 10y agoThe latter phrase seems like a reasonable expression of the article, so we've replaced the (arguably baity and misleading) title above. If anyone suggests a better (accurate and neutral) title, we can change it again.