5 ms·
Thinking about where the computation is bound is only one half of the space of use cases: you also have to consider the shared/non-shared axis. There's a 2x2 ma
by garethrees 10y ago
Thinking about where the computation is bound is only one half of the space of use cases: you also have to consider the shared/non-shared axis. There's a 2x2 matrix of task types:
non-shared shared
compute-bound multiprocessing ???
I/O-bound either coroutines
I put ??? in the upper right because as far as I know Python doesn't have any general solutions here.
Also, if "most tasks you've seen" are compute-bound at some points, then of course it makes sense to use Celery! But you should bear in mind that other people may have seen other kinds of task.
- mdomans 10y agoWell, the point I'm making is that I don't see why we we can't have a simple API in Python to define work to do, e.g. access something over the web, without making 30 decisions about the implementation details. And it's not that I think Celery is the best. It has huge drawbacks. I merely suggesting that the model Celery and GCD enforce is much better for writing CPU, I/O and CPU+I/O bound code.
- sunkencity 10y agowell GCD is an evented model and Celery is not so what's your point? Queue and Thread is available in python and work well up to a point. Twisted (and sugar @inlineCallbacks) or similar makes for much more concise code.
- garethrees 10y agoYou're not addressing the point about shared data. How do I program my Celery tasks to operate safely on shared data? Well, I need a system that is responsible for maintaining the shared data, and I have to program my tasks to send it queries and updates, and receive the results, and so probably I'm going to need serializers and deserializers for all my data structures, and if the tasks have transactions (which they almost certainly do) then I'm going to have to implement locks to maintain the data integrity. Whatever it is, this system looks a lot like a database. Whereas with coroutines each task just updates ordinary data structures in memory. No need for communication, queries, serialization, or locks.
- lqdc13 10y agoFor shared compute-bound you can use MPI4Py http://mpi4py.readthedocs.io/en/stable/tutorial.html http://mpi4py.readthedocs.io/en/stable/tutorial.html
- fovc 10y agoFor your upper right quadrant, I think that's where the biggest problem is today. PyParallel is a (long-term) attempt at solving it. Reading through the presentations, it appears having proper OS support is also necessary, so it's not just a Python GIL problem. [1] http://pyparallel.org/ http://pyparallel.org/
- Symmetry 10y agoAnd just as an elaboration, shared only counts if you're sharing mutable state. We have a fair amount of shared sensor data but only need to return a relatively small plan so thanks to the wonders of copy on write multiprocessing works just fine for us.