5 ms·
multiprocessing only works fine when you're working on problems that don't require 10+ GB of memory per process. Once you have significant memory usage, you rea
by ynik 3y ago
multiprocessing only works fine when you're working on problems that don't require 10+ GB of memory per process.
Once you have significant memory usage, you really need to find a way to share that memory across multiple CPU cores. For non-trivial data structures partly implemented in C++ (as optimization, because pure python would be too slow), that means messing with allocators and shared memory. Such GIL-workarounds have easily cost our company several man-years of engineer time, and we still have a bunch of embarrassingly parallel stuff that we still cannot parallelize due to GIL and not yet supporting shared memory allocation for that stuff.
Once the Python ecosystem supports either subinterpreters or nogil, we'll happily migrate to those and get rid of our hacky interprocess code.
Subinterpreters with independent GILs, released with 3.12, theoretically solve our problems but practically are not yet usable, as none of Cython/pybind11/nanobind support them yet. In comparison, nogil feels like it'll be easier to support.
- pillusmany 3y ago"Ray" can share python objects memory between processes. It's also much easier to use than multi processing.
- ptx 3y agoHow does that work? I'm not familiar with Ray, but I'm assuming you might be referring to actors [1]? Isn't that basically the same idea as multiprocessing's Managers [2], which also allow client processes to manipulate a remote object through message-passing? (See also DCOM.) [1] https://docs.ray.io/en/latest/ray-core/walkthrough.html#calling-an-actor https://docs.ray.io/en/latest/ray-core/walkthrough.html#call... [2] https://docs.python.org/3/library/multiprocessing.html#managers https://docs.python.org/3/library/multiprocessing.html#manag...
- pillusmany 3y agoShared memory: https://docs.ray.io/en/latest/ray-core/objects.html https://docs.ray.io/en/latest/ray-core/objects.html
- ptx 3y agoAccording to the docs, those shared memory objects have significant limitations: they are immutable and only support numpy arrays (or must be deserialized). Sharing arrays of numbers is supported in multiprocessing as well: https://docs.python.org/3/library/multiprocessing.html#sharing-state-between-processes https://docs.python.org/3/library/multiprocessing.html#shari...
- ebiester 3y agoAnd I guess what I don't understand is why people choose Python for these use cases. I am not in the "Rustify" everything camp, but Go + C, Java + JNI, Rust, and C++ all seem like more suitable solutions.
- oivey 3y agoNotably, all of those are static languages and none of them have array types as nice as PyTorch or NumPy, among many other packages in the Python ecosystem. Those two facts are likely closely related.
- samatman 3y agoIf only there were a dynamic language which performs comparably to C and Fortran, and was specifically designed to have excellent array processing facilities. Unfortunately, the closest thing we have to that is Julia, which fails to meet none of the requirements. Alas.
- rmbyrro 3y agoIf only there was a car that could fly, but was still as easy and cheap to buy and maintain :D
- abdullahkhalids 3y agoPython is just the more popular language. Julia array manipulation is mostly better (better syntax, better integration, larger standard library) or as good as python. Julia is also dynamically typed. It is also faster than Python, except for the jit issues.
- jononor 3y agoI think that 90 or maybe even 99% of cases has under 1GB of memory per process? At least it has been the case for me the last 15 years. Of course, getting threads to be actually useful for concurrency (GIL removed) adds another very useful tool to the performance toolkit, so that is great.