4 ms·
Well it depends on your use-case. We have a large (1-20 GB), complex in-memory C++ data structure that's getting read by Python code. "Just use multiprocessing
by dgrunwald 5y ago
Well it depends on your use-case.
We have a large (1-20 GB), complex in-memory C++ data structure that's getting read by Python code.
"Just use multiprocessing" just doesn't work in this situation. We can't afford to multiply our memory consumption by 128.
We've already invested many months of developer time just to allow some parts of the data structure to be allocated in shared memory. But this has been highly complex (e.g. no more std::string for us), and it only allowed us to parallelize about half of the Python scripts; and the gains on the overall execution time have been fairly limited.
I'm starting to think we will have to rewrite 50000+ lines of Python code in another language, just so that we can parallelize something that would already be embarrassingly parallel in any other language.
I'm hoping that the "subinterpreters with separate GILs" will work out at some point; but I'm doubtful this can work as long as Python doesn't break compatibility with existing C extensions.
And yes, on Linux `fork()` trivially solves the "share complex C++ data structure for Python multiprocessing" issue. Unfortunately most of our customers use Windows :(
- hutrdvnj 5y ago> Unfortunately most of our customers use Windows :( Maybe you can run your application in WSL2 with `fork()` on Windows.
- _dps 5y agoI don't know what kind of latency vs throughput tradeoff you're facing, but given these constraints (and being on windows) I might try running multiple python processes each with the C++ structure in shared memory, and then broker any inter-python shared state through something like Redis on localhost (so as long as the inter-python shared state is small and rarely updated, you only pay a small IPC cost to get it from Redis instead of having it in-process).