4 ms·
Multiprocessing is great as a first pass parallelization but I've found that debugging it to be very hard, especially for junior employees. It seems much easie
by jw887c 3y ago
Multiprocessing is great as a first pass parallelization but I've found that debugging it to be very hard, especially for junior employees.
It seems much easier to follow when you can push everything to horizontally scaled single processes for languages like Python.
- flakes 3y agoDepends on the workflow. For one off jobs or client tooling, parallelism makes sense to have rapid user feedback. For batch pipelines on that work many requests, having a serial workflow has a lot of the advantages you mention. Serial execution makes the load more predictable and makes scaling easier to rationalize.
- uniqueuid 3y agoI agree. The main problems aren't syntax, they are architectural: Catching and retrying individual failures in a pool.map, anticipating OOM with heavy tasks, understanding process lifecycle and the underlying pickle/ipc. All these are much more reliably solved with horizontal scaling. [edit] by the way, a very useful minimal sugar on top of multiprocessing for one-off tasks is tqdm's process_map, which automatically shows a progress bar https://tqdm.github.io/docs/contrib.concurrent/ https://tqdm.github.io/docs/contrib.concurrent/
- xapata 3y agoHow is coordinating between different machines any different than coordinating between different processes? A multiprocessing implementation is a good prototype for a distributed implementation.
- uniqueuid 3y agoTo be honest, I don't think both are similar at all. Parallelizing across machines involves networks, and well, that's why we have jepsen, and byzantine failures, and eventual consistency, and net splits, and leadership election, and discovery - so in short a stack of hard problems that in and of itself is usually much larger than what you're trying to solve with multiprocessing.
- deleted 3y ago[deleted]
- xapata 3y agoTrue, the networking causes trouble. I usually rely on a communication layer that addresses those troubles. A good message queue makes the two paradigms quite similar. Or something like Dask (https://www.dask.org/ https://www.dask.org/). Having your single-machine development environment able to reproduce nearly all the bugs that arise in production is a wonderful thing.
- Terretta 3y agoFrom the linked Mpire readme: Suppose we want to know the status of the current task: how many tasks are completed, how long before the work is ready? It's as simple as setting the progress_bar parameter to True: with WorkerPool(n_jobs=5) as pool: results = pool.map(time_consuming_function, range(10), progress_bar=True) And it will output a nicely formatted tqdm progress bar.
- dr_kiszonka 3y agoParsl has quite good debugging facilities built in, which include automatic logging and visualizations. https://parsl.readthedocs.io/en/stable/faq.html https://parsl.readthedocs.io/en/stable/faq.html https://parsl.readthedocs.io/en/stable/userguide/monitoring.html https://parsl.readthedocs.io/en/stable/userguide/monitoring....
- wheelerof4te 3y agoOr just use numpy's arrays, which have their integrated multiprocessing.