3 ms·
For a single workstation and the task you describe, the pool.map() functionality of the multiprocessing module should be perfectly adequate. Not sure how schedu
by juanjgalvez 8y ago
For a single workstation and the task you describe, the pool.map() functionality of the multiprocessing module should be perfectly adequate. Not sure how scheduling overhead would compare between charmpy and multiprocessing, but for this task it shouldn't matter (I assume you need at least a second to convert one file, and even if the conversion is faster, you can chunk the tasks anyway to mask overhead). I would say the big difference for this task is if you want to run it in parallel on multiple hosts, which pool.map can't do. With charmpy we can provide a distributed parallel map offering the same or similar API as pool.map. There is a simple example in 'examples/parallel-map/par-map.py', but we are working on offering a library on top of charmpy with more features and a solid API.
- p1esk 8y agoOh, good point about batching - my files were really small (audio samples for speech recognition), so a conversion of a single file took a lot less than a second. I looked at the par-map.py example, however I can't quite understand where do I enter a server IP or something like that. The whole process is fuzzy to be honest. What do I need to do if I want to run my conversion task on two local workstations? E.g. I install CharmPy on both, then what?
- juanjgalvez 8y agoYou don't actually have to specify hosts or addresses in your application code. When the application starts, the runtime will know how many processes there are and on which hosts. The key is to use a job launcher. For the par-map.py example, suppose you want to run it on 4 hosts and 8 processes per host. One way to do this is by launching the application with "charmrun". First, install charmpy on all hosts like you said. Then you would create a nodelist file with the names or addresses of the 4 hosts. Finally, launch like this: `$ charmrun +p32 par-map.py ++nodelist mynodelist.txt` I have updated the "Running" section of the docs to try to explain this better, also pointing to the charmrun manual. Hopefully things are clearer now.
- p1esk 8y agoThank you, now it's a lot clearer. I will try it on multiple workstations next time I need to run a large job.
- samtwhite 8y agoAnd on most clusters and supercomputers you don't need to manually create the nodelist at all, charmrun can do that automatically for you by parsing the batch scheduler's list of allocated hosts.