4 ms·
I've found splitting the dataframe, and using multiprocessing module with `apply` to compute chunks of data to be quite efficient. One can use `group_by` method
by prashnts 9y ago
I've found splitting the dataframe, and using multiprocessing module with `apply` to compute chunks of data to be quite efficient. One can use `group_by` method for that, or just slice the dataframe.
For example:
concurrency = 4 # Num of cores
pool = multiprocessing.Pool(processes=concurrency)
results = pool.map(fn, df.group(...)) # fn would be a callable for computing on a chunk.
pool.close()
return pd.concat(results)