3 ms·
Agreed RE GroupBy being challenging, especially compared to dplyr. As I've worked on a port of dplyr to python over the past year, though, I've realized the dt
by closed 7y ago
Agreed RE GroupBy being challenging, especially compared to dplyr.
As I've worked on a port of dplyr to python over the past year, though, I've realized the dtype issue (like you said), indexes, and GroupBy being difficult are likely connected. Basically,
* dplyr can chop up a dataframe into 50,000 groups and apply arbitrary functions to it--no problem.
* custom pandas grouped applies are very slow
There are basically three reasons for slow pandas apply methods...
1. creating an index for each subgroup is slow (will not be a RangeIndex)
2. initializing a series for each group is slow (mostly due to type inference being re-run; could be avoided)
3. AFAIK more type inference is run when concatenating results
This leads to a world where grouped calculations can't be run using arbitrary expressions (e.g. lambdas), but have to go through specific SeriesGroupBy methods.
I wrote a bit on how I tried to work around that, to enable fast dplyr-like syntax over grouped data in python. Would definitely be interested in your take! There are other libraries, like ibis that do a good job with it, too!
https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
- has2k1 7y agoI have been through this path of trying to figure out why pandas Groupby is slow. It is a big bottleneck when building packages on top of pandas.
- closed 7y agoWell, I'm a big fan of plotnine, and plydata was part of the inspiration for siuba! I think at its core, the groupby issue is a really big problem, and am devoting most of this year to working on it. So if you ever want to pair to work on pandas / pandas wrapping libraries send me an email (link in profile)!