4 ms·
What I think would be a nice approach would be someone creating a new Dataframe class that merely inherits from a Pandas dataframe but just implementing .filter
by jphoward 6y ago
What I think would be a nice approach would be someone creating a new Dataframe class that merely inherits from a Pandas dataframe but just implementing .filter(), .summarise(), .select() etc. as methods, using the dplyr syntax. If they return their own dataframes then chaining follows naturally.
I know that isn't how dplyr works, but it feels more Pythonic, and this solution isn't entirely analogous to dplyr anyway.
This approach seems a little complicated, though I'm sure with some use I could learn to enjoy it.
- closed 6y agoHey, author of siuba here, I totally agree that subclassing would be a natural choice in python. One challenge there is that users will often get a DataFrame back (e.g. from pd.read_csv), so it requires a lot of casting to the child class. Right now, a compromise I've been exploring is just attaching siuba's DF functions to a pandas DataFrame, e.g. df.siu_mutate(...). This seems to be what pandas wants people to do [1]! One obstacle here is that the DataFrame has 300+ methods, which can be overwhelming to learners. I've spent a lot of time wondering whether the piping syntax feels like too much vs chaining. It's still an open question in my mind, so it's really helpful to hear what feels most natural! https://pandas.pydata.org/pandas-docs/stable/development/extending.html#registering-custom-accessors https://pandas.pydata.org/pandas-docs/stable/development/ext...
- ellisv 6y ago> What I think would be a nice approach would be someone creating a new Dataframe class that merely inherits from a Pandas dataframe but just implementing .filter(), .summarise(), .select() etc. as methods, using the dplyr syntax. If they return their own dataframes then chaining follows naturally. > I know that isn't how dplyr works, but it feels more Pythonic, and this solution isn't entirely analogous to dplyr anyway. I think I agree. Pandas has a lot of overhead/baggage that I don't want 90% of the time. Being able to chain _simple_ verbs on a data frame would be great -- something like mtcars.groupby(cyl).summarize(avg_hp = hp.mean())
- closed 6y agoI'm still debating chaining vs piping, but you can do.. from siuba import _ from siuba.data import mtcars # mtcars is a pandas DataFrame mtcars \ .groupby("cyl") \ .siu_summarize(avg_hp=_.hp.mean())