5 ms·
This looks neat, the _ trick is similar to the Self [0] of fastcore (from fastai). However many things are possible with vanilla pandas. I use it a lot for dat
by rogue7 6y ago
This looks neat, the _ trick is similar to the Self [0] of fastcore (from fastai).
However many things are possible with vanilla pandas. I use it a lot for data munging, usually with the fluent interface (method chaining) style [1], e.g.:
df.loc[lambda f: ...].groupby(...).agg(["mean", "count"])
It also plays nicely with the black autoformatter.
Anonymous functions are verbose and limited in Python, but you can still do many things and use a regular function when a lambda won't do.
I guess one of the thing I need the most when doing data analysis is column name autocompletion, inside groupby, lambdas, for column selection...
I wonder if one could do it in IPython, similar to the string autocompletion of file/directory paths. Basically a parsing of dataframe column names in order to autocomplete strings.
[0]: https://fastcore.fast.ai/basics.html#Self-(with-an-uppercase-S) https://fastcore.fast.ai/basics.html#Self-(with-an-uppercase...
[1]: https://tomaugspurger.github.io/method-chaining https://tomaugspurger.github.io/method-chaining
- ZeroCool2u 6y agoIf the necessary information is there, as in you've mentioned a column name in a dataframe at least once, PyCharm will now do column name autocompletion. It's actually pretty solid in my experience.
- alexilliamson 6y agoDidn't know pycharm would do that... thanks for the info! Yeah the best things about dplyr in opinion are 1) less verbose than pandas 2) much better autocomplete.
- bobbylarrybobby 6y agoIPython already does tab completion of data frame column names. E.G., `df[“col<tab>` will do what you’d hope.
- lordgrenville 6y agoAs long as there isn't a space in the column name. This is riding on a Pandas trick of making the column name accessible as an attribute of the dataframe, which breaks down when there's a space.
- closed 6y agoHey, thanks for pointing out Self--I definitely need to dig into fastcore more! One motivation for developing siuba is that the grouped agg you show requires users specify only one operation on one column. E.g. 1. Calculate mean of x However, common operations like demeaning a column are multiple operations: 1. Calculate mean of x 2. Subtract result of (1) from x In siuba you can just write mutate(res = _.x -_.x.mean()). This isn't possible from something like gdf.x.agg("mean"), and from what I can tell deeply confusing to analysts :/. In vanilla pandas I really like to use the chaining method you laid out, and siuba to me is mostly a utility library for making the approach a little more succinct / performant[1]. siuba has experimental autocompletion (thanks to Tim Mastny!), and there's a pretty hefty technical write up on how it uses IPython machinery for that in siuba's architectural desicion record folder[2]. [1]: https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou... [2]: https://github.com/machow/siuba/blob/master/examples/architecture/006-autocompletion.ipynb https://github.com/machow/siuba/blob/master/examples/archite...
- JPKab 6y agoFastcore is great. I've been going ape with it recently for a work project, and I've particularly fallen in love with patch. Made it extremely easy to assemble a collection of pure functions into pseudo methods attached to an existing class.