8 ms·
What would it take to recreate dplyr in Python? (2020)
- mint2 5y agoThat looks neat and was very interesting. Pandas is sometimes a dark art to know how to do something fast enough. Port the functionality of the R package but try to keep it python. Run flake8.
- lysecret 5y agoHey regarding groupby operations in pandas. I have been moving more and more (for too complicated groupbys) to running them something like: out_rec = [] for id, group in data_frame.groupby("id"): ladidida.... result = f(group) out_rec.append(result) in my experience it isn't much slower than a groupby.apply.
- ellisv 5y agoPerformance aside, I think one of the main criticism of `groupby` in pandas is that they are not composable and become verbose. I'd argue that's the case with your example too.
- sweezyjeezy 5y agogroupby.apply is often the thing that is already way too slow though based on use-case - because I think it pretty much does what you wrote rather than something that makes stronger optimisation assumptions.
- closed 5y agoAuthor here. My concern in the article isn't that you can't do groupby fast, but that the approach above is not composable. * If f() converts grouped data to something ungrouped, then you can't use a similar function f2(f(group)) * If f() returns a grouped object, then you can't do basic operations like f(grouped) + 1, because DataFrameGroupBy, SeriesGroupBy do not define basic operators like addition. Let alone operations against other grouped data. A lot of this is worked out in siuba now, and this doc explains a bit more: https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
- lysecret 5y agoYea makes sense. I actually come from R and the deplyr world. Was really happy to read your article :). I might give siuba a try.
- sweezyjeezy 5y agoAlways been interested to know why pandas implemented index the way it did. I generally find myself doing .reset_index on everything by default because it's just one less thing to think about, but it's clear that pandas devs are very fond of it based on the API. Where it still feels weird is when e.g. groupby/pivot by default return everything with a custom index, when I've given no indication that it needs to be treated differently to a column, but then e.g. merge doesn't do this? It's also written out by default in .to_csv, like just... why? Not useful for any csv that is to be used outside pandas. God help you if you end up needing to use a multi-index for something - deeply unpleasant. Was this just a high level (possibly misguided) paradigm that the pandas devs fell in love with - or is there a good, performance related reason to embed it so deeply in the API?
- lysecret 5y agoI have been using indexes more and more if you do essentially series based operations but you need to re-associate the results back. It enables you to avoid using dataframes as inputs to your data transformations.
- sweezyjeezy 5y agoI am experienced with pandas and I understand its uses - but it definitely makes learning it much more confusing for new people - like the first time you see a multi-index you're just like 'OMG what'. It also feels like it silghtly breaks the mental model of a dataframe for me - like why am I treating these columns as special all of a sudden? Sometimes that can make things slicker, but 90% of the time for me it just necessitates the need for .reset_index or index=False or equivalent. If I want to use index to optimise something - that should be a deliberate act, not something pushed on me by the API.
- cardosof 5y agoAs someone with years of working experience with dplyr who had to learn pandas, 100% agree. And thanks for that post, I thought it was just me.
- 5y ago
- mgradowski 5y agoDuckDB and Polars are my bets in the Python data-wrangling space. I grew tired of Pandas' weird-ass API.
- sweezyjeezy 5y agoI would love to switch to something else, but it feels like pandas is lingua-franca in data science now, to switch puts a burden on everyone else.
- mytherin 5y agoYou can use DuckDB as a processing engine on top of Pandas [1], while continuing to use Pandas as a data storage/data interchange format. [1] https://duckdb.org/2021/05/14/sql-on-pandas.html https://duckdb.org/2021/05/14/sql-on-pandas.html
- mgradowski 5y agoThat's what I do at $dayjob whenever I have to do windowing &c. Figuring out this stuff in Pandas is a waste of time. Before I discovered DuckDB, I would re-learn the API every damn time. I came up with a little utility function, which you can implement yourself :) ``` def sqldf(df: DataFrame, query: str) -> DataFrame: ... ```
- kristjansson 5y agoYears of unpicking others use of Rs sqldf (which by default used to copy the entire data frame to a SQLite db, run the query, the copy the result set back) when they complained their R code was to slow has taught me a visceral, negative to the name and pattern. Glad to to see duckDB delivering, finally, on the promise of running SQL against in-memory dataframes
- mgradowski 5y agoTIL there's an actual 'botched' library with the same name; I actually came up with it independently on a lazy office afternoon :^)
- bobolito 5y agoCheck out Tidypolars
- mrtranscendence 5y agoNice, but honestly, Polars offers a nice enough API that I don't miss dplyr as much when using it. I probably wouldn't bother with another API on top of Polars unless it were particularly feature-rich.
- elforce002 5y agoSame. I'd keep an eye on tidypolars but polars is good for now.
- usermi 5y agoThere is a project called Datar (https://github.com/pwwang/datar https://github.com/pwwang/datar), which mimics dplyr in Python.
- malshe 5y agoThanks for posting it here. I was going to look for this exact package to post here but I forgot its name!
- peatmoss 5y agoDplyr (and its spin-off dbplyr) is to me a fantastic practical example of the power of lispy metaprogramming. While R gets knocked a lot for not being a "real" programming language, or being weird, or just generally being different than e.g. Python, I don't know that I've seen a cooler example of high-level programming language ideas expressed as cleanly. Ditto the rest of the tidyverse.
- mbrudd 5y agoAgree 100%!
- kristjansson 5y agoR’s waxing popularity is a testament to that power. dplyr and friends effectively reject much of the R-the-language and substitute their own friendlier, more popular syntax without rejecting R-the-platform.
- closed 5y agoAuthor here--happy to answer questions :) Siuba has come a long way since I wrote this, and now can optimize for fast grouped operations!: * https://github.com/machow/siuba https://github.com/machow/siuba * https://siuba.readthedocs.io/en/latest/developer/pandas-group-ops.html https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
- dunefox 5y agoAs a comparison to another language: for Julia there is DataFrames.jl: https://dataframes.juliadata.org/stable/ https://dataframes.juliadata.org/stable/ Comparison to dplyr, ...: https://dataframes.juliadata.org/stable/man/comparisons/#Comparison-with-the-R-package-dplyr https://dataframes.juliadata.org/stable/man/comparisons/#Com... Comparison to Pandas: https://dataframes.juliadata.org/stable/man/comparisons/#Comparison-with-the-Python-package-pandas https://dataframes.juliadata.org/stable/man/comparisons/#Com...
- sundarurfriend 5y agoUsing DataFramesMeta.jl [1], the code from the first example turns out pretty similar: using DataFramesMeta, Statistics @chain mtcars begin @select :Cyl :HP groupby(:Cyl) @transform _ begin :dumb_result = myarbitraryfunc.(:HP) :demeaned = :HP .- mean(:HP) end end (mtcars is from RDataSets.jl) [1] https://juliadata.github.io/DataFramesMeta.jl/stable/dplyr/ https://juliadata.github.io/DataFramesMeta.jl/stable/dplyr/
- armanboyaci 5y agoI am distracted with the example provided in the post. I am pretty sure that most pandas users will just put course_ids on the columns. I mean the shape of the user_courses dataframe is not suitable for the task. (user_courses .set_index(["student_id", "course_id"]) .unstack() .apply(lambda x: x+1))
- psimm 5y agoThis article compares dplyr syntax with pandas, siuba, polars, ibis and duckdb: https://simmering.dev/blog/dataframes/ https://simmering.dev/blog/dataframes/ As other have said, escaping pandas is hard. Many visualization and data manipulation, validation and analysis libraries expect pandas input. Siuba is really cool in that it offers a convenient syntax on top of pandas (and SQL databases) without requiring its own data format.
- tpoacher 5y agoMeanwhile, every time I use an external library (including pandas) I still think numpy does everything you need, and does it well, and people just haven't bothered to learn it properly and keep reinventing the wheel. (no disrespect to to the package in the article or OP who I know is active in this thread. just a general motif that I keep coming across in python).
- closed 5y agoI used numpy a lot for my PhD research, and think you're right, in the sense that things could have been much more numpy array centric (just like how in R the dataframe is a tiny structure around arrays). There are a lot of problems you encounter when using arrays for data analysis, like some of their funky behavior with strings [0], but it seems like extending arrays, or building new types of numpy arrays would have been better than new data structures like the pandas Series. (pandas folks thought a lot about these problems so I could be very wrong). [0]: https://mchow.com/posts/pandas-has-a-hard-job/ https://mchow.com/posts/pandas-has-a-hard-job/