7 ms·
The ergonomics of grouping and aggregation in R are really much better because libraries can make of its non-standard evaluation[^0] (which in other cases also
by aorist 4y ago
The ergonomics of grouping and aggregation in R are really much better because libraries can make of its non-standard evaluation[^0] (which in other cases also makes the language a nightmare to deal with).
Compare:
pd_df.groupby(['date'])['failure'].count() # pandas
pl_df.groupby(pl.col('date')).agg(pl.count('failure')) # polars
dt[, .N, date] # R data.table
In both Pandas and Polars, the specification of the date has to be a string inside a list or method call, but in R it can be a bare token.
[^0]: http://adv-r.had.co.nz/Computing-on-the-language.html http://adv-r.had.co.nz/Computing-on-the-language.html
- barumrho 4y agoWhile it's shorter, it seems more magical? How does it know to count `failure`?
- deleted 4y ago[deleted]
- aorist 4y agoIt doesn't count `failure` — just the number of rows. But neither does the pandas version: `pd_df.groupby(['date'])['failure'].count()` and `pd_df.groupby(['date']).count()` are the same except the former returns a single `pd.Series` with the count and the latter produces a `pd.DataFrame` where each column has the same count (not super useful). e.g. > iris.groupby('species').count() sepal_length sepal_width petal_length petal_width species setosa 50 50 50 50 versicolor 50 50 50 50 virginica 50 50 50 50 vs. > iris.groupby('species')['sepal_length'].count() species setosa 50 versicolor 50 virginica 50
- iamlemec 4y agoI believe `count` will only give you the number of non-null rows, so the numbers from the first command could differ by column if there were null values. You can also use the `size` command to get the total number of rows, and that will return a `pd.Series` with or without a column specifier.
- wokwokwok 4y agoHaving migrated 1000s of lines of legacy r, all I can say is… yes, but then you have you to use r. (: R is not a replacement for pandas. R is it’s own special little painful ecosystem, loved by people who don’t have to maintain the code they write. You can complain all you like about pandas, but at the end of the day, it’s python. Python tooling works with it. The python ecosystem works with it. It’s not without faults, but at least you’ll have a community of people to help when things go wrong. R, not so much. (Spoken as jaded developer who had to support r on databricks, which is deep in the hell of “well, it’s not really a tier one language” even from their support team)
- jasonpbecker 4y agoHaving written tens of thousands of lines of R code that I've been maintaining and using for production pipelines for 9 years... Sounds like you worked with (or wrote) really bad R code.
- wokwokwok 4y agoThe point I’m making isn’t that you can write bad r code; you can write bad code in any language. …the point I’m making is that when you already have bad r code, it’s a massive pain in the ass. Bad python is terrible too, but lots of people know python, and it’s easy to find help to unduck bad python code and turn it into maintainable code. That has not been my experience with r. Ever. At any organisation. Your experience may vary. (:
- nequo 4y ago> when you already have bad r code, it’s a massive pain in the ass. Do you think it’s because R code tends to be written by statisticians and stats-adjacent domain experts who don’t necessarily know how to write clean code while Python code has at least some input from actual programmers? Or is this really down to the language itself?
- bovinejoni 4y agoCan you elaborate please? This hasn’t been my experience whatsoever, curious what the issues have been
- minimaxir 4y agodata.table syntax is indeed more concise but harder to parse later, which is more important in a collaborative environment. dplyr syntax IMO is the best balance between clarity and nonredundancy as evident in pandas/polar code.
- tryptophan 4y agoR is easy and fun to use. R is also impossible to understand whta is actually going on.
- miohtama 4y agoR is Perl for math
- mutant_self 4y agoI would do the Pandas and Polars examples differently: ``` pd_df['date'].value_counts() # pandas pl_df.select(pl.col('date').value_counts()) # polars ``` Note: we could also do it the Pandas way in Polars but square bracket indexing in Polars is not recommended