4 ms·
This particular pipe comes from F#. Let's say I have a table called dat. dat %>% filter(col_a == 'Good') %>% group_by(col_b) %>% summariz
by sin7 9y ago
This particular pipe comes from F#.
Let's say I have a table called dat.
dat %>%
filter(col_a == 'Good') %>%
group_by(col_b) %>%
summarize(n = n(), sum_c = sum(col_c))
To do this in traditional R, I would have to:
dat <- dat[dat$col_a == 'Good']
dat_n <- aggregate(col_a ~ col_b, dat, length)
dat_sum <- aggregate(col_c ~ col_b, dat, sum)
merge(dat_n, dat_sum, by = "col_b")
I think the piped version is more readable. At least there are less variable to track.
- extr 9y agoThat's because you're comparing to base-r. The data.table way would be: dat[col_a == 'Good', .(Length = .N, Sum = sum(col_c)), col_b]
- jcheng 9y agoThat's all well and good for operations that are built into data.table; pipes can be used with anything that's a function call (like ggpage).
- nerdponx 9y agoData.table is great for SQL-style code but it imposes some annoying limitations, namely that the output is always coerced to a data.table.
- jhbadger 9y agoThe problem with data.table is that in practice your data gets converted to something else when you pass it through packages -- many functions will return a data.frame or matrix, others in the Hadleyverse will return a tibble, and so on. So you have to constantly force your data back into a data table. R has so many datatypes that basically represent a spreadsheet/database table.