4 ms·
It is particularly infuriating in R, because lm(y ~ ., data = my_dataframe) already means "regress the variable y on all other columns in `my_dataframe`."
by grayclhn 6y ago
It is particularly infuriating in R, because
lm(y ~ ., data = my_dataframe)
already means "regress the variable y on all other columns in `my_dataframe`." For big, interactive regresions, it's really natural to write
my_original_dataframe %>%
do_a_bunch_of_tranformations() %>%
select(...) %>% # Pull out just the columns you want
lm(y ~ ., data = .)
and god knows how that last line is going to be interpreted. So disambiguating through some mechanism is necessary anyway. A lambda is much better than some temporary variable that just holds the formula `y ~ .`.
- torfason 6y agoThe zfit package is intended to address this issue, with the zlm() and comparable functions that are very thin wrappers around lm() and friends. The ony thing they do is flip the argument order so the data comes first, making exactly this use case much simpler. So you can do: cars %>% zlm(dist ~ speed) (or now) cars |> zlm(dist ~ speed) https://github.com/torfason/zfit https://github.com/torfason/zfit
- grayclhn 6y agoTbh, I would 1000% rather my coworkers write a lambda function or closure where it's necessary than add a new package depencency just to change the order of arguments in widely used functions. Plus, I still wouldn't trust the code cars %>% zlm(dist ~ .) to necessarily work the way I want, or to work the same way across package versions.