4 ms·
R certainly expanded my programming views! Haskell did too, but the lessons of Haskell didn't stick the way that R's lessons did. Here are some of the things I
by datastoat 5y ago
R certainly expanded my programming views! Haskell did too, but the lessons of Haskell didn't stick the way that R's lessons did. Here are some of the things I learnt from R (though they can be found in other languages of course).
* Multiple dispatch. Before learning R, I knew about polymorphism in Java and C++, and multiple dispatch in R broadened my mind and turns out to be very handy.
* The idea of "frames". In R, when you invoke `lm(height~sex*age, data=mydataframe)`, the first argument (the formula) doesn't get evaluated until the lm command asks it to be evaluated, and lm can set up the "frame" for that evaluation, i.e. the place where variables are looked up, however it likes. In fact, lm sets it up to include variables from both the scope in which you invoked lm, and also from mydataframe. This is what makes R so wonderfully concise for modelling in data science, compared to e.g. Python + pandas. I knew about frames from interactive debuggers, but until R it never occurred to me that the programming language could manipulate them.
* "Held" arguments. In R, when you invoke `plot(x, y1+y2)`, it doesn't just evaluate the arguments and then call the plot function -- it leaves the arguments unevaluated, and invokes plot. Plot then (1) decides when to evaluate them, (2) gets access to the language expression `y1+y2`, which means that it can print "y1+y2" on the plot label, (3) it can even define extra variables to include in the scope when y1+y2 gets evaluated. (I knew about held arguments earlier, from Mathematica, but they only clicked when I read the R documentation.)
I've read that R is a descendent of Scheme, and that that's where it gets all its "manipulate language expressions" from. I don't know any Scheme, nor Lisp, and I should definitely learn them -- but in the meantime, my experience has been that R's ability to manipulate language expressions is what makes it such a wonderful sweet spot as a data modelling language. I mostly use Python + pandas nowadays, but it feels such a slog in comparison.
- bookofsand 5y agoSyntactic forms ('frames', 'held arguments') are reasonably useful, but have two flaws: A. Understanding how to implement functions using syntactic forms is a steep learning curve. I remember running out of dplyr and having to implement a udf. Fairly unpleasant experience (enquo, !!, perhaps other unusual constructs). Felt like programming C macros. B. "A function can decide where variables are looked up however it likes" is a significant obstacle in understanding how even basic constructs like function calls actually work. There is a non-trivial amount of hard-to-debug dark magic lurking behind every corner. A middle ground has never been achieved. For example, `plot(expr(x), expr(y1 + y2))`, where the system limits the dark magic to explicit uses of the `expr()` construct, and `expr(x)` always means `{vars: vars(x), expr: (vars(x)) => x}`. Instead of patching interpreter environments, simply call a lambda function.
- datastoat 5y agoI completely agree about the steep learning curve and the feeling of dark magic -- how many times have I had to relearn what deparse(substitute(x)) means -- but oh the satisfaction of broadening my programming horizons. For me it didn't feel like C macros, it felt like "This must be what it feels like to have the power of Lisp"! That's the weird thing about R. All this dark magic is hiding under the hood, but the core R team hid it so deftly that to the casual statistician it's a straightforward data modelling language that "just works". I'm not sure that it's possible to get rid of the dark magic and retain that data-modeller friendliness.