3 ms·
> We found R has a really hard time integrating into data pipelines and was best used as a standalone tool by individuals In my org we have several 100% R team
by civilized 3y ago
> We found R has a really hard time integrating into data pipelines and was best used as a standalone tool by individuals
In my org we have several 100% R teams (including mine) that have been developing and maintaining business-critical, data-intensive applications for a decade now. We don't find R difficult to integrate into data pipelines. We write our data pipelines in R, and we find it very efficient to do so. They talk to databases, APIs, command line tools, etc without issue.
Doing what we do in Python is unimaginable, especially if pandas is the tabular lingua franca in the team. I vehemently agree with this article on the clunkiness of pandas from a sister comment: https://www.sumsar.net/blog/pandas-feels-clunky-when-coming-from-r/ https://www.sumsar.net/blog/pandas-feels-clunky-when-coming-.... Compared to dplyr and the tidyverse, pandas very noticeably gets in your way rather than being a tool of thought. (For what it's worth, there are other teams in my org that use Python for entirely justified reasons, and they use polars these days, not pandas.)
If I had to complain about anything in R these days, it would be the increasing complexity and illegibility of error messages. Tidyverse tracebacks are often dozens or hundreds of lines. This is made much worse if you have a web app in the Shiny framework, as Shiny seems to mangle and garble what little useful information you can get (my kingdom for an error with a file name and line number). Even outside of advanced packages like Shiny, the reporting of error messages suffers from some clunkiness and irregularity.
As an expert user, I can usually squint at the error barrage and infer what is really going on, but it's probably quite confusing and off-putting to newer users.
Overall though, I'm not seeing any competition for R in our space. My fondest hope is that in the coming decades there arises a new, thoughtfully designed language with the Lispy flexibility of R, but also optional type safety and static analysis affordances. I'm not sure if that's even possible, but I hope the computer science geniuses figure out a way.
- samstave 3y ago[flagged]
- hadley 3y agoIf you have specific issues around error messages and tracebacks please feel free to let me know directly or to file issues on Github. We really do care about the legibility of errors and tracebacks and me and my team have put a lot of effort into them in the last few years. But there's always room to do better and I'd love to know where the pain points are. (The intersection of tidyverse and shiny tracbacks are a known pain point that's hard to resolve. Unfortunately shiny and tidyverse did a bunch of parallel work that took us in slightly different directions and now it's hard to re-align.) One thing we are missing is a guide to reading traceback for newer users. Often experts can get a good sense of where the problem is, but we've failed to teach newer users how to get the most value from a traceback.
- civilized 3y agoThere's clearly been a ton of progress in this area; the only issue is that feature development is even faster :) I'll keep an eye out for specific issues that seem helpful to raise. The biggest one I have right now is a little niche, but probably useful to address. Moderately complex dbplyr pipelines on wide tables have a tendency to generate very long queries, and if there's an error, the generated SQL returned tends to overflow some text or line limit allotted to show the error at the command prompt. My workaround is to use sink() to dump the error to a file, which is a little painful as the sink() API and documentation are not the most straightforward or intuitive. (Hmm, I wonder if a withr wrapper would help me make something simpler to use...)
- hadley 3y agoHmmmm, I think that's something we could probably help with in dbplyr by providing something like `last_sql()` that would return the most recent SQL sent to the database. (By analogy to ggplot2::last_plot() and httr2::last_request()/last_response()). I filed an issue so I don't forget about this: https://github.com/tidyverse/dbplyr/issues/1471 https://github.com/tidyverse/dbplyr/issues/1471
- xapata 3y agoWhy Polars and not Dask?
- ramblenode 3y ago> My fondest hope is that in the coming decades there arises a new, thoughtfully designed language with the Lispy flexibility of R, but also optional type safety and static analysis affordances. I think many of us saw Julia as the successor to R. Unfortunately, the package ecosystem---one of R's strongest points---still has a long way to go.
- civilized 3y agoI was excited about Julia too but it now seems to be a relatively niche HPC language. It's about saving CPU time more than user time. My sniff test for a successor language to R is whether it can replicate the tidyverse API with 100% fidelity. The API is already optimal for tabular data analysis, especially the dplyr core. It can be thought of as a specification for other languages to implement. There is a great deal about how R works that is negotiable. But if the language can't implement dplyr to spec, or somehow doesn't "want to", it's not the language for the audience served by the tidyverse.
- hadley 3y agoI love this framing :)
- civilized 3y agoFollow-up: I did a little legwork and found some Julia folks who were also inspired to implement tidyverse as faithfully as possible: https://github.com/TidierOrg https://github.com/TidierOrg Here's how dplyr-style chains look in their system: using TidierData using RDatasets movies = dataset("ggplot2", "movies"); @chain movies begin @mutate(Budget = Budget / 1_000_000) @filter(Budget >= mean(skipmissing(Budget))) @select(Title, Budget) @slice(1:5) end Not a character-for-character match to dplyr, but gets much closer than most other attempts!