8 ms·
R, and by R I mean R+tidyverse, is the world's best graphing calculator attached to an OK scheme. To which I mean R is a highly optimized, well-oiled machine i
by tel 5y ago
R, and by R I mean R+tidyverse, is the world's best graphing calculator attached to an OK scheme.
To which I mean R is a highly optimized, well-oiled machine if you're using it for its highly-optimized, well-oiled purposes. I tend to have notebooks full of tiny fragments like this
dat_min %>%
group_by(ymd = make_date(year(date), month(date), day(date))) %>%
summarize(vol_btc=sum(vol_btc), vol_usdt=sum(vol_usdt), tradecount=sum(tradecount)) %>%
ungroup() %>%
pivot_longer(cols=c(-ymd)) %>%
ggplot(aes(ymd, value)) +
geom_line() +
facet_grid(name ~ ., scales="free_y")
It's madness if you're not familiar with the tidyverse, but 3 dozen fragments like this is enough to eviscerate a fresh data set. Almost any question you can dream of is a 3-20 line set of transforms away from a beautiful plot or analysis answering your question. Very notably, this includes some of the finest modeling tools available today.
Terseness here is a huge advantage as well because in many data analysis workflows you are rerunning that same 10 line snippet over and over, making small changes, adjusting to eventually visualize the thing you're looking for perfectly. Having all of that in the same small block is ideal.
Finally, for the non-trivial number of folks in this specific scenario, the integration between Stan and R/RStudio is top-notch and makes using both tools very pleasant.
You can replicate all of this in Python, but optimal Python/Jupyter is still a far cry away from R/RStudio for these specific sorts of tasks.
- CapmCrackaWaka 5y agoI also use R for any heavy data manipulation, but I primarily use the data.table package. The efficiency that both of these packages unlock is absolutely unparalleled in any other tabular data manipulation library, in any other language that I have used. And R has the top 2!! My skin writhes every time I need to type: table.loc[(table.column > 2) | (table.column2 < 3)].reset_index(drop=True) when I want to subset a table.
- alexilliamson 5y agoNot to mention the auto complete that comes with RStudio. Is there any way to get equivalent functionality in Jupyter?
- CapmCrackaWaka 5y agoI use pycharm which has decent autocomplete. Pycharm has its own issue though, it fills out its autocomplete info by looking at the function that created the object, not the object itself. So if a function can return different types, autocomplete won’t work. That’s caused me quite a bit of pain.
- discordance 5y agoIf you set up use the Jupyter extension[0] and open your notebooks in VS Code you get Intellisense (code completion, method info and hints etc). 0: https://marketplace.visualstudio.com/items?itemName=ms-toolsai.jupyter https://marketplace.visualstudio.com/items?itemName=ms-tools...
- claytonjy 5y agoIME this is strictly worse than the RStudio experience; most of the time I hit tab in a VSCode notebook I get way too many options that IMO are clearly not what I want, though at some level it's more pandas' fault (too many methods & attributes even before attaching every column name as an attribute) than VSCode or Intellisense.
- ogogmad 5y agoIn addition to the sibling answers: If you know the class of an object, let's say Class, but the object has yet to be "constructed" so that IPython can correctly infer its type, you can type `Class.[TAB]` in IPython and look at its methods. For example, in Sympy, you have a matrix type called Matrix. You can do `(A * B).diagonalize()`, or alternatively you can do `Matrix.diagonalize(A * B)`, which has some advantages because doing `(A * B).[TAB]` does nothing useful because Python can't infer types. You can also do the same for modules. `ModuleName.[TAB]` To be honest though, I found the experience smoother in R for some reason.
- sntscy 5y agoCan't get around resetting the index, as far as I know, but for the filtering you can also do, `table.query("column > 2 and column2 < 3")`
- claytonjy 5y agodo any IDEs help you autocomplete the string argument? that's a bit part of the non-standard evaluation magic in the tidyverse these days; you can use bare, _unquoted_ names, _and_ get excellent autocomplete, at least in RStudio.
- _dain_ 5y ago.loc lets you supply a callable, so you can write: table.loc[lambda df: df["column"].between(2, 3, inclusive="neither")] this is useful when your dataframe has a long name, or when you have some long method chain and you need to subset at the end: table.foo().bar().baz().loc[lambda df: ...] it is still more verbose, but I actually prefer always providing column names as strings. it's more explicit. I don't like R's environment-manipulation metaprogramming magic where you can give column names as symbols. as for resetting the index all the time, this is something of an antipattern. if you set up your index right beforehand it isn't necessary so often.
- qudat 5y agoMy wife is a researcher and started delving into doing her own statistical analysis. It's be fun (and frustrating) learning R with her. I agree that dplyr and the tidyverse are some fantastic packges for a software engineer who thinks about spreadsheets as SQL tables. I would say the most frustrating part about RStudio is that it is a workbook where you can execute code based on your cursor. For my wife, these workbooks become a total mess because things aren't necessarily run sequentially.
- dash2 5y agoHere are two good strategies: 1. Always run from the top, using the "run previous chunks" button. When this gets too slow, you know that it's time to think harder about your workflow. For a more extreme version of the same idea, regularly restart R using Ctrl-Shift-0, and run from the top. It'll ensure your code is working right. 2. Have a setup chunk that always gets you to the same state. Make sure every other chunk works directly after calling the setup chunk. Then just alternate between "run setup chunk" and "run current chunk".
- da39a3ee 5y ago0. Don’t use notebooks. Neither RStudio nor Jupyter. They prevent people from developing good programming skills and good version control skills, and they encourage making a huge mess.
- dash2 5y agoI mean that's fair, but if you have to write an academic paper, your options are more or less use a notebook, or copy and paste your results. Of course, the question is how much code you should write in the notebook, versus having it in a more organized set of functions and libraries. It's very easy to end up with a huge bloated document which contains thousands of lines of spaghetti.
- da39a3ee 5y agoI absolutely agree with this: > the question is how much code you should write in the notebook, versus having it in a more organized set of functions and libraries. It's very easy to end up with a huge bloated document which contains thousands of lines of spaghetti. But this isn't correct: > your options are more or less use a notebook, or copy and paste your results. What's wrong with writing scripts that write images to disk? That's how millions of academic papers were written before the advent of notebooks. You could use Makefiles if you like, or you could even use a technology such as Sweave to automatically mix images with LaTeX output. I mean this as politely as possible but the fact that you think that the options are "use a notebook or copy and paste" I think shows that you've caught a notebook mentality disease! The fundamental point I'm trying to make is that you we don't need to do everything interactively from REPLs. REPLs are great for trying things out, but when it comes to producing the images for your paper, those should be produced by scripts, not by commands entered into a REPL, or notebook. An those scripts should evolve via version control, which is the basis of evolving any good and correct software. And the scripts for producing images for a paper should be good and correct software.
- peatmoss 5y ago> best graphing calculator attached to an OK scheme. I discovered "How To Design Programs" somewhere late in my first year of using R. Like most beginning R coders with nominal experience in other languages, I wrote a lot of monolithic scripts in a very imperative style. HtDP gave me a mental framework for decomposing larger problems into bite-sized chunks. The lispy roots of R lent itself particularly well to the model of thinking presented in that book. Ever since then, I've pined for the graphing calculator parts in a more modern Scheme. When ggplot and then the tidyverse (neé hadleyverse) came on the scene, I was even more convinced that Scheme, especially Racket, was the ideal future for data science. If R could support a large ecosystem like tidyverse, just imagine what the metaprogramming facilities of Racket could do! But I think those graphing calculator parts are hard to reproduce. Attempts to clone ggplot2 fall short year after year, because most other languages don't have grid graphics to build on top of. R is a deep ecosystem on "an OK scheme," which is damned hard to beat. Aside: my first year with R, was in an urban planning masters program and I was terrified of my first big kid statistics course (taught in SPSS). I decided I'd give myself bonus work by learning R. While it was absurd to be doing my stats homework in SPSS, then R, then reviewing HtDP on top of the rest of my course load, I did ace that stats course. :-)
- jdougan 5y agoThere is an better alternate universe where xlisp-stat doesn't fall behind and S doesn't happen.
- peatmoss 5y agoInterestingly, there appears to be an attempted reboot: https://lisp-stat.dev/ https://lisp-stat.dev/ My first reaction, is "why not on a modern Scheme as opposed to Common Lisp" and in so thinking, I have demonstrated exactly why no lisp / scheme has ever achieved critical mass :-)
- jdougan 5y agoObvious answer: 3rd party library support. Same thing that makes R ubiquitous.
- hadley 5y agoI love the phrase "eviscerate a fresh data set" :D
- stadeschuldt 5y agoGreat stuff. Could you maybe share a bit of these "3 dozen" scripts? This could be super helpful.
- ekianjo 5y agoLooks like your notebooks focus on crypto currencies :-) Good use case for data analysis!
- delusional 5y ago> R is a highly optimized, well-oiled machine if you're using it for its highly-optimized, well-oiled purposes. This hits home for me. We are just starting to use R for risk modeling where I work. R, more than any language I've ever used, makes me appreciate "worse is better". From a theoretical "aesthetic" perspective R is a mess. Yet for data processing all those theoretical concerns don't matter. It just works. It's honestly kind of humbling that something so theoretically messy can be so practically coherent. It makes me question my assumptions about simplicity.
- bachmeier 5y agoI hadn't thought about R as a "worse is better" language, but that's a good way to think about it. Makes sense, too, since it came from the place that inspired worse is better.
- dash2 5y agoR comes from New Zealand, no?
- bachmeier 5y agoR is an implementation of S. John Chambers worked on it at Bell Labs starting in 1975. https://en.wikipedia.org/wiki/S_%28programming_language%29 https://en.wikipedia.org/wiki/S_%28programming_language%29
- edgyquant 5y agoI’m trying to figure out if you were actually asking a question or if it was rhetorical and you were calling shots.
- deehouie 5y agoBravo! This is exactly right.
- canjobear 5y agoR "just works" now because a huge amount of effort has gone into improving the language over the last 10 or so years, in part spurred by the tidyverse movement, although not restricted in scope to tidyverse. When I was starting grad school around 2010, if someone sent you some R code, the chances that you would be able to "just run" it were basically zero: there would be weird version mismatches in how functions worked, file paths would be specified in inconsistent ways in different parts of the script, all kinds of crazy impenetrable errors were the norm. Now there are several R code snippets posted in these HN comments that will run without trouble. If I could have gone back in time and told myself that this is how R would develop, I would have been shocked (and happy).
- deleted 5y ago[deleted]
- dm319 5y agoThis is just a quick example - I would be grateful if people could recreate this brief look at UK COVID figures in another language: library(tidyverse) library(scales) download.file(url = "https://api.coronavirus.data.gov.uk/v2/data?areaType=overview&metric=covidOccupiedMVBeds&metric=newAdmissions&metric=newCasesBySpecimenDate&metric=newDeaths28DaysByDeathDate&metric=newPeopleReceivingFirstDose&format=csv", destfile = "./data.csv", method = "wget") read_csv("./data.csv") %>% pivot_longer(names_to = "Data", cols = c(newCasesBySpecimenDate, covidOccupiedMVBeds, newAdmissions, newDeaths28DaysByDeathDate)) %>% mutate(Data = factor(Data)) %>% mutate(Data = recode_factor(Data, newCasesBySpecimenDate = "New Cases", newAdmissions = "Admissions", newDeaths28DaysByDeathDate = "Deaths", covidOccupiedMVBeds = "Ventilated")) %>% ggplot(aes(y = value, x = date, colour = Data))+ geom_point(size = 1, colour = "gray", alpha = 0.6)+ geom_smooth(type = "LOESS", span = 0.1)+ labs(y = "Daily rate", x = "Date", colour = "UK COVID-19")+ scale_x_date(date_breaks = "months", date_labels = "%b-%y")+ scale_y_log10(labels = comma(10 ^ (0:5), accuracy = 1), breaks = 10 ^ (0:5))+ theme(axis.text.x = element_text(angle = 45, hjust = 1))
- azalemeth 5y agoYou don't need the download file command -- read CSV works with URLs directly :-)
- dm319 5y agoOh nice, that's very useful to know!
- nuclearnice1 5y agoI think these kind of common task challenges are great for comparison. You see a few approaches to a single task and you can do a more aligned and detailed comparison. Unfortunately, it’s also a lot of work. In this case, you’ve posted an intermediate stage artifact from R. If one of the many Python programmers reading this want to produce a comparable artifact they need to understand or run that code. That alone reduces your likelihood of getting any substantial replies. Maybe add a link to an image of the resulting plot?