4 ms·
It's also worth noting that R becomes much more pleasurable with the Tidyverse libraries. The pipe alone makes everything more readable. I'm also coming from m
by bokstavkjeks 8y ago
It's also worth noting that R becomes much more pleasurable with the Tidyverse libraries. The pipe alone makes everything more readable.
I'm also coming from more of an office setting where everything is in Excel. I've used R to reorganize and tidy up Excel files a lot. Ggplot2 (part of the Tidyverse) is also fantastic for plotting, the grammar of graphics makes it really easy to make nice and slightly complex graphs. Compared to my Matplotlib experiences, it's night and day. Though I'd expect my experience with programming to be quite different from others' though, mainly because any code I write is basically an intermediary step before the output goes back in Excel.
That said, if anyone's interested in learning R from a beginner's level, I can recommend the book R for Data Science. It's available freely at http://r4ds.had.co.nz/ http://r4ds.had.co.nz/ and the author also wrote ggplot2, RStudio, and several of the other Tidyverse libraries.
EDIT: I'm also currently writing my master's thesis in RMarkdown with the Thesisdown package. It's wonderful, it allows for using Latex without really knowing Latex which is great for us in business school.
- riskneutral 8y agoTidy features (like pipes) are detrimental to performance. The best things R has going for it are data.table, ggplot, stringr, RMarkdown, RStudio, and the massive, unmatched breadth and depth of special-purpose statistics libraries. Combined, this is a formidable and highly performant toolset for data analytics workflows, and I can say with some certainty that even though “base Python” might look prettier than “base R,” the combination of Python and NumPy is not necessarily more powerful or even a more elegant syntax. The data.table syntax is quite convenient and powerful, even if it does not produce the same “warm fuzzy” feeling that pipes might. NumPy syntax is just as clunky as anything in R, if not worse, largely because NumPy was not part of the base Python design (as opposed to languages like R and MATLAB that were designed for data frames and matrices). What is probably not a good idea (which the article unfortunately does) is to introduce people to R by talking about data.frame without mentioning data.table. Just as an example, the article mentions read.table, which is a very old R function which will be very slow on large files. The right answer is to use fread and data.table, and if you are new to R then get the hangs of these early on so that you don’t waste a lot of time using older, essentially obsolete parts of the language.
- curiousgal 8y agoFrom your experience what makes data.table so useful?
- vijucat 8y agoAnswering questions in a rapid, interactive way (, while using C to be efficient enough that one can run it on millions of rows): # Given a dataset that looks like this… > head(dt, 3) mpg cyl disp hp drat wt qsec vs am gear carb name 1: 21.0 6 160 110 3.90 2.620 16.46 0 1 4 4 Mazda RX4 2: 21.0 6 160 110 3.90 2.875 17.02 0 1 4 4 Mazda RX4 Wag 3: 22.8 4 108 93 3.85 2.320 18.61 1 1 4 1 Datsun 710 # What's the mean hp and wt by number of carburettors? > dt[, list(mean(hp), mean(wt)), by=carb] carb V1 V2 1: 4 187.0 3.8974 2: 1 86.0 2.4900 3: 2 117.2 2.8628 4: 3 180.0 3.8600 5: 6 175.0 2.7700 6: 8 335.0 3.5700 # How many Mercs are there and what's their median hp? > dt[grepl('Merc', name), list(.N, median(hp))] N V2 1: 7 123 # Non-Mercs? > dt[!grepl('Merc', name), list(.N, median(hp))] N V2 1: 25 113 # N observations and avg hp and wt per {num. cylinders and num. carburettors} > dcast(dt, cyl + carb ~ ., value.var=c("hp", "wt"), fun.aggregate=list(mean, length)) cyl carb hp_mean wt_mean hp_length wt_length 1: 4 1 77.4 2.151000 5 5 2: 4 2 87.0 2.398000 6 6 3: 6 1 107.5 3.337500 2 2 4: 6 4 116.5 3.093750 4 4 5: 6 6 175.0 2.770000 1 1 6: 8 2 162.5 3.560000 4 4 7: 8 3 180.0 3.860000 3 3 8: 8 4 234.0 4.433167 6 6 9: 8 8 335.0 3.570000 1 1 I used slightly verbose syntax so that it is (hopefully) clear even to non-R users. You can see that the interactivity is great at helping you compose answers step-by-step, molding the data as you go, especially when you combine with tools like plot.ly to also visualize results.
- extr 8y agoCompletely agree. dplyr is nice enough but the verbose style gets old fast when you're trying to use it in an interactive fashion. imo data.table is the fastest way to explore data across any language, period.