15 ms·
Comparison – R vs. Python: head to head data analysis
- acaloiar 11y agoI have always considered R the best tool for both simple and complex analytics. But, it should not go unmentioned that the features responsible for R's usability often manifest as poor performance. As a result, I have some experience rewriting the underlying C code in other languages. What one finds under the hood is not often pretty. It would be interesting to see a performance comparison between Python and R.
- pjmlp 11y agoGiven that R folks are porting it to the JVM, I guess performance on the R side will improve thanks to Hotspot and Graal/Truffle. http://www.renjin.org/ http://www.renjin.org/ http://www.oracle.com/technetwork/java/jvmls2013vitek-2013524.pdf http://www.oracle.com/technetwork/java/jvmls2013vitek-201352... Then there is PyPy as well. I also think they should probably add Julia and Wolfram/Mathematica to these comparisons.
- jtth 11y agoI would say they're both as limited as Python, Julia far more so. R's stats packages get ported to Julia faster, though. Mathematica still can't do mixed generalized linear modeling, and no other language (other than SAS and Stata) has a package for analyzing simple effects within them.
- pjmlp 11y agoThanks for the overview, I don't use them. It is more my language geek side speaking louder. :)
- acaloiar 11y agoI have found Renjin quite useful in the past, and I love the motivation behind the project. I know that the guys at Bedatadriven hope to improve upon its performance, however it does not always (or often, depending on how you use R) outperform GNU R. Some great changes have been made lately (http://www.renjin.org/blog/2015-06-28-renjin-at-rsummit-2015.html http://www.renjin.org/blog/2015-06-28-renjin-at-rsummit-2015...), so I hope to see Renjin's performance progress beyond GNU R across the board. I actually contributed Renjin's current PRNG – a Java translation of GNU R's – which was my first experience getting under R's hood. The Purdue project you linked looks quite interesting. Unfortunately, development appears to have stagnated: https://github.com/allr/purdue-fastr https://github.com/allr/purdue-fastr [edit] Another important aspect that Renjin contributes is the packages ecosystem: http://packages.renjin.org/ http://packages.renjin.org/
- huac 11y agoR being single-threaded internally may also result in performance hits.
- Mikeb85 11y agoR also has tools to spread tasks over multiple cores or over a cluster quite effortlessly. In practice, I can create a Fortran or C++ module, then use R to apply it over multiple cores, and get fantastic performance for certain tasks.
- sweezyjeezy 11y agoThis is just a series of incredibly generic operations on an already cleaned dataset in csv format. In reality, you probably need to retrieve and clean the dataset yourself from, say, a database, and you you may well need to do something non-standard with the data, which needs an external library with good documentation. Python is better equipped in both regards. Not to mention, if you're building this into any sort of product rather than just exploring, R is a bad choice. Disclaimer, I learned R before Python, and won't go back.
- The13thDoc 11y agoI agree. Once you incorporate the other necessary work and preparation, a well-documented, object oriented language is a better way to go.
- vegabook 11y agoI have to agree that Python is more powerful, and I am indeed doing more and more in Python. Python was my first language, before R. However when the dataset is medium sized (i.e.: fits into your computer's memory / 2) R crushes Python (and Pandas) for the 80% of the time you'll be spending wrangling. The reason is that R is vector-based from the ground up. Pandas does everything that R does, but does it in a less-consistent, grafted-on way, whereas the experienced R person who "thinks vectors" is way ahead of the Python guy before the analysis has even started (i.e., most of the work). I know both really well. I use Python when I want to "get (semi) serious" production wise (I qualify with "semi" because if you're really serious about production, you're probably going to go to Scala). But when it comes to taking a big chunk of untidy data and bashing it around till it's clean and cube-shaped, will parse, and has no no obvious errors, R is miles ahead of Python. R is where you do your discovering. Python can do it too, but I would estimate the cognitive overhead as double. By the way, that's why people who "think time series" all day long (i.e., vectors, not objects), and who want to implement their algos, not think CS, will first typically build it in R, which is why CRAN beats Python all the time and every time for off-the-shelf data analysis packages. Data people go to R, computer-people go to Python (schematizing). R is slow. That's its main problem. And that's saying something when comparing it to Python! But the gem of vector-everything makes it a much more satisfying language than imperative, OO, Python, when it comes to the world of data first, code second. Finally I'd add that Python 3.x is arguably distancing itself from the pragmatism which data science requires, and 2.x provided, towards a world of CS purity. It's not moving in a direction which is data science friendly. It's moving towards a world of competition with Golang and Javascript, and Java itself.
- zitterbewegung 11y agoThis is not just interesting for comparison but its interesting for people that know R/Python how to go from one to the other.
- jtth 11y agoKind of, but the R code is written a little oddly to my eye.
- CJKinni 11y agoHow so? As someone familiar with Python but not R, I've always been hesitant to jump in. This code was very readable and made me think that it might be a far more accessible language than I'd previously assumed.
- geomark 11y agoOne example in the section titled "Split into training and testing sets" would be to use the createDataPartition() function from the caret package for creating training and testing sets. He says "In R, there are packages to make sampling simpler, but aren’t much more concise than using the built-in sample function" but using caret is more concise. Added: Later in the section on random forests he says "With R, there are many smaller packages containing individual algorithms, often with inconsistent ways to access them." Which is why you want to use the caret package as it makes accessing many machine learning packages consistent and easy.
- Pinatubo 11y agoMe too. Why for example did they use an sapply for column means when they could have just used colMeans with na.rm=T?
- compbio 11y agoThat is a major difference between these two languages. Python: There should be one, and only one, preferable way to do things. Though this may not be obvious at first. R: Every author has a different style of doing things, reflecting in the code. As for the comparison in general: You can call R from within Python. So Python is at least as powerful as R. The rest (BeautifulSoup, Compression, Game development etc.) is icing on the cake.
- deleted 11y ago[deleted]
- acomjean 11y agoI work with biologists. R which seems strange to me they seem to take to. I think some of it is Rstudio the ide, which shows variables in memory on the side bar, you can click to see them. It makes everything really accessible for those that aren't programmers. It seems to replace excel use for generating plots. I've grown to appreciate R, especially its plotting ability (ggplot).
- mbreese 11y agoRstudio is R for a lot of people. I'm a computational biologist in a group. Our PI is trying to get the postdocs to learn R themselves, but it's an uphill battle. I eventually warmed up to it - primarily for the plotting. But a few weeks back he asked me how to do some kind of data sorting / manipulation in R. My answer was that it was a 10 line Python script and I gave him the code. Alas, he couldn't figure out how to save the script and run it from a command-line. You can't underestimate at how important Rstudio is to the popularity of R for non-programmers.
- collyw 11y agoIt amazes me how few biologists / bioinformticians use an IDE.
- kyllo 11y agoI think some of it is Rstudio the ide, which shows variables in memory on the side bar, you can click to see them This. Most programming IDEs show the code but hide the data. Excel shows the data but hides the code. RStudio is awesome because it shows both the code and the data.
- mbreese 11y agoThis is interesting, but not really an R vs. Python comparison. It's an R vs. Pandas/Numpy comparison. For basic (or even advanced) stats, R wins hands down. And it's really hard to beat ggplot. And CRAN is much better for finding other statistical or data analysis packages. But when you start having to massage the data in the language (database lookups, integrating datasets, more complicated logic), Python is the better "general-purpose" language. It is a pretty steep learning curve to grok the R internal data representations and how things work. The better part of this comparison, in my opinion, is how to perform similar tasks in each language. It would be more beneficial to have a comparison of here is where Python/Pandas is good, here is where R is better, and how to switch between them. Another way of saying this is figuring out when something is too hard in R and it's time to flip to Python for a while...
- halflings 11y agoR has pandas/numpy/scipy integrated in the language (for the most used features at least), but that doesn't make much of a difference because any person that wants to use these tools will do a quick "pip install" to grab them. (which is pretty fast with the new Wheels system) Out of curiosity, why do you consider CRAN to be much better than PyPI?
- mbreese 11y agoI'm only thinking about CRAN > PyPI in terms of statistical packages. CRAN is where new statistical analysis techniques / packages are initially published. If you're lucky they might get ported to Python after the fact. I didn't even mention Bioconductor, which is another beast entirely. There isn't an equivalent of Bioconductor for Python at all. And the last time I checked, "pip install numpy" could be quite a pain, especially if you needed to compile dependencies. Rstudio makes it ridiculously easy to install R and add packages. However - for all other types of packages, PyPI is obviously superior. The breadth of packages on PyPI is much better than CRAN. It about choosing the right tool for the job.
- sandGorgon 11y ago
- k8tte 11y agoi tried help my wife who use R in school, only to get quickly lost. also attended ~1 hour R course on university. to me, R was a waste of time and I really dont understand why its so popular in academia. if you already have some programming knowledge, go with Python + Scipy instead EDIT: R is even more useless without r studio, http://www.rstudio.com/ http://www.rstudio.com/. and NO, dont go build a website in R!
- untothebreach 11y agoMaybe you didn't mean it this way, but to me your comment reads as, basically, "I tried R for an hour and didn't immediately grok it, therefore it is a waste of time." That may not be what you meant, so I haven't downvoted yet, but it doesn't seem to be an attitude that is helpful for the conversation.
- k8tte 11y agoThanks for your explanation. It seems my ability to communicate is getting worse every year :-/. What I meant to say was that I helped my wife during her master thesis (~6 months) with R, in addition to spending an hour in one of the classes. Her teachers also were novices of both R and Excel, and we had several issues with everything from how R processes csv:s, to just figuring out the proper syntax to have R do what we wanted. Sorry if my comment wasnt helpful, i was merely attempting to add some reflections from personal experience to the discussion.
- evandev 11y agoI disagree with R being more useless without r studio. I'm not a fan of R overall, but I run everything in tmux+vim and R is the same way. I prefer it to Rstudio. It's popular, because it makes a few choices which are different many programming languages to be geared towards writing scripts for statistics. (e.g. index 0, assignments)
- peatmoss 11y agoI'll second the utility of alternative environments to RStudio. For me, I love RStudio, but I spend too much time in Python (and occasionally dabbling in others) to use it all the time. So, for me it's Emacs Speaks Statistics, which is fantastic. As a side benefit, the first time I tried dabbling in Julia, I was pleasantly surprised to have a familiar mature environment work with it out of the box.
- The13thDoc 11y agoThe "cheat sheet" comparison between R and Python is helpful. The presentation is well done. The conclusions state what we already know: Python is object oriented; R is functional. The Last Word appropriately tells us your opinion that Python is stronger in more areas.
- danso 11y agoI spent a few weeks a few months ago learning R. It's not a bad language, and yes, the plotting is currently second-to-none, at least based on my limited experience with matplotlib and seaborn. There's scant few articles on going from Python to R...and I think that has given me a lot of reason to hesitate. One of the big assets of R is Hadley Wickham...the amount and variety of work he has contributed is prodigious (not just ggplot2, but everything from data cleaning, web scraping, dev tools, time-handling a la moment.js, and books). But that's not just evidence of how generous and talented Wickham is, but how relatively little dev support there is in R. If something breaks in ggplot2 -- or any of the many libraries he's involved in, he's often the one to respond to the ticket. He's only one person. There are many talented developers in R but it's not quite a deep open-source ecosystem and community yet. Also word-of-warning: ggplot2 (as of 2014[1]) is in maintenance mode and Wickham is focused on ggvis, which will be a web visualization library. I don't know if there has been much talk about non-Hadley-Wickham people taking over ggplot2 and expanding it...it seems more that people are content to follow him into ggvis, even though a static viz library is still very valuable. [1] https://groups.google.com/forum/#!topic/ggplot2/SSxt8B8QLfo/discussion https://groups.google.com/forum/#!topic/ggplot2/SSxt8B8QLfo/...
- revorad 11y agoHadley is actively working on ggplot2. In fact, he just tweeted a list of improvements - https://twitter.com/hadleywickham/status/654283936755904512 https://twitter.com/hadleywickham/status/654283936755904512 https://github.com/hadley/ggplot2/blob/master/NEWS.md https://github.com/hadley/ggplot2/blob/master/NEWS.md
- danso 11y agoThanks...I didn't know that (though I had been paying attention to bug fixes)...but my point exactly, he's prodigious, so maybe "maintenance mode" to him is "major features every 3 months instead of 2) :). Also worth pointing out, he's actively working on a new book for ggplot2, which, AFAICT, he's providing for free (you just have to run the build tools) https://github.com/hadley/ggplot2-book https://github.com/hadley/ggplot2-book I think if someone were to run an analysis of Wickham's Github activity, it would produce a freakishly busy chart.
- willpearse 11y agoVery picky, but beware constantly using "set.seed" throughout your R scripts. Always using the same random number is not necessarily helpful for stats, and makes the R code look a lot trickier than it need be
- xname2 11y ago"data analysis" means differently in R and Python. In R, it's all kinds of statistical analyses. In Python, it's basic statistical analysis plus data mining stuff. There are too many statistical analyses only exist in R.
- fsiefken 11y agoIt would be nice to compare JuliaStats and Clojure based Incanter with Python Pandas/NumPy/SciPy. http://juliastats.github.io/ http://juliastats.github.io/
- mojoe 11y agoThe one thing that sometimes gets overlooked when people decide whether to use R or Python is how robust the language and libraries are. I've programmed professionally in both, and R is really bad for production environments. The packages (and even language internals sometimes) break fairly often for certain use cases, and doing regression testing on R is not as easy as Python. If you're doing one-off analyses, R is great -- for anything else I'd recommend Python/Pandas/Scikit.
- vegabook 11y agoor Scala, Clojure, or indeed C. R's great strength is finding the interesting bits of the data. Testing the Algo. Doing the R&D basically. Better than Python. Once that's done, why stop at Python? If your game is production, Python will do it, but others will do it so much better, faster, more efficiently.
- CuriouslyC 11y agoOne nice thing about Python is that you can make a piecewise transition from Python -> C, as it is fairly trivial to wrap C code for use in Python. On the other hand, Java's C interface system JNI is pretty much universally reviled.
- tareef 11y agoThe same can be said about R. Rcpp makes it super easy for you to drop right into C++ for bits of code that need that level of performance.
- nrpprn 11y agoYou can beat scala and approach c in python, with python syntax, using numba. It compiles numerical python code.
- vegabook 11y agoGood point, but personally I am thinking about the future of clustered data analysis, and this seems to be a JVM world and Scala seems to be the language of choice. Flink / Storm / Spark etc.
- bigtunacan 11y agoR is certainly a unique language, but when it comes to statistics I haven't seen anything else that compares. Often I see this R vs Python comparison being made (not that this particular article has that slant) as a come drink the Python kool-aid; it tastes better. Yes; Python is a better general purpose language. It is inferior though when it comes specifically to statistical analysis. Personally I don't even try to use R as a general purpose language. I use it for data processing, statistics, and static visualizations. If I want dynamic visualizations I process in R then typically do a hand off to JavaScript and use D3. Another clear advantage of R is that it is embedded into so many other tools. Ruby, C++, Java, Postgres, SQL Server (2016); I'm sure there are others.
- dagw 11y agoUsing R from within Python works pretty well for all those unique R packages which don't have a python equivalence.
- bigtunacan 11y agoThanks; I suspected support for embedding R within Python already existed as well, but I wasn't sure about that one.
- dagw 11y agoRpy2 is the python library you probably want: http://rpy2.readthedocs.org/ http://rpy2.readthedocs.org/
- devty 11y agoCould you provide an example in stat analysis where python is clearly inferior? In the article, R seems to have an advantage of having many useful stat functions baked in vs having to import specific modules in python. im wondering if your proficiency in R is being weighed in your evaluation of R - maybe python's statistical analysis tool has many to offer, but you are more aware of R's toolsets.
- 11y ago
- vegabook 11y agoPython's main problem is that it's moving in a CS direction and not a data science direction. The "weekend hack" that was Python, a philosophy carried into 2.x, made it a supremely pragmatic language, which the data scientists love. They want to think algorithms and maths. The language must not get in the way. 3.x is wanting to be serious. It wants to take on Golang. Javascript, Java. It wants to be taken seriously. Enterprise and Web. There is nothing in 3.x for data scientists other than the fig leaf of the @ operator. It's more complicated to do simple stuff in 3.x. It's more robust from a theoretical point of view, maybe, but it also imposes a cognitive overhead for those people whose minds are already FULL of their algo problems and just want to get from a -> b as easily as possible, without CS purity or implementation elegance putting up barriers to pragmatism (I give you Unicode v Ascii, print() v print, xrange v range, 01 v 1 (the first is an error in 3.x. Why exactly?), focus on concurrency not raw parallelism, the list goes on). R wants to get things done, and is vectors first. Vectors are what big data typically is all about (if not matrices and tensors). It's an order of magnitude higher dimensionality in the default, canonical data structure. Applies and indexing in R, vector-wise, feels natural. Numpy makes a good effort, but must still operate in a scalar/OO world of its host language, and inconsistencies inevitably creep in, even in Pandas. As a final point, I'll suggest that R is much closer to the vectorised future, and that even if it is tragically slow, it will train your mind in the first steps towards "thinking parallel".
- evanpw 11y agoIf you only have time to learn one language, learn Python, because it's better for non-statistical purposes (I don't think that's very controversial). If you need cutting-edge or esoteric statistics, use R. If it exists, there is an R implementation, but the major Python packages really only cover the most popular techniques. If neither of those apply, it's mostly a matter of taste which one you use, and they interact pretty well with each other anyway.
- xixi77 11y agoI'd say, if most of your job is analyzing the data yourself and trying to make sense of it, R wins hands down. Particularly if statistical graphics or advanced statistical methods may be needed, but it's still the case even if they won't. If most of your job is going to be implementing data analysis techniques that you or someone else has done earlier and putting things into production, then Python will quite possibly be more suitable.
- blumkvist 11y agoR does not mean only esoteric statistics. You have many more utilities in the R packages to diagnose and select models. Fitting a model is like 1% of the work, diagnostic is the more important part and R has much more to offer than Python ever will.
- nrpprn 11y agoStatsmodels has tons of model diagnostic...and there is no R equivalent to Pymc3 (stan has less capability and worse API)
- roel_v 11y ago"If you only have time to learn one language, learn Python, because it's better for non-statistical purposes (I don't think that's very controversial)." Actually, it is. When someone has only 3 or 4 years to finish their thesis and learning how to program is secondary at best, and they have to do it in a math-heavy department or field, there is no time or use to learn Python.
- daveorzach 11y agoIn manufacturing Minitab and JMP are used for data analysis (histograms, control charts, DOE analysis, etc.) They are much easier to use and provide helpful tutorials on the actual analysis. What features or workflow does R or Pandas/Numpy offer to manufacturing that Minatab & JMP can't?
- dbbolton 11y agoR, Numpy, and Pandas are all FOSS. Probably not much of a practical concern, but it might be preferable in some cases. I don't know anything about Minitab/JMP scripting myself, but my understanding is that R is generally the most intuitive of all the aforementioned (although that would basically boil down to individual preference). Here's a review including Minitab and R that might be of interest: http://www.prostatservices.com/statistical-consulting/articles-of-interest/a-review-of-the-top-five-statistical-software-systems http://www.prostatservices.com/statistical-consulting/articl...
- falicon 11y agoLanguage comparisons are equiv. to religion comparisons...you aren't going to find a universal answer or truth, it's an individual/faith sort of thing. That being said - all the serious math/data people I know love both R and Python...R for the heavy math, Python for the simplicity, glue, and organization.
- vineet7kumar 11y agoIt would be nice to also have some notes about performance of both the languages for each of the tasks compared. I believe pandas would be faster due to its implementation in C. The last time I checked R was an interpreted language with its interpreter written in R.
- hadley 11y agoAnd like pandas, many of the performance bottlenecks in R have been re-written in C. See dplyr and data.table for packages that solve a similar problem to pandas with similar speed (and for some scenarios they're actually faster!)
- vineet7kumar 11y agoLooks interesting! Thanks for the information.
- xixi77 11y agoReally, syntax "nba.head(1)" is not any more "object-oriented" than "head(nba, 1)" -- it's just syntax, and the R statement is in fact an application of R's object system (there are several of them). IMO, R's system is actually more powerful and intuitive -- e.g. it is fairly straightforward to write a generic function dosomething(x,y) that would dispatch specific code depending on classes of both x and y.
- illumen 11y agoSingle-dispatch generic functions are easy in python too: https://www.python.org/dev/peps/pep-0443/ https://www.python.org/dev/peps/pep-0443/
- xixi77 11y agoThat's good to know, thanks :) Although, for single dispatch, the S3 system of R is kinda hard to beat -- you just name your function print.myclass and you are done :)
- cfv 11y agoWould you please be so kind as to stop comparing a library to a language? It's misleading and I personally tend to take those as attempts to mislead, which might not be your intended message and thus makes me misinterpret your article
- phillipamann 11y agoR is a wonderful language if you chose to get used to it. I love it. I've even used R in production quality assurance to check for regressions in data (not the statistical regressions). I see countless R posts where people try to compare it to Python to find the one true language for working with data. Article after article, there clearly isn't a winner. People like R and Python for different reasons. I think it's actually quite intuitive to think about everything in terms of vectors with R. I like the functional aspects of R. I wish R was a bit faster but I am pretty sure the people who maintain R are working on that. You can't beat the enormous library that R has.
- baldfat 11y agoI also LOVE R. Plus the fact that Microsoft and other corporations are supporting R will help more and more. With Hadly Wickham's universe it is a great place to do all your work.
- Mikeb85 11y agoYup. R is supported by MS, Oracle, IBM and others, and companies like Twitter and even the Python shop that is Google use it.
- wesm 11y agoI hope you all know that the people who have invested most in actually building this software care the least about this discussion.
- Mikeb85 11y agoThe reason I like R - it just makes data exploration and analysis too damn easy. You've got R Studio, which is one of the best environments ever for exploring data, visualisation, and it manages all your R packages, projects, and version control effortlessly. Then you've got the plethora of packages - if you're any of the following fields: statistics, finance, economics, bioinformatics, and probably a few others, there's packages that instantly make your life easier. The environment is perfect for data exploration - it saves all the data in your 'environment', allows you to define multiple environments, and your project can be saved at any point, with all the global data intact. If I want some extra speed, I can create C++ modules from within R Studio, compile and link them, as easily as simply creating a new R script. Fortran is a tiny bit more work, still easy enough however. Want multicore or to spread tasks over a cluster? R has built in functions that do that for you. As easy as calling mcapply, parApply, or clusterApply. Heck, you can even write your function in another language, then R handles applying that over however many cores you want. Want to install and manage packages, update them, create them, etc...? All can be done from R Studio's interface. Knitr can create markdown/HTML/pdf/MS Word files from R markdown, or you can simply compile everything to a 'notebook' style HTML page. And all this is done incredibly easily, all from a single package (R Studio) which itself is easy to get and install. Oh yeah, visualisation, nothing really beats R. And while there are quirks to the language, for non-programmers this isn't really an obstacle, since they aren't already used to any particular paradigm. As for Python, I'm sure it's great (I've used it a little), but I really don't see how it can compare. R's entire environment is geared towards data analysis and exploration, towards interfacing with the compiled languages most used for HPC, and running tasks over the hardware you will most likely be using.
- ggrothendieck 11y agoFor R: (1) instead of `sapply(nba, mean, na.rm = TRUE)` use `colMeans(nba, na.rm = TRUE)`. (2) instead of `nba[, c("ast", "fg", "trb")]` use `nba[c("ast", "fg", "trb")]`, (3) instead of `sum(is.na(col)) == 0` use `!anyNA(col)`, (4) instead of `sample(1:nrow(nba), trainRowCount)` use `sample(nrow(nba), trainRowCount)` and (5) instead of tons of code use `library(XML); readHTMLTable(url, stringsAsFactors = FALSE)`
- hogu 11y agoIF you do your stuff in R, how do you move it into production? Or do you not need to
- Mikeb85 11y agoThere are packages for that (web servers and such). Or you can call it from Java/Python/whatever. Most R tasks that people use exit. Typical data science task is: gather data, apply an operation over said data, analyse results.
- c3534l 11y agoI like Python better as a language, but Python's libraries take more work to understand and the APIs aren't very unified. R is much more regular and the documentation is better. Even complicated and obscure machine learning tasks have good support in R. BUT the performance for R can be very, very annoying. Assignment is slow as all hell and it can often take work to figure out how to rephrase complicated functions in a way that R can figure out how to do efficiently. I think being much more functional than Python works well for data. I mean the L in LISP stands for list! Visualizations are also easier and more intuitive in R, too, IMO. Especially since half the time you can just wrap some data in "plot" and R will figure our which one it should use. I think the conclusion of the article is correct. R is more pleasant for mathier type stuff, while Python is the better general-purpose language. If your jobs involves showing people powerpoint presentations of the mathematical analysis you've done,you'd probably want to use R. If, on the other hand, you're prototyping data-driven applications, Python would probably be better. That said, I really like Julia, but can't justify really diving into it at this point. :\
- baldfat 11y ago> prototyping data-driven applications, Python would probably be better I would disagree. Python's libraries are really reimplementing R in Python (Mainly Pandas). I find R to be very flexible and especially in the last 5 years with Hadley Wickham's libraries things are concise and very powerful. Please look at dplyr and see how this new way fo doing R works. Especially with piping with %>%. https://cran.rstudio.com/web/packages/dplyr/vignettes/introduction.html https://cran.rstudio.com/web/packages/dplyr/vignettes/introd... Code in R can look like this beautiful code (If you don't code in R and I would expect anyone can see what is happening) This is why I disagree that prototyping in Python would be better.: flights %>% group_by(year, month, day) %>% select(arr_delay, dep_delay) summarise( arr = mean(arr_delay, na.rm = TRUE), dep = mean(dep_delay, na.rm = TRUE)) %>% filter(arr > 30 | dep > 30) Python has .pipe but I find it strange it goes to the new line before the items. Python Code: >>> (df.pipe(h) ... .pipe(g, arg1=a) ... .pipe((f, 'arg2'), arg1=a, arg3=c) ... )
- avdempsey 11y ago
- thebelal 11y agoThe rvest implementation was the main thing that seemed like an R port of the python implementation rather than best use of rvest. An alternate (simpler) implementation of the rvest web scraping example is at https://gist.github.com/jimhester/01087e190618cc91a213 https://gist.github.com/jimhester/01087e190618cc91a213 It would be even simpler but basketball-reference designs it's tables for humans rather than for easy scraping.
- baldfat 11y ago>seemed like an R port of the python implementation End of the github for rvest: Inspirations Python: Robobrowser, beautiful soup.
- jkyle 11y agoCaret is a great package for a lot of utility functions and tuning in R. For example, the sampling example can be done using Caret's createDataPartition which maintains the relative distributions of the target classes and is more 'terse'. > data(iris) > library(caret) > data(iris) > idx <- caret::createDataPartition(iris$Species, p = 0.7, list = F) > summary(iris$Species) setosa versicolor virginica 50 50 50 > summary(iris[idx,]$Species) setosa versicolor virginica 35 35 35
- andyjgarcia 11y agoThe comparison is R to Python+pandas. The equivalent comparison should be R+dplyr to Python+pandas. Base R is quite verbose and convoluted compared to using dplyr. Likewise data analysis in Python is painful compared to using pandas.
- dekhn 11y agoIn general, if I have to chose between two languages, one of which was designed specifically for statistics, and one that was more general, I will chose the more general one. R's value is in the implementation of its libraries but there is no technical reason a really OCD person couldn't implement such high quality of libraries in Python.
- Myrmornis 11y agopython < world > csv R < csv > analysis