4 ms·
I love R, and it was actually the first language I really learned to program with (for obvious reasons, I wouldn't ever recommend this). I can identify a lot w
by vikp 13y ago
I love R, and it was actually the first language I really learned to program with (for obvious reasons, I wouldn't ever recommend this). I can identify a lot with the "one-off scripts" problem of R. When I look back at some of my R code, I find that it is just a giant mess of commands mixed together semi-randomly.
As I learned to program properly, I solved the reuse and "good code" problem more by moving towards using Python than by making R packages. I occasionally use R for data exploration and visualization, but Python has support for almost all of the machine learning and statistical functions that I need.
I am very interested to hear how other people have solved this problem. Do you only use R, or R in combination with Python/Julia/Java? If you only use R, are you using it purely academically?
- RA_Fisher 13y agoI'm just now getting into Python. Primarily because I'm abandoning Fisherian methods for Bayesian methods and all of the books seem to have examples in Python. I use R in a business setting. I'm a data scientist at Treehouse (teamtreehouse.com)
- jasonpbecker 13y agoThere are great tools that interface R with BUGS that may be worth checking out. I would highly recommend getting this for the shelf: http://www.amazon.com/Analysis-Regression-Multilevel-Hierarchical-Models/dp/052168689X http://www.amazon.com/Analysis-Regression-Multilevel-Hierarc... It's one of the most readable books on data analysis I've come across and does a great job presenting both frequentist and Bayesian techniques with tons of R sample code. There are a lot of advantages and nice things in Python, but I do think folks tend to toss out R a bit too casually. Each tool has areas they excel in. I don't even do particularly complex analysis, but have run into areas where Python is woefully lacking in fairly common (social science) models.
- RA_Fisher 13y agoThanks! I'm planning on digging into Cam Davidson-Pilon's Bayesian Methods for Hackers soon:https://github.com/CamDavidsonPilon/Probabilistic-Programming-and-Bayesian-Methods-for- https://github.com/CamDavidsonPilon/Probabilistic-Programmin... Thanks for the link to the book. I could see it being really helpful to me to have frequentist examples alongside Bayesian examples. I'm hoping to learn Python so I have both tools. R is so domain-specific that I doubt I'll ever completely stop using it in my career. Who knows, though!?
- chubot 13y agoIMO the Unix philosophy of reuse is vital for data analysis. I recommend that people in this space learn to use the interactive shell and shell scripts well. I saw a comment that said "Shell is a REPL for C", which I think is quite pithy. People think too much about reusing R packages or Python packages. In my experience you can get things done a lot faster if you factor things into programs and not libraries. This lets you use multiple languages, and all real world analysis problems need multiple languages. (If you're only using one language, then you're likely only working on part of the problem). Right my toolset is Python, C++, and R, coordinated with shell scripts. I still like R for quick plotting, and of course data frames are essential, but I'm playing with Pandas now, which seems impressive. I don't think Python will ever catch up to R in terms of statistical functions, but in terms of plotting/munging/data frames it might. In the future there will be more languages, not less (Julia will add to the number of languages, not replace any). So being able to decompose a problem into separate, reasonably generic, programs is an important skill IMO. Most data analysis pipelines are a huge mess, but they don't have to be.
- vikp 13y agoThat's an interesting way to go about it. Any reason for using shell scripts to coordinate the flow instead of things like Cython and RPy? (I don't shell script a lot, so this may be a silly question) These days, I mostly seem to be able to get away with using single languages for applied machine learning, like a Python webserver that runs background machine learning tasks, or an Android app in Java that connects to a webserver written in Python. But back when I was doing less applied side stuff, I was similar to you in that things were spread across R and Python.
- LukeShu 13y agoThe Unix Shell is basically an IPC language. It's sole purpose is to be able to combine programs, and allow them to communicate. Instead of trying to make a package in one language work with another, it is a lot simpler to use Shell to connect together programs in any language.
- chubot 13y ago
- vdaniuk 13y agoShould one learn R if one already has some experience in data analysis with Python?
- collyw 13y agoI would like to know the same. I work with a lot of bioinformaticians. The ones from a biology background (i.e. code quality is not usually important) seem to like R. The ones with a computer science background seem to go for Python, or Perl if they are old school. So far that has put me off R.
- hadley 13y agoI think R is actually a rather great first language, as the focus is usually on doing cool stuff, rather than developing a rigorous framework for programming. It's more important to put motivation before rigour, since if you focus only on rigour most people will lose interest before getting to the fun stuff.
- vikp 13y agoGreat point. I still wish I coded in R more, just because the MTTC (mean time to cool) is so much lower. But, I feel like you can do the same with sklearn/ipython notebook/pandas if you ignore the parts of the "rigorous framework", and just focus on the syntax for interesting things.
- oddthink 13y agoI'm slowly moving my elaborate data-cleaning system from R to python. I never got the hang of R packages and ended up rolling my own "import" scheme, but now it all seems like I've passed some critical complexity point. I'm still doing all of my model estimation in R, and data.table + ggplot2 is still my go-to solution for interactive exploration and plotting, but I'm putting more and more of the process-csv, write-hdf5, remove out-of-bounds values, backfill, stratified sampling/bootstrap, etc., logic into python. I may end up moving more of the estimation itself into python, now that statsmodels and sklearn are more mature, but R does that pretty well. I've not tried Julia, but I have experimented with Java for some of the data-processing, but it's never seemed worth it, over python. Basic CSV processing is faster in python than in Java (although the python version uses much more CPU), so I haven't felt the need. (I'm probably limited more by disk bandwidth than CPU, but I've not done the tests required to prove that.) Julia seems nice, but it seems intent on copying all the questionable design choices of matlab, rather than using ideas from the numpy/kdb+/apl worlds.