6 ms·
Does anybody use R in production services or just for exploratory work? It seems that once you figure out a good model in R, its almost always rewritten into e
by vasaulys 10y ago
Does anybody use R in production services or just for exploratory work?
It seems that once you figure out a good model in R, its almost always rewritten into either Scala or Java for real production work.
- 0x001E84EE 10y agoPart of that may stem from R and most (all?) of its libraries being licensed under GPL.
- baldfat 10y ago> It seems that once you figure out a good model in R, its almost always rewritten into either Scala or Java for real production work. I wouldn't say 1% of programs in R written need that speed. I personally use it for small projects (Besides a few Spark side projects) and I am out putting Reports. I really would like someone to show an actual example of this happening in 2016.
- nerdponx 10y agoI do it at my company. I prototype in R, and then end up having to rewrite chunks of it in Python so it can be worked into our application, which right now is exclusively Python. It's not a matter of performance, it's just because it would be an enormous amount of engineering overhead to start calling R from inside the Python app
- baldfat 10y agoThat seems like you could simply use http://jupyter.org/ http://jupyter.org/ and just run the script with R code inline. http://blog.revolutionanalytics.com/2016/01/pipelining-r-python.html http://blog.revolutionanalytics.com/2016/01/pipelining-r-pyt... Also why not just switch to Pandas it really is a pretty close R clone.
- blahi 10y agoHow much experience do you have in statistical computing, out of curiosity?
- nerdponx 10y agoIt has nothing to do with interoperability on my machine. I use notebooks (and Pandas) all the time, and I consider myself fluent in bith R and Python. It's because R is a substantial engineering dependency. As I said, our entire stack is Python and Node. Yes, you can call R from Python using Rpy2, but that's a pro-bono project maintained largely by one person. It's great for casual use, but there is far too much risk to start talking about building critical business code around it.
- baldfat 10y agoSo why not Pandas?
- nerdponx 10y agoPersonal preference. I switch back-and-forth based on the project. R data frames are native and feel native. Pandas data frames are non-native and can be a pain in the ass to work with. That, and there is a lot mpre to the decision than just which data frame implementation I like better.
- kgwgk 10y ago"Pretty close" as long as you stay within the region of common functionality. I wouldn't say it's a clone.
- baldfat 10y agoThat is true. I actually started my journey with Pandas and then switched to R for the ecco-system and zero based for data science drove me nuts. But I do feel that the goal is a clone. "Python has long been great for data munging and preparation, but less so for data analysis and modeling. pandas helps fill this gap, enabling you to carry out your entire data analysis workflow in Python without having to switch to a more domain specific language like R." http://pandas.pydata.org/ http://pandas.pydata.org/
- RA_Fisher 10y agoCheck out opencpu.org, it's an R web api. Really cool stuff.
- nerdponx 10y agoAfaik Bloomberg uses it extensively for internal data visualization tools.
- madenine 10y agodoesn't Bloomberg have a custom, in-house R IDE?
- vegabook 10y agoI have 20k lines of (my own) R code running in production (used intensively by a salesforce of up to 20 people who price bonds with it) and it's an unmitigated nightmare to manage. Slow as crazy. No threading to manage concurrency so constant batch jobs everywhere. Memory hog. On Windows (this is finance), unfortunate fairly frequent crashes. No real time feeds due to the horrible architecture of the interpreter. That said, beautiful charts! Just Say No. It'll sap your mojo. Am moving the whole thing to a blend of C, Python, and a distributed computing framework (thinking of Flink or Concord.io).
- blahi 10y agoThat sounds like bad coders, not that R is bad. Evidenced by: >No threading to manage concurrency R is used in production at EA, Activision, Ebay, Trulia, Google, Microsoft and many, many more. Those are just the ones I've seen give talks about scoring >1TBs regularly with R. Every time somebody says R can't do be used for large data sets or is slow, I ask for more details and almost universally the programmer's complete lack of initiative is the weak link.
- kgwgk 10y agoExcel is used in production very widely, but I'm sure we all agree it has its limitations.
- vegabook 10y agoR just does not have robust software engineering tools for anything that even begins to resemble scale and anybody who says otherwise is denying reality. R can certainly be used in production but the skeleton framework cannot be R. RPC only in my experience with all the structure with something else. R is intrinsically single user / batch with maybe shared database but say goodbye to anything that even starts to approach real time, or multi-node dependent. In my experience the only people who insist that R is robust for production, inevitably have a vested interest. Any objective programmer can see its greatness but also its glaring flaws.
- blahi 10y ago
- apohn 10y agoI used to work in the consulting arm of a software firm and we wrote and deployed R code in production at many Fortune 500 companies. We worked in almost every industry. I spent quite a bit of time refactoring bad R code so it could run reliably in a production environment. There is a ton of bad R code out there that barely works for exploratory analysis, let alone a production environment. So yes, R is used in production environment in a lot of places.
- vijucat 10y agoDid you guys separate out the R process (or multiple processes?) from the rest of the transaction-processing / other server infrastructure or embed the REngine (which sounds like a bad idea to me; incorrect data serialization can easily crash the whole process)? What is a stable way to connect (and reconnect!) to R, assuming it was a separate process? I would think that an indirect communication path, such as Server <--> Database <--> R would work best, but I'd love to hear your battle hardened take on it.
- sandGorgon 10y agoi have the same question - how do you use R in production ?
- apohn 10y agoWe used separate workflows depending on if the data was streaming or batch oriented (e.g. on-demand or triggered by a user). First I'll talk about batch oriented jobs. The company I worked for had a tomcat based product that exposed R via a RESTful API. It was similar to what you get from AzureML now, except it was on-premise. So basically we would call out to this and configure it to restart R sessions if they crashed or timed out. In an ideal situation we would isolate this server from the rest of the processing as much as possible. To be honest our server was pretty basic - it basically served to queue jobs (if needed) and manage RSessions if the server was configured to run multiple sessions. For serious failover we had a second server. We did try to do as much as possible outside of R such as data pipelining an ETL. That was done for the obvious reasons, but also because many customers had SQL and Data people, but not R people. So if one of their Data people understood the data ETL, they could fix it without calling us. For many customers they'd never let R connect to a Database directly. So They'd have a separate process pull data and write it to disk. Then an R script would be triggered and would pick this data up. I never saw major crashing issues with R in production with batch oriented jobs unless there was something unexpected with the size or type of data. Typically as long as there was time between jobs, R's garbage collector would sort things out and be ready for the next job. Also by the time something made it into production we'd hardened the script, frozen the CRAN package versions, etc. So some small issue wouldn't cause a major issue. Streaming data presented it's own adventure. To get data into/out of R as quickly as possible, you need to embed the REngine and talk to it via rJava. If we streamed data through R very quickly it would do fine for a while - then you'd see the memory usage go up and the time for each transaction started to vary greatly. Then it would crash. The solution to this was multiple Rsessions and a lot of telemetry. We would track how long each transaction took through R. As soon as we started seeing a lot of variance in the time we'd restart the engine. By running the multiple Rsessions in round-robin we'd delay the onset of this instability, and it didn't matter when R sessions needed to be restarted. Another trick we used was to cache data in an in-memory database so if something crashed the whole service would restart and pull from the in-memory database instead of trying to fetch old data from the server.