6 ms·
To be honest I've been using R a bit lately for my work and while I like it I don't find it at all innovative. That's not a criticism of R: the libraries it has
by nanairo 16y ago
To be honest I've been using R a bit lately for my work and while I like it I don't find it at all innovative. That's not a criticism of R: the libraries it has are amazing, as well as the mindshare among people who care of statistics.
But I wonder why R actually needs to exist as its own language. It seems it could be recast in Ruby for example or one of the latest functional languages.
So I am kind of pleased my this news... if there's gonna be a need for R to have its own language, speed seems to be the most important distinguishing feature. A bit like Fortran is still used in science.
(incidentally... don't let people tell you otherwise: Fortran(90+) is a very nice language... much more pleasurable than C to use and gives you better performance (unless you know a lot about compilers and compiling flags... but most scientist don't ^_^))
- ewjordan 16y agoBut I wonder why R actually needs to exist as its own language. It seems it could be recast in Ruby for example or one of the latest functional languages. Indeed - it would be a shame for them to start over from scratch and end up coming up with a brand new language, brand new syntax, brand new quirks, brand new performance problems, etc., while they could have simply searched around a bit for something that's already mature and somewhat optimized as well as suiting their needs. If they wanted to add on to or modify an existing language (for instance, to provide more concise syntax for some of the things that are more important in statistics than in general purpose programming), that would be just fine, it could become a dialect of some other language, but starting fresh seems like an awful waste of energy... Something that ran on the JVM would be awesome, they'd have no trouble at all rebuilding the massive library of contributions.
- equark 16y agoFear not, some of us are working on this very hard.
- knowtheory 16y agoo? what is it that you're doing?
- equark 16y agoBuilding a next generation of statistical computing.
- phren0logy 16y agoWould you be willing to be more specific? That's not very helpful.
- equark 16y agoI can't be yet, unfortunately. But there are definitely several different groups of very talented people working on this problem (think MIT and Harvard PhDs). I suspect some very interesting work will come out in the next year or two. Whether they are open source remains another question.
- filobloomz 16y agoAustin Heap? Is that you?
- knowtheory 16y agoI was hoping you'd be at least the slightest bit more specific :) are you working in a proprietary startup? are you working in academia? presumably you're working on a codebase with the intent to distribute it to users at some point?
- deleted 16y ago[deleted]
- bigfudge 16y agoPerhaps I'm wrong, but scipy and numpy seem to do a lot of what they want, and if they started there a lot more of their effort could go into porting the libraries.
- earl 16y agoYes, but (to my shock), their speed makes even R look fast. Which I didn't think as possible. My problem was taking a matrix market formatted matrix, loading it, turning it into a sparse column vector representation, then computing norms of the columns. It was running for roughly 8 hours on 12MM columns. R of all things was running faster. I reimplemented in java and it takes < 1 second.
- studer 16y agoGiven that numpy/scipy is basically a collection of C/C++/Fortran primitives, chances are that you managed to write a program that spent very little time actually computing things, and a lot of time doing something else. Not sure what that could be, though; even low-level algorithms usually run 10-100x slower than C speed if naively coded in plain Python, so a 30000x slowdown using a specialized library sounds rather odd. I assume you checked for memory leaks and swapping. Did you do any profiling?
- cdavid 16y agoTo be fair, the sparse matrix package is still very rough. I am highly skeptical of the 1sec in java vs 8 hours in scipy, though. In general, numpy/scipy is quite faster than R: it does not have the pass by value semantics for once. I am also skeptical about writing a "new" R: the main value of R is in the R packages. Any new language would threw that away.
- earl 16y agocdavid -- see http://news.ycombinator.com/item?id=1688694 http://news.ycombinator.com/item?id=1688694
- Quiark 16y agoFor JVM, there is Incanter which is a statistics library written in Clojure. It is backed by Parallel Colt for the heavy number lifting. Note that I'm not trying to say that Clojure would be a scientist-friendly language :)
- nanairo 16y agoHow does Incanter work on HPC? R is pretty awful from that point of view, and if there ever will be room for a specialised statistical language it's got to be able to do massive number crunching. I've heard bad things of JVM for tightly coupled jobs on HPC (though I know there's been some improvement: e.g. a lot of work done by EPCC in Edinburgh). Does Clojure manages to offer a good parallel implementation on top of the JVM or has no work been done in this area?
- Quiark 16y agoAccording to the website, Parallel Colt supports multicore machines. There's no mention of MPI though. As far as I know, Incanter only wraps it, so it does not influence the performance that much..
- piggybox 16y agoA lot of people blame Lisp for too many brackets, but I generally feel they are not the issue after reading Lisp for a while. What may make it scientist-unfriendly is to write math formulas in infix grammar.
- regularfry 16y agoIncanter has the $= macro for writing infix expressions, so I don't think that need be an issue.
- _delirium 16y agoBut I wonder why R actually needs to exist as its own language. It seems it could be recast in Ruby for example or one of the latest functional languages. I agree if we're talking about defining a new R (like the blog post discusses), but the existing R makes sense to me to exist as its own language. It wasn't really invented from scratch gratuitously, but began as an open-source reimplementation of the Bell Labs "S" language, which had already become close to a de-facto standard in the statistics community. After 30 years of writing S and R code, I think there's going to be a big uphill effort if you want to convince statisticians to read and write Ruby code instead. You'd also lose the ability to run snippets of code from the thousands of existing papers that include R/S code in an appendix. One in-between possibility could be to retain the standard syntax/semantics but target an existing VM with a bigger development community. Ihaka seems to think that's impossible (he briefly discusses attempts to compile R as futile), but lots of weird and highly dynamic languages now have more efficient implementations than most people would've thought possible 10 years ago.
- fhars 16y agoR and S already have completely different scoping rules (S is dynamically scoped, while R has static scope, although quite broken according to the article). So it might be totally possible to develop a third language with mostly the same syntax (so snippets from papers that don't explore the darker corners of the existing semantics can be reused or easily adapted) and this time consistent scoping rules. If it would be worthwhile, I don't know.
- messel 16y ago"One in-between possibility could be to retain the standard syntax/semantics but target an existing VM with a bigger development community" I think that's a great move for a number of languages as they lose popularity over time. The Scheme on lisp machines was outpaced by version on compiled machines, and is the version we use today.
- equark 16y agoThe problem is a huge number of R packages use compiled code that won't be portable to this new VM.
- earl 16y agoI think the interviewee has, elsewhere, been pretty explicit about his desire for the R committee to get away from maintaining their own language and to build on a platform maintained elsewhere. Basically, it seems like he thinks R should be a package and some libraries, but not in the runtime/compiler business.
- EdiX 16y agoIndeed, R "the language" has been the main barrier to my adoption of R.
- earl 16y agoR needs it's own language because one of the most important factors in it's widespread use is the language itself. First, it is very similar to an older language that is one of the first widely used statistical packages. Second, the syntax is very easy to use at a simple level. For example, to read a csv file and compute a linear regression with betas, p-values, the works, all you have to do is: # read a csv file with headers into ram data <- read.csv(file='blah.csv', header=T) # compute a linear model with dependent variable income, explanatory variables gender, education, and ethnicity, with automatic creation of (n-1) indicator variables as appropriate for categorical data model <- lm(income ~ gender + education + ethnicity + age) # if instead, I want to use a glm family of models, this is also trivial.. # note this is nonsensical statistically with respect to my data, but I just thought of this off my head, and it demonstrates how easy it is to do sophisticated things in R model.glm <- glm(income ~ gender + education + ethnicity + age, family=binomial, link=cloglog) That's it. So you can see why stats people love it -- it's easy to pick up and use to get productive work done. It's easy to show students. For what it's meant to do originally -- desktop statistics -- the language works quite well. It lacks as a programming language in many regards, don't get me wrong, but simplicity and ease of use, particularly in the beginning, are critical. Also, the data frame is the single best data structure I've ever used for manipulating tabular data. Finally, the excellent repl makes working in R and exploratory analysis absolutely awesome.
- nanairo 16y agoCould you tell me a bit more about why you like data.frame? Is it not just a 2D matrix?
- Quiark 16y agoI'm pretty sure you could do a lot of this in Python as well.
- hadley 16y agoIt seems that way on the surface, but their are a lot of subtleties to R (lazy evaluation of function arguments, access to both static and dynamic scope) that would make implementation in another language difficult. R has many features that are not that desirable from a programming language stand point (and that do make optimisation difficult), but do make it a very expressive, flexible language for statistical computing.
- hadley 16y agoI'd be interested to learn how to add missing value support into an existing language. It's pervasive in R, so important for statistics, and it seems like it would be hard to patch onto an existing language.
- equark 16y agoHow important is the distinction between NaN versus NA? I agree this is a subtle issue.
- cdavid 16y agoIt is important - you cannot know whether NaN is coming from a computation or is really a missing value otherwise. In Numpy, we have the MaskedArray implementation to do this.
- equark 16y agoWhat does a NA value become when you extract it to a float? i.e. What is the behavior of X[0]? It is somewhat confusing that python base types and numpy differ in behavior, for instance when dealing with inf or divide by zero exceptions. I think this gets to hadley's point that it will be hard to bolt on R to an existing language.
- cdavid 16y agoIf X[0] is masked, it will return the value mask, of type MaskedConstant. As for python vs numpy differences: yes, those can be confusing, and that's inherent to the fact that we use a "real" language with a library on top of it instead of the language designed around the domain. If you want to do numerical computation, you do want the behavior of numpy in most cases, I think. There is the issue of "easiness" vs what scientists need. You regularly have people who complain about various float issues, and people with little numerical computation knowledge advising to use decimal, etc... unaware of the issues. Also, python will want to "hide" the complexities, whereas numpy is much less forgiving. As for the special case of divide by 0 or inf, note that you can get a behavior quite similar to python float. You can control how FPU exceptions are raised with numpy.seterr: import numpy as np a = np.random.randn(4) a / 0 # puts a warnings, gives an array of +/- inf np.seterr(divide="raise") a / 0 # raise divide by 0 exception
- lambda 16y agoIf you want to do scientific computing in a modern general purpose language, NumPy/SciPy are pretty popular. Haven't used them much myself, but they look pretty decent.