61 ms·
Regarding R for BigData, do you think R is going to continue to be a stable of the analytical toolset as data sizes continues to grow? I use R to prototype mac
by alook 14y ago
Regarding R for BigData, do you think R is going to continue to be a stable of the analytical toolset as data sizes continues to grow?
I use R to prototype machine learning techniques on a small set of data, or visualize some summary statistics. But if I want to run K-Means Clustering or Support Vector Machine algorithms on 1,000,000,000 rows of data, I've found that running R on Hadoop is tricky. There are some libraries out there ( for example, RHadoop https://github.com/RevolutionAnalytics/RHadoop/wiki/rmr https://github.com/RevolutionAnalytics/RHadoop/wiki/rmr ) but they require writing your algorithm in such a manner that algorithms must be adapted to run within map() and reduce() functions. My understanding is that the built-in functions that make R so useful will often not adapt well to a mapreduce algorithm.
From what I've seen, once an algorithm is prototyped in something like R/Matlab, if the data size warrants it, it's best to re-write the algorithm in Java MapReduce or use Apache Mahout.
- aheilbut 14y agoor C.