4 ms·
I am a data scientist (and use both R and Python regularly). I would take the plot in your link with a grain of salt - RapidMiner is not frequently used (at al
by perturbation 8y ago
I am a data scientist (and use both R and Python regularly). I would take the plot in your link with a grain of salt - RapidMiner is not frequently used (at all) with any of my friends / coworkers, though maybe I'm just in a bubble.
I like R a lot the more I use it. Tidyverse libraries (especially purrr + dplyr + ggplot2) make R a joy to use. I would argue that R is ahead of Python in terms of libraries for everything but deep learning (which is, after all, only part of ML).
- claytonjy 8y agocouldn't agree more; I think working in R is a vastly different experience now than it was 5+ years ago, and has changed (for the better) much more quickly than Python has. I'm primarily thinking of the tidyverse here; the three you mention are so much more intuitive than loops + pandas + matplotlib. With Max Kuhn's new tidymodels stuff, I think R has a real shot at providing a nice alternative to sci-kit learn, though there's a lot of catching-up to do. For the deep learning stuff, Keras in R is about as nice as Keras in python, but I'm not holding my breath for pytorch-like workflows in R.
- perturbation 8y agoMLR and caret do a lot in the 'unified API for models' approach, though (IMO) they're not as streamlined / mature as sklearn, especially if trying to do modeling using sparse matrices (they mostly expect dense matrices as input).
- claytonjy 8y agoI haven't played with MLR; how do you like it, esp. compared to caret? I never really caught on to caret as it felt rigid and clunky and non-idiomatic, but I've been using Max's newer rsample (setup CV) + recipes (like sklearn pipeline) + yardstick (metrics) packages to good effect lately. parsnip, which will handle the core model-fitting, seems promising but is too early to use yet. I don't expect sparse matrix support to get any better, as the core model functions would have to be rewritten entirely to avoid rehydrating them, which AFAIK nobody is seriously working on :(
- perturbation 8y agoI've only used it a little in side projects (xgboost + mlr); I think it works better than caret, though (not as brittle). Most of my day-to-day is text data (which mlr isn't really well suited for). What originally drew me to it was a blog post [1] about using their model-based optimization framework for tuning hyperparameters. It's a lot more sophisticated than anything I've seen for Python, including hyperopt / hyperband. [1]: http://mlr-org.github.io/How-to-win-a-drone-in-20-lines-of-R-code/ http://mlr-org.github.io/How-to-win-a-drone-in-20-lines-of-R...
- claytonjy 8y ago> Most of my day-to-day is text data Ah the concern about sparse matrices makes even more sense now! I love that tidytext can produce sparse tf-idf matrices...but hardly any models can use them :(
- stared 8y agoI am a data scientist as well (ML/DL). This thing with RapidMiner surprised me a bit (I don't know what it is THB), but I guess it may be a totally different strand of analytics or stuff. > would argue that R is ahead of Python in terms of libraries for everything but deep learning Certainly there are more implementations of various statistics. Still I jump to R because there is something which isn't implemented in Python (some decision tree models, and things related to Item Response Theory - in my case). And just for ggplot2. :) Yet, with "being ahead" it is tricky. For example, scikit-learn have a limited collection of algorithms, but once it covers something, it is with the same API and typically well-implemented.
- shoguning 8y agoI had not heard of RapidMiner either. I looked up its website on Alexa: https://www.alexa.com/siteinfo/rapidminer.com https://www.alexa.com/siteinfo/rapidminer.com It looks like most users are in China.