3 ms·
MLR and caret do a lot in the 'unified API for models' approach, though (IMO) they're not as streamlined / mature as sklearn, especially if trying to do modelin
by perturbation 8y ago
MLR and caret do a lot in the 'unified API for models' approach, though (IMO) they're not as streamlined / mature as sklearn, especially if trying to do modeling using sparse matrices (they mostly expect dense matrices as input).
- claytonjy 8y agoI haven't played with MLR; how do you like it, esp. compared to caret? I never really caught on to caret as it felt rigid and clunky and non-idiomatic, but I've been using Max's newer rsample (setup CV) + recipes (like sklearn pipeline) + yardstick (metrics) packages to good effect lately. parsnip, which will handle the core model-fitting, seems promising but is too early to use yet. I don't expect sparse matrix support to get any better, as the core model functions would have to be rewritten entirely to avoid rehydrating them, which AFAIK nobody is seriously working on :(
- perturbation 8y agoI've only used it a little in side projects (xgboost + mlr); I think it works better than caret, though (not as brittle). Most of my day-to-day is text data (which mlr isn't really well suited for). What originally drew me to it was a blog post [1] about using their model-based optimization framework for tuning hyperparameters. It's a lot more sophisticated than anything I've seen for Python, including hyperopt / hyperband. [1]: http://mlr-org.github.io/How-to-win-a-drone-in-20-lines-of-R-code/ http://mlr-org.github.io/How-to-win-a-drone-in-20-lines-of-R...
- claytonjy 8y ago> Most of my day-to-day is text data Ah the concern about sparse matrices makes even more sense now! I love that tidytext can produce sparse tf-idf matrices...but hardly any models can use them :(