9 ms·
AutoML toolkit for neural architecture search and hyper-parameter tuning
- hestefisk 8y agoThis is very cool.
- wongarsu 8y ago> We support Linux (Ubuntu 16.04 or higher), MacOS (10.14.1) in our current stage. No Windows support in a Microsoft product. Curious. This looks very useful for tuning hyper-parameters, and the fact that the tuned algorithm is treated as a black box makes this very flexible.
- yeahhhhh 8y agoActually, they will support in Windows later. Due to many developers usually train their deep learning model in Linux, so they support Linux and Max first.
- sandGorgon 8y agointeresting - there's no scikit support, which for long has been the mainstay for data scientists everywhere. Are people migrating from scikit to tensorflow in production for non-deep learning usecases ?
- mmq 8y agoI think it should probably support scikit as well as any other library, since it's only making suggestions of hyper-parameters based on recorded/historical observations or random evaluations. At least that's the behaviour of the platform[1] I am working on. [1]: https://github.com/polyaxon/polyaxon#hyperparameters-tuning https://github.com/polyaxon/polyaxon#hyperparameters-tuning
- mmq 8y agoUPDATE: Looking at the docs, there's an example[1] using this library with scikit-learn. [1]: https://nni.readthedocs.io/en/latest/sklearn_examples.html https://nni.readthedocs.io/en/latest/sklearn_examples.html
- pplonski86 8y agoI think it all depends on the purpose of the library and who is a target user. The NNI is a package for tuning neural networks models, it will be mostly used in use cases that require deep neural networks, like image classification or voice recognition. BTW, I think all autoML solutions forget about end users. They all require too much engineering knowledge from the user. I think it will be nice to have an autoML solution that can be used by citizen data scientist.
- minimaxir 8y ago> BTW, I think all autoML solutions forget about end users. They all require too much engineering knowledge from the user. I think it will be nice to have an autoML solution that can be used by citizen data scientist. This is the approach of a project I am currently working on. (and am now explicitly making clear in the README!)
- pplonski86 8y agoCould you provide some link to the project?
- human_scientist 8y agoWhat about approaches like auto-sklearn [1]? With these it is basicaly: >>> automl = autosklearn.classification.AutoSklearnClassifier() >>> automl.fit(X_train, y_train) >>> y_hat = automl.predict(X_test) [1] https://automl.github.io/auto-sklearn/stable/ https://automl.github.io/auto-sklearn/stable/
- ayidnelm 8y agoThere's also auto scikit-learn https://github.com/automl/auto-sklearn https://github.com/automl/auto-sklearn if you haven't already come across that.
- cuchoi 8y agoI think that for Neural Networks scikit has not been the "go to" library, in particular AutoML advertises that they automate neural architecture search which I don't think scikit allows a lot of flexibility for that.
- samcodes 8y agoHave you seen TPOT [0]? AutoML library that uses genetic algorithms to write scikit code for you. So fun. [0] https://github.com/EpistasisLab/tpot https://github.com/EpistasisLab/tpot
- williamsmj 8y agoThere is scikit support. There's an example in the docs. https://nni.readthedocs.io/en/latest/sklearn_examples.html https://nni.readthedocs.io/en/latest/sklearn_examples.html
- streetcat1 8y agoscikit learn is a different type of search, hence it will not be supported by this tool or any DNN search tool. DNN require an architecture search, I.e. the building block are full layers, depth of the network, optimizer etc. scikit learn search a parameter space, I.e. the algorithm weight are much much simpler and few. So to sum up, DNN search involve big building blocks while scikit learn search (or for that reason any "classical ML" algorithm) is more of a parameter search. [ The actual sci kit learn search would also include pre processing steps, which can be seen as a separate block] Also, note that that DNN search is much more expensive than scikit learn search (100X) ]
- williamsmj 8y agoThis tool absolutely supports scikit-learn. Please see the docs. https://nni.readthedocs.io/en/latest/sklearn_examples.html https://nni.readthedocs.io/en/latest/sklearn_examples.html.
- human_scientist 8y agoAutomatically building a scikit learn estimator might include many conditional hyperparameters and also a very large amount of them (<100) [1]. However, performing joint architecture and hyperparameter search can be framed to be on a much simpler search space, e.g., for a recent paper that aims to automate the design of RNA molecules, we formulated a 14 dimensional search space which includes very little conditional hyperparameters [2]. The tools included in the repository are very broadly applicable and only a few of them are specifically targeted at neural architecture search. [1] https://www.kdnuggets.com/2016/08/winning-automl-challenge-auto-sklearn.html https://www.kdnuggets.com/2016/08/winning-automl-challenge-a... [2] https://openreview.net/forum?id=ByfyHh05tQ https://openreview.net/forum?id=ByfyHh05tQ
- scottlegrand2 8y agoAt a previous gig we tried to do this: port a computational graph that wasn't a neural network to tensorflow. It was a disaster. Tensorflow is very tightly optimized for the things Google think are important. if you fall off of those paths tensorflow is a god-awful slow tool to use. We saw a ~20x regression in performance. in contrast, when we wrote bespoke GPU code for the graph, we saw a ~25x performance increase over relying on CPU plus MKL. I am being deliberately vague here and I cannot give further detail.
- ec109685 8y agoYou are somewhat uniquely qualified to do so: > possibly the world's first or second (full-time) CUDA programmer, with 14 filed patents, and the world's fastest implementations of molecular Dynamics (CUDA ports of Folding@Home and AMBER).
- scottlegrand2 8y agoYes, compared to someone who insists on doing all of their computation from python alone, I have a unique (and in my opinion absurd) advantage. Because I think that's insane. It's one thing if you don't care about speed and you care more about time-to-market. It's another thing if you're complaining about things being too slow but you're not willing to learn about anything that would let you do anything about it. I run into far more of the latter.
- nurettin 8y agoDo we need a hyper-parameter tuner tuner for this?
- mlthoughts2018 8y agoStuart Geman (one of the inventors of Gibbs Sampling) always used to say, “Parameters are the death of an algorithm.”
- nurettin 8y agoEnvironmental constraints (like width, height) are not bad. I would have argued Mr. Stuart.
- mark_l_watson 8y agoI manage a machine learning team for a large financial services company and AutoML tools, Microsoft’s NNI included, are on our radar. I think the `future of work` for machine learning practitioners will quickly separate into two groups: a very small and elite group that performs research and a much larger groups that use AutoML but whose jobs also deal more with data preparation (which gets automated also) and ML devops, supporting models in production.
- mlthoughts2018 8y agoThis sounds like parody to me. There are so many problems in applied statistics, and neural networks are not helpful for most of them. Consider Bayesian analysis for very small data sets as an example (just the tip of the iceberg). In financial services in particular, there are tons of time series and regression problems on small data such that a neural network (beyond perhaps some super small MLP) would be a ridiculous thing to try. I think the breakdown of workload you described will only happen in business departments where there is a need for large scale embedding models, enhanced multi-modal search indices, computer vision and natural language applications, and maybe a handful of things that eventually productize reinforcement learning. I could also see this happening in businesses that can benefit from synthetically generated content, like stock photography, essays / news summaries / some fiction, website generators, probably more. What I described above is a tiny drop in the ocean of applied statistics problems that business have to solve.
- mjburgess 8y agoThe problem is "Applied Statistics" became "Machine Learning" which became "AI" which became "Deep Learning". Throw away all the BS. and, yes, it's obvious.
- DebtDeflation 8y agoIt's another example of the FAANG + Bay Area Startups world versus the other 99% of Corporate America. In the latter world, most of the "machine learning" in production is traditional stuff like Random Forest, SVM, and more recently Gradient Boosting. Hell, Marketing departments across the country are still running old school decision tree (CART and CHAID) models and logistic regression models written in SAS 20+ years ago. DL/NN is a minuscule proportion of production ML in the enterprise space.
- sgt101 8y agoI don't understand - isn't this model fishing? How is it different?
- glial 8y agoYes, but that's not necessarily bad. You want a model that effectively captures the structure present in your dataset. There are currently only rules-of-thumb in model architecture, and it makes sense to explore the model space to determine which architecture and hyper parameters are suitable to the needs at hand. Two things save this from being a statistical sin: one, the final evaluation set is typically different than the validation set, and evaluation is only performed at the end of the 'fishing expedition', thus providing a reliable measure of the model's ability to generalize. Second, we're doing engineering here, not science, and our goal is to capture the structure of observations and not make a scientific claim about values of latent parameters.
- thanatropism 8y agoWith training, test and validation sets. In good old fashioned statistics there's the idea of the jackknife: for the i-th sample run a regression on all the data except i, and store statistics of interest (coefficients, predictions, etc). This gives you an ipso facto sampling distribution for the statistics of interest. Similar and more common in econometrics is the bootstrap: run your model in like 1999 subsamples (with repetition) of the data and get sampling distributions. With said sampling distributions, whether from the jackknife or the bootstrap, you're able to test whether your model is valid -- what's the probability that it'll have significant coefficients or an r2/mae/mape score indicating predictive capacity. Cross-validation (and even scikit-learn is starting to default to five folds not three) is a "lazy" version of this. You don't get a sampling distribution but at least you're able to know that a given model appears good because it grips the data with all its might and doesn't work out-of-sample. sklearn even offers the jackknife under some ML-y name like "one at a time scoring".
- williamsmj 8y agoI'd be interested in the creator's thoughts on this paper, "Random Search and Reproducibility for Neural Architecture Search", https://arxiv.org/abs/1902.07638 https://arxiv.org/abs/1902.07638, posted on the arxiv last week. Among other conclusions, they find: "Our results show that random search with early-stopping is a competitive NAS baseline, e.g., it performs at least as well as ENAS, a leading NAS method, on both benchmarks" ENAS, the specific algorithm that they find does no better than chance, is in this library. My understanding is that the results are pretty generic though, i.e. NAS is very far from a solved problem. (Hyperparameter tuning for "classical" models are another matter. That's commoditized and available as a service at this point, see tpot, DataRobot, etc., etc.)
- perturbation 8y agoTheir example with LightGBM (https://nni.readthedocs.io/en/latest/gbdt_example.html https://nni.readthedocs.io/en/latest/gbdt_example.html) is very cool - I wanted to put together a custom script with mlflow + catboost + mlrMBO to do something similar, but this puts everything together in one package. I think this does everything MLFlow does and more (besides maybe helping with deployment?)
- angel_j 8y agoDoes it test against and prevent over-fitting?
- yzh 8y agoI'm working on auto hyper-parameter tuning and network optimization, I always think that people have put too much focus on NAS, which aims to create a whole new network from scratch, but not nearly enough on hyper-parameter tuning and local structural optimizations for an existing network, which I think is more demanding at least in the industry. Looks less cool than NAS though, maybe that's the reason.