4 ms·
Is there a good reason not to regard this as a standard few-parameter no-gradient optimisation problem, and use something like Nelder-Mead on it?
by improbable22 7y ago
Is there a good reason not to regard this as a standard few-parameter no-gradient optimisation problem, and use something like Nelder-Mead on it?
- bigred100 7y agoI think many people (including the DFO community) already do that. People also consider the notion of multiple objectives important here I believe
- improbable22 7y agoThanks. What's DFO? And what do you usefully do with multiple objectives, besides minimise some total?
- jcagalawan 7y agoDFO is derivative free optimization. With multiple objectives you try to find different solutions given different weightings to the objectives for the Pareto front and pick one depending on the domain.
- Zephyr314 7y agoThis NVIDIA post goes into extending Bayesian Optimization to multiple metrics [0]. It shows how you can use efficient optimization to find a good Pareto Frontier[1]. [0]: https://devblogs.nvidia.com/sigopt-deep-learning-hyperparameter-optimization/ https://devblogs.nvidia.com/sigopt-deep-learning-hyperparame... [1]: https://en.wikipedia.org/wiki/Pareto_efficiency https://en.wikipedia.org/wiki/Pareto_efficiency
- tictacttoe 7y agoBayesian parameter estimation typically trains an emulator to reproduce the objective function using a limited number of design point (order 10 per dimension). Once the emulator is trained, you could of course use a multi dimensional minimization function of your choice to find the best fit point. However, constructing and sampling the Bayesian posterior using MCMC methods has several advantages. Sometimes you can have a local minimum which is essentially flat, so the optimal hyperparameter is unstable. You'll see this in the posterior distribution. Or you could have two parameters which are correlated so it's their sum that's constrained not their individual values. All this information provides important context when understanding your model's uncertainty.
- improbable22 7y agoThanks. Do I understand right that the Bayesian gaussian-process things people do here use only the fully trained loss as input, i.e. just one number L(W), being minimised over W? As opposed to something more detailed about the model, viewed as generating probabilities perhaps, or having training history. Big nearly-flat areas aren't really a new feature of hyperparameter problems... I guess the exact choice of algorithm would depend on how common they are, and maybe Nelder-Mead would be a poor choice. (And I'm not sure how easy it is to parallelise.)