4 ms·
Thanks. Do I understand right that the Bayesian gaussian-process things people do here use only the fully trained loss as input, i.e. just one number L(W), bein
by improbable22 7y ago
Thanks. Do I understand right that the Bayesian gaussian-process things people do here use only the fully trained loss as input, i.e. just one number L(W), being minimised over W? As opposed to something more detailed about the model, viewed as generating probabilities perhaps, or having training history.
Big nearly-flat areas aren't really a new feature of hyperparameter problems... I guess the exact choice of algorithm would depend on how common they are, and maybe Nelder-Mead would be a poor choice. (And I'm not sure how easy it is to parallelise.)