3 ms·
It's black-box optimization. This means that we just have an objective function, without access to derivatives or whatever other information. This is not releva
by oteytaud 8y ago
It's black-box optimization. This means that we just have an objective function, without access to derivatives or whatever other information. This is not relevant for training weights in deep learning for image classification, or other things for which the gradient works well.
- lostmsu 8y agoThere was a recent paper from Uber, that GA works well for weights, so I wouldn't drop that area right away.
- mi_lk 8y agoWhat’s GA here?
- oteytaud 8y agoGA stands for genetic algorithms.
- oteytaud 8y agoSure GA can be great for weights as well - but mainly when gradient is unreliable. I would not use Nevergrad for training the weights of a convolutional network for image classification for example; whereas I use Nevergrad for WorldModels.
- lostmsu 8y agoDoesn't the model Uber used begin with a bunch of convolutional layer sets, since it processes raw images?