3 ms·
You can run a k steps of adam-sgd and differentiate through the learning rate and scalarising parameters of a composite loss function in order to meta-learn the
by nmca 6y ago
You can run a k steps of adam-sgd and differentiate through the learning rate and scalarising parameters of a composite loss function in order to meta-learn them.
Support for general n-th total derivatives is rather good :)