4 ms·
Good question. That was an issue with tanh as activation function, and before residual connections and normalization layers. Tanh as a normalization but with ot
by imjonse 2y ago
Good question. That was an issue with tanh as activation function, and before residual connections and normalization layers. Tanh as a normalization but with other activations and residual present apparently is ok.
- tsurba 2y agoProper initialization is more important. Batch norm and others are important for faster convergence due to forcing the model to focus creating second and higher order nonlinearities, as a simple shift in mean/std is normalized out, and thus the gradient does not point in a direction that would only change those properties of the output distribution.