3 ms·
Nit: it seems it's more like a smooth approximation to maxarg than max. Yeah it makes sense that this is a super important function, but I still feel like one
by throwaway080383 8y ago
Nit: it seems it's more like a smooth approximation to maxarg than max.
Yeah it makes sense that this is a super important function, but I still feel like one could just remember the principle that "exponentiation followed by normalization is a smooth approximation to maxarg."
- p1esk 8y agoBasic building blocks of most deep learning models are convolutional layer, pooling layer, fully connected layer, and softmax layer. How do you propose we call "softmax layer" instead?
- TeMPOraL 8y agoNormalization layer? This opens up possibility of using something else than softmax in there.
- p1esk 8y agoWell, there are other building blocks, such as batch normalization layer, or local contrast normalization layer (not to mention a dozen of batchnorm alternatives, e.g. group normalization, weight normalization, layer normalization, instance normalization, etc). If you just say "normalization layer" how am I supposed to know which normalization you're talking about?