4 ms·
Coming from pure math, I often feel this way now learning statistics and ML. In pure math, it feels like the threshold for how a novel a concept should be befor
by throwaway080383 8y ago
Coming from pure math, I often feel this way now learning statistics and ML. In pure math, it feels like the threshold for how a novel a concept should be before it gets its own word is much higher.
E.g, we have "regression" and "classification" instead of "supervised continuous prediction" and "supervised discrete prediction".
- p1esk 8y agoIf you don't undestand where the name "softmax" came from, you don't really understand what it is. Softmax is a differentiable approximation of the max function. Plot max(0, x) and softmax(0, x) functions, and it should become clear.
- throwaway080383 8y agoNit: it seems it's more like a smooth approximation to maxarg than max. Yeah it makes sense that this is a super important function, but I still feel like one could just remember the principle that "exponentiation followed by normalization is a smooth approximation to maxarg."
- p1esk 8y agoBasic building blocks of most deep learning models are convolutional layer, pooling layer, fully connected layer, and softmax layer. How do you propose we call "softmax layer" instead?
- TeMPOraL 8y agoNormalization layer? This opens up possibility of using something else than softmax in there.
- p1esk 8y agoWell, there are other building blocks, such as batch normalization layer, or local contrast normalization layer (not to mention a dozen of batchnorm alternatives, e.g. group normalization, weight normalization, layer normalization, instance normalization, etc). If you just say "normalization layer" how am I supposed to know which normalization you're talking about?