4 ms·
" The facts speak loudly for themselves; in most cases, deep learning-based solutions lack mathematical elegance and offer very little interpretability of the f
by linux_devil 9y ago
" The facts speak loudly for themselves; in most cases, deep learning-based solutions lack mathematical elegance and offer very little interpretability of the found solution or understanding of the underlying phenomena."
I don't agree with this statement, a simple look at cs231n lecture series will show you how much math is involved. A lot of articles/people claim its a black box, but while writing a small network architecture you realise it's not. Methods like stride, padding, activations, learning rate, optimisations, drop outs etc. give you "aha" moment which is followed by a mathematical explanation. One should study the topic thoroughly before criticising it.
- matt4077 9y ago"Interpretability", in this case, is concerned with the problem domain, not ML itself. A hypothetical example: A trained decision tree algorithm for a medical decision may place the gender, or the age, of the patient at the root. That lends itself to a quick interpretation as to the relevance of factors for treatment, whereas with a Neural Net, you'll get millions of arbitrary floats that do not impart any meaning just by looking at them. That's not to say NNs don't sometimes give tantalising insights, as the article points out. I've seen a view visualisations of generative character models that were fascinating–such as finding individual neutrons tracking sentence length or nesting depth. Same for some interesting patterns emerging in the intermediate layers of object recognition networks: oh, I never knew ears were so important for face recognition.
- reader5000 9y agoThere is no understanding of why NN training converges in reasonable time and does not over/underfit. (Unlike for say convex SVMs)
- gjulianm 9y agoHaving a lot of math involved does not mean that it is mathematically elegant. As a mathematician, I ask myself several questions. First of all, what is really a neural network? Is it an approximating function? Is it a geometric separation on a space, such as SVM? Is it a manifold classificator? (see http://colah.github.io/posts/2014-03-NN-Manifolds-Topology/ http://colah.github.io/posts/2014-03-NN-Manifolds-Topology/, which is very interesting) Also, what are we approximating? Continuous functions? Non-continuous functions? Are they even functions and not probability measures? Are those functions arbitrary or do they represent something like a manifold? And the most important: how well are we approximating whatever we want to approximate? The universal approximation theorem gives uniform convergence for measurable functions, but do not specify at which rate or depending on which parameters. It is a strong theorem but not that surprising from the mathematical standpoint, where you already know that you can approximate any function by continuous, compactly supported functions. Finally, how do you mathematically define the problems that arise in neural networks? What is overfitting? How does the learning algorithm affect the results? The fact that some techniques are justified by mathematical explanations does not mean that it is mathematically elegant. For it to be mathematically elegant you should have at least clear definitions of the objects of study and the problems you want to solve. I don't think this is the case in neural networks.
- linux_devil 9y ago"Having a lot of math involved does not mean that it is mathematically elegant" I agree to your point, but at the same time how much effort is being made to understand the intricacies of deep learning is my question. In my opinion, those who dont understand these techniques bluntly say its a black box but its not entirely a black box. I am confident if more reearchers start to peel it off layer by layer , a lot more insights will be generated given the field is relatively new .
- deleted 9y ago[deleted]
- linux_devil 9y agoNot sure why I am downvoted to put my view upfront