5 ms·
No. In all supervised learning, the algorithm learns a model from the data and generalise it for unseen data. Better generalisation means better model. In unsup
by Kip9000 10y ago
No. In all supervised learning, the algorithm learns a model from the data and generalise it for unseen data. Better generalisation means better model. In unsupervised learning, (mostly clustering, density estimation etc) the same thing happens but we don't tell the algorithm what to learn.
Deep learning is not machine learning plus something else. It is a collection of techniques that overcomes the scalability problem of feed forward neural networks. NNs are very difficult to scale over number of layers. Standard training method of back propagation can't handle many layers because of vanishing gradient and the computational infeasibility brought on by the explosive growth of connections.
NNs are very difficult to scale with additional classification targets you may require (for example, you have a classifier for categorising 10 classes, but to scale it up to 20, requires a lot of topological changes and qualitative analysis.)
Deep learning addresses the scaling over layers with various techniques coupled with hardware acceleration (GPUs). Currently this stand at about 150 layers.
- Houshalter 10y agoEven experts often use the "feature learning" analogy. I don't think it's wrong, or at least a bad way of explaining it. The difference between (deep) neural networks and shallow machine learning, is that NNs can learn arbitrary features. Yes clustering doesn't require feature learning. But it is also super limited in the kinds of features it can learn. Neural nets can learn arbitrary circuits, and other types of functions.
- nightski 10y agoClustering also isn't a supervised learning technique. Even though you might say DNNs can be unsupervised (autoencoders), it generally is not the case in practical systems. So it's not a good comparison at all.
- Houshalter 10y agoI don't care about the supervised/unsupervised distinction. I'm saying they can learn features automatically (with or without supervision.)
- YeGoblynQueenne 10y agoWell, feature learning is unsupervised by necessity, otherwise you're not learning features, you're learning a mapping between features and labels. Deep nets used in the way you say are first trained unsupervised to extract features, then the features are used in supervised learning, to learn a mapping from those new features to labels. You can also do this "by hand" using unsupervised learning techniques like clustering, Principal Component Analysis etc: you make your own features then, and train a classifier afterwards, on the features you extracted in that way. Deep nets just sort of automate the process.
- halflings 10y agoGaussian Processes can learn any arbitrary function. Is it "shallow" machine learning? I think the point that the parent makes is valid: Most of the advantages of deep learning when using a "simple" feed-forward topology is advances related to scaling learning and solving problems encountered at with difficult tasks like image recognition, etc. I do not know enough about neural nets to say if that is all there is to it, but one thing is sure: it's not just about "learning features", although it was shown that the output at every layer abstracts some sort of higher-level features (in the case of image recognition)
- Houshalter 10y agoGPs require exponentially many parameters though. They can't learn arbitrary functions, they just stupidly memorize a lookup table.
- YeGoblynQueenne 10y ago>> it's not just about "learning features" So, I'm in no position to prove this, but my intuition is that any machine learning algorithm can be configured in a semi-supervised learning set-up, like deep nets have. You could train a decision forest classifier for instance to learn in an unsupervised manner. An algorithm I'm developing for my MSc dissertation is essentially unsupervised recursive partitioning, a.k.a. decision trees (only, first-order rather than propositional). Well, possibly not _any_ algorithm. But I get the feeling that many classifiers in particular could be adapted to unsupervised learning with a bit of elbow grease, at which point you could connect them to their own input and, voila, semi-supervised learning. But like I say, I don't reckon I'll be in a position to prove this any time soon.
- YeGoblynQueenne 10y agoThe way I know it goes along the lines of: "multilayer perceptrons with no more than three layers can learn any function to arbitrary precision given a large enough number of inputs". But like halflings say, neural nets are not alone in this. Decision Trees can learn any binary decision diagram I guess (they can encode arbitrary disjunctions of conjunctions). I'm pretty sure there are similar results for other algorithms also. In any case, you can represent a function as a set-theoretical relation and enumerate its parameters- and there you go, learning done with arbitrary precision. That's not what makes neural nets impressive. So what is it? "Shallow machine learning" is a worrying neologism. "Shallow" and "deep" only apply to neural networks, really. You couldn't very well distinguish between shallow and deep K-NN classifiers, say. Or shallow and deep k-means clustering. I mean, what the hell?
- etatoby 10y agoHoly cr*p on a cracker! 150 layers? It boggles the mind. How do you even start propagating over 150 layers? Do you assign specific functions / targets to some of the inner layers?
- Kip9000 10y agoDeep Residual Learning for Image Recognition https://arxiv.org/abs/1512.03385 https://arxiv.org/abs/1512.03385 Abstract: Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers---8x deeper than VGG nets but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers. The depth of representations is of central importance for many visual recognition tasks. Solely due to our extremely deep representations, we obtain a 28% relative improvement on the COCO object detection dataset. Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation. And some good answers here: https://www.quora.com/How-does-deep-residual-learning-work https://www.quora.com/How-does-deep-residual-learning-work
- pedrosorio 10y agoAlso, very deep NN without residuals: https://arxiv.org/abs/1605.07648 https://arxiv.org/abs/1605.07648
- PeterisP 10y agoHighway layers http://arxiv.org/abs/1505.00387 http://arxiv.org/abs/1505.00387 help with this propagation.
- phinance99 10y agoI agree (except for the first word), however I read the question with emphasis on "usual", as in, "What makes DNNs special?" There's pure performance (ex., in a Kaggle competition [http://blog.kaggle.com/2012/11/01/deep-learning-how-i-did-it-merck-1st-place-interview/ http://blog.kaggle.com/2012/11/01/deep-learning-how-i-did-it...] or on a standard data set [http://yann.lecun.com/exdb/mnist/ http://yann.lecun.com/exdb/mnist/], [http://blogs.microsoft.com/next/2015/12/10/microsoft-researchers-win-imagenet-computer-vision-challenge/#sm.001t3igmw1ctxfjwsr62l6ghxb23w http://blogs.microsoft.com/next/2015/12/10/microsoft-researc...] ), but that's what makes any ML method better than another. I think the deeper awesomeness is that DNNs so good at Feature Learning from raw data. On vision, NLP, and speech problems [nice overview by Andrew Ng: https://m.youtube.com/watch?v=W15K9PegQt0 https://m.youtube.com/watch?v=W15K9PegQt0] DNNs have achieved superior performance to the combination of expertly-engineered features + some usual ML algorithm. Where a "usual ML" pipeline might look like (1) engineer features through manual effort by studying raw data and the problem domain, (2) apply ML to those features, a new DNN pipeline might look like (1) Apply DNN to raw data. First off, removing the feature engineering step could be a huge savings in human time spent. Second, there's the potential to get a better answer (!) when you're done. But more than that, the DNN pipeline holds the promise of more regular, systematic improvement. We (as engineers) don't have to wait for a bright idea about how to construct a feature from the data. Instead, we can focus on (1) collecting more and better data, (2) improving the optimization algorithms, and (3 acquiring more computing resources. These latter tasks, I suspect, are easier to define and evaluate than the task "discover a new feature".
- nightski 10y agoYou might not call it feature engineering, but let's face it - most DNN models vary dramatically in structure based on the problem at hand.
- YeGoblynQueenne 10y agoYep. Have a look at DNNs for image recognition, or LSTM RNN. They're the results of some furious architectural work by researchers and not at all simple to come up with (though they may be simple enough to understand now someone's created them).