15 ms·
Back in around 2008, SVMs were all the rage in computer vision. We would use hand designed visual features and then a linear SVM on top. That was how object det
by quantombone 8y ago
Back in around 2008, SVMs were all the rage in computer vision. We would use hand designed visual features and then a linear SVM on top. That was how object detectors were built (remember DPM?)
Funny how SVMs are just max-margin loss functions and we just took for granted that you needed domain expertise to craft features like HOG/SIFT by hand.
By 2018, we use ConvNets to learn BOTH the features and the classifier. In fact, it’s hard to separate where the features end and the classifier begins (in a modern CNN).
- wpietri 8y agoCould you say a little more about this? I ask because when we're training human to understand things, there are a variety of benefits to separate feature-understanding from the classifiers. In particular, you get gains in flexibility, extendability, and debuggability. I get why people are happy to take the ConvNet gains and run with them for now. But have you seen any interesting work to get the benefits of separation in the new paradigm? (Or, alternately, is there a reason why those concerns are outmoded?)
- soraki_soladead 8y agoThat's actually closer to how deep learning started. Initially, deep learning mostly consisted of unsupervised (task independent) features with a linear classifier on top. We had to fit an unsupervised model (e.g. autoencoder) layer by layer before using the feature layers in a supervised task. This was because we didn't understand how to train a deep model end-to-end until later. When we learned how to make that end-to-end training work it tended to perform better because the learned features were task specific. You can still learn general features in a bunch of ways, in addition to the older method using autoencoders. For one example, multiple supervised heads with auxiliary losses can learn more generalize features.
- etrain 8y agoIt’s also hard to separate the design of the neural architecture from the definition of the feature extractor.
- yters 8y agoIf you use the right sort of kernel for an SVM it becomes a neural network with automatic architecture derivation. See slide 7: http://www.cs.rpi.edu/~magdon/courses/LFD-Slides/SlidesLect26.pdf http://www.cs.rpi.edu/~magdon/courses/LFD-Slides/SlidesLect2...
- dplavery92 8y agoSignificantly, it becomes a simple, 2-layer neural network. The power of the advances of neural networks in the past decade have largely relied on "deep" architectures with many layers. Very deep networks effectively learn the features from the data, rather than learn a decision surface over a set of hand-crafted features, as in learning with SVMs or shallow neural networks.
- nightski 8y agoI thought it had been proven that a two layer neural network has the same power as a deep one (obviously with a much greater width). It's just that deep neural networks are a lot more practical to train in practice. So I'm not sure how important that distinction is.
- slashcom 8y agoAn infinitely sized 2 layer NN is universal in the same way a Turing machine is universal — sure you can write any program; God help you if you try.
- dplavery92 8y agoThis is something of an academic factoid that has nothing to do with the practice of training and using neural networks, or with the merits of deep networks that I was describing above. Shallow feed-forward networks are "universal function approximators" [0] when the number of hidden neurons is finite but unbounded. Of course, the width of that layer grows exponentially in the depth of the deep network that you might wish to approximate [1]. The statement that "[i]t's just that deep neural networks are a lot more practical to train" (emphasis mine) sounds somewhat reductive; it's not only that depth is a nice trick or hack for training speed, but that depth makes the success of deep networks in the past decade at all possible. We live in a world with bounded computing resources and bounded training data. You cannot subsume all deep networks into shallow networks, and shallow networks into SVMs in the real world. So I am pretty sure of how important that distinction is. And what's more, depth extracts a hierarchy of interpret-able features at multiple scales[2], and a decision surface embedded within that feature space, rather than a brittle decision surface in an extremely high dimensional space with little semantic meaning. One of these approaches generalizes better than the other to unseen data. [0] https://en.wikipedia.org/wiki/Universal_approximation_theorem https://en.wikipedia.org/wiki/Universal_approximation_theore... [1] https://pdfs.semanticscholar.org/f594/f693903e1507c33670b89612410f823012dd.pdf https://pdfs.semanticscholar.org/f594/f693903e1507c33670b896... [2] https://distill.pub/2017/feature-visualization/ https://distill.pub/2017/feature-visualization/
- MAXPOOL 8y agoYou still may want to replace softmax layer with support vector machine for classification sometimes.
- a-dub 8y agoSo the pitch is that you don't have to do feature engineering... but then instead it seems people do network structure engineering with featurish things like convolutions. The performance is still better in most cases but I often have to wonder, are people just doing feature engineering once removed and is the better performance just the result of having WAY more parameters in the model?
- a-dub 8y agoI guess one upshot to the SVM approach is that there's math for quantifying how well a given model will generalize, subject to some assumptions. Is there anything like that in the ANN world?
- computerex 8y agoIn short, no. Not for practically large models used in common tasks like image classification or speech to text.
- nl 8y agoare people just doing feature engineering once removed and is the better performance just the result of having WAY more parameters in the model? Not really, or sort of, depending on how you think. A deep neural network does work - at least to some extent - because of the large number of parameters. However, it is practical because it can be trained in a reasonable amount of time. Things like ResNets are useful because they allow us to train deeper networks. You can create a SVM with the same number of parameters[1], and in theory it could be as accurate (this is basically the no free lunch theorem[2]). But you won't be able to train it to the same accuracy. [1] Of course there are practical concerns about what you do for features, since hand created features just aren't as good as neural network ones. One thing people do now is use the lower layers of a deep neural network as a feature extractor and then put a SVM on top of them as the classifier. This works quite well, and is reasonably fast to train. [2] https://en.wikipedia.org/wiki/No_free_lunch_theorem https://en.wikipedia.org/wiki/No_free_lunch_theorem
- 8y ago
- im3w1l 8y agoTo be fair there is a lot of domain knowledge embedded in the use of a convolutional architecture. There is a fascinating paper where the authors don't even train the weights of the convolutional layers and are still able to achieve good performance. https://arxiv.org/pdf/1606.04801v2.pdf https://arxiv.org/pdf/1606.04801v2.pdf
- cscurmudgeon 8y agoI feel like we pushed features into the architecture and called it a day. Otherwise, why we would we need a gazillion architectures for different problems (or even the same exact problem)?