5 ms·
> Deep Neural Networks can learn features from essentially raw data. Usual machine learning starts with features engineered manually. What does this mean? What
by eshvk 10y ago
> Deep Neural Networks can learn features from essentially raw data. Usual machine learning starts with features engineered manually.
What does this mean? What is "raw data" and what is a "feature engineered manually"?
- eegilbert 10y agoOne example is this recent paper on learning (somewhat) high-level attributes of text from character streams alone (i.e., without telling the convolutional networks that things like words and punctuation exist). https://arxiv.org/abs/1502.01710 https://arxiv.org/abs/1502.01710
- YeGoblynQueenne 10y ago>> We show that temporal ConvNets can achieve astonishing performance Yay, astonishing performance! I'm totally gonna waste half an hour of my life to read about what awesome badassery convnets are! Because that sounds so objective! /snark
- hacker42 10y agoTens of thousands of engineers (audio, vision, linguists etc.) spent millions of hours for billions of dollars in the past 30 years to invent algorithms that reliably tell us something about a bunch of data. For example, an corner feature algorithm (such as SIFT) can extract the locations of corners in an image and characterize them. This is essential to many kinds of information processing tasks because we want to apply the same algorithm to different data (generalization), so we kind of need an interface to the data. This interface is called a feature (or feature algorithm, feature extractor or feature-descriptor). All of this work (some of these papers have on the order of ten thousands of citations) is now obsolete because you can start with a random initialization of the weights of a neural network and iteratively improve the weights using backprop for any kind of task. All you need a measure of improvement that is relatively smooth and differentiable with respect to the network weights. What is surprising is that the circuits and programs within reach of backprop training of fully connected neural networks are actually astonishingly good at what they do. But ultimately, this is maybe not so surprising given that our brains do something similar all the time.
- Kip9000 10y ago>All of this work (some of these papers have on the order of ten thousands of citations) is now obsolete because you can start with a random initialization of the weights of a neural network and iteratively improve the weights using backprop for any kind of task Hardly correct. You can't magically learn any kind of task. You can't add arbitrary number of layers and hope for the back prop to do its magic. It is difficult. Deep learning techniques are what makes it somewhat feasible. SIFT is not obsolete because of NNs. They all have their pros and cons. You have to select the right tool for the job. BTW SIFT is not an edge detector (That's the Canny Transform). It describes images using salient features in scale invariant manner.
- hacker42 10y agoTypos fixed. "All" was hyperbole of course, but I think it definitely does not look good for the majority of the work done on features. SIFT was recently outperformed PN-Net for example.
- nightski 10y agoSIFT is also quite old. It's amazing a single technique has retained so much value. Isn't it curious that modern convnets use convolution. On top that, they do convolutions at multiple scales (pooling). Starting to sound very familiar...
- hacker42 10y agoActually the neural net approaches are older than SIFT. Neural nets learn the distribution and even causal factors in the data. To me it seems that this distribution is often just too complex for it to be robustly captured by something that doesn't learn. Learning causal factors critically depends on learning along the depth of the network of latent variables which is a particularly opaque process, but this is what MLPs seem to do quite canonically (convnet being just a restricted special case of MLPs). I mean discerning causal factors is pretty much canonically the act of accumulating evidence with priors (weighted summation), deciding whether it is sufficient evidence and signaling how much it is (non-linearity).
- Homunculiheaded 10y agoThe Neural Network Playground is great for understanding this[0]! The default example is classification of a circle of one class surrounded by a donut of another. There are two features x_1 and x_2 (this is the "raw data"). One solution to this problem is to use a single layer and a single neuron but engineer features manually. These manually engineered features are x_1*x_2, x_1^2,x_2^2, sin(x_1) and sin(x_2). Here's a link to this model (long url)[1]. This model performs very well at learning to classify the data just by combining these manual features with a single neuron. The problem is a human needs to figure out these features. Try removing some and observe the different performance given different manual features. You'll see how important it is to engineer the correct ones. Alternatively you can have 2 layers of 4 neurons [2]. In nearly the exact number of iterations this network also learns to classify the data correctly. This is because the non-linear interactions between neurons are actually transforming the inputs the appropriate ways. That is to say the networks is learning to engineer the features itself. Try removing layers/nodes and you'll find that a simpler network will have a harder and harder time at this. I recommend playing around with the various tradeoff between manually engineered features and network complexity. The interesting thing you will observe is that in some cases the manual features are much faster to learn a simplier model than the network. The big issues comes up when we can't simply "see" the problem in 2d so we have no idea what features may and may not be useful. [0] http://playground.tensorflow.org/ http://playground.tensorflow.org/ [1] http://playground.tensorflow.org/#activation=tanh&batchSize=10&dataset=circle®Dataset=reg-plane&learningRate=0.03®ularizationRate=0&noise=0&networkShape=1&seed=0.66504&showTestData=false&discretize=false&percTrainData=50&x=true&y=true&xTimesY=true&xSquared=true&ySquared=true&cosX=false&sinX=true&cosY=false&sinY=true&collectStats=false&problem=classification&initZero=false http://playground.tensorflow.org/#activation=tanh&batchSize=... [2]. http://playground.tensorflow.org/#activation=tanh&batchSize=10&dataset=circle®Dataset=reg-plane&learningRate=0.03®ularizationRate=0&noise=0&networkShape=4,4&seed=0.66504&showTestData=false&discretize=false&percTrainData=50&x=true&y=true&xTimesY=false&xSquared=false&ySquared=false&cosX=false&sinX=false&cosY=false&sinY=false&collectStats=false&problem=classification&initZero=false http://playground.tensorflow.org/#activation=tanh&batchSize=...