9 ms·
I think it is key to remember that the accessibility of huge amounts of labeled data is behind most of this. ImageNet is 1.2TB (~1.2 million images!). Convoluti
by kastnerkyle 12y ago
I think it is key to remember that the accessibility of huge amounts of labeled data is behind most of this. ImageNet is 1.2TB (~1.2 million images!). Convolutional neural nets have been around for a long, long time (early 90s? or maybe late 80s) and they just needed more data (and a few training tricks :) ).
As we enter the world of "big data" more and more companies are locking this level/size (or bigger) data away in closed vaults. This is a competitive edge in business, but also strangles the ability for academics to do "real world" research, which is often a huge criticism of academia at large.
Academic research can have a role in tech R&D if given a chance, so I hope that companies will open their data to private researchers to continue this kind of improvement - even if under secrecy clauses or some other safeguard.
All that said, the technological improvements have been astounding, and I hope this is only the beginning! There have been some incredible results using related techniques in NLP for machine translation recently, and speech has always been a great playpen.
If anyone is interested in playing with these things in Python, myself and another researcher have recently created a library for using these type of neural networks as a black-box WITHOUT ANY TRAINING!, in a similar style to scikit-learn. See [1] for more. We support only a few nets currently, but plan to integrate support for caffe and pylearn2 in the near future.
[1] http://sklearn-theano.github.io/ http://sklearn-theano.github.io/
- agibsonccc 12y agoGreat work guys! Are these going to only be computer vision nets? I think the huge problem with most deep learning frameworks out there are they are either hard to use or limited in scope. There's nothing wrong with that per se, but I'd be curious to see what you guys intend. It's great that you guys are giving an sklearn interface to it. Edit: I should add my criticisms come from a biased perspective: I write http://deeplearning4j.org/ http://deeplearning4j.org/ I love comparing notes with others regardless. Either way: Wish you the best of luck with it.
- kastnerkyle 12y agoSupport is planned for audio and hopefully text - I am working on building a million song dataset to recreate the work Sander Dieleman did for Spotify, and have had some possible support in October for the weights of a trained speech network! So yes, feature extraction from other domains should be on the horizon. We are specifically trying to make it easy to say: I want to transform an image (or audio, etc.) using a pretrained net. Download the weights, extract the features for me, and give me the feature vectors so I can do something with those. This seems to be really, really, really hard in all the tools I have used and usually involves training yourself, which is not very useful for things on the level of ImageNet. Of course, having nice examples and good docs is one of the great parts of scikit-learn, and is actually one of the things I have been working on most recently. Our docs aren't to that level yet, but I hope they can be one day. DeCAF (precursor to Caffe) and OverFeat binaries were really kind of the first in this regard (about 1 year ago now), but IMO one of the limitations is that they lose interaction with the rest of the Python ecosystem for data munging, and simple algorithms for exploration. By wrapping the weights, we hope to leverage the support of the Python ML ecosystem easily, while still being able to use the power of these networks. Right now the most compelling use case is as part of a scikit-learn pipeline i.e. make_pipeline(OverfeatTransformer, LinearSVC) or whatever. Feed images in, get predictions out. I am also working on a demo of "writing your own twitter bot" similar to https://twitter.com/id_birds https://twitter.com/id_birds , written by Daniel Nouri . I like sloths, so it will definitely be a slothbot. I also hope to support recurrent architectures from Groundhog (https://github.com/lisa-groundhog/GroundHog https://github.com/lisa-groundhog/GroundHog) as several researchers here at the LISA lab have been using it to get pretty amazing results in NLP and audio, both of which are potential targets in the future. If we can leverage their work, it would be a very nice way for people to immediately play with SOTA architectures in different applications. In any case, just loading in weights and extracting image features easily is nice, and was a benefit for both Michael and myself in a research project this summer.
- agibsonccc 12y agoGreat to hear! That's exactly what I'm trying to replicate as well. I'm mainly trying to do it for industry myself. Not a lot of people like making this stuff for the JVM ecosystem (understandble of course...I love python as well) I also agree about caffe as well. The python ecosystem is amazing and should be leveraged which also increases adoption. As I said before, being able to do this at scale for people where their data is stored on the JVM should help it make it more accessibble to a lot of people. Re: Twitter bot. This looks really cool.! I'll be keeping an eye on developments here. Good stuff!
- aswanson 12y agoThis post reminds me why I fell in love with HN. I use Theano, awesome contribution!
- michaelochurch 12y agoConvolutional neural nets have been around for a long, long time (early 90s? or maybe late 80s) and they just needed more data (and a few training tricks :) ). The state of the art has become a lot better in terms of training deep (multi-layer) neural networks. The naive approach (start at a "random" weight setting, use gradient descent) that works on a convex error surface fails catastrophically on deep, complex neural nets. (Shallow neural net training is a non-convex problem as well, but seems to be "less non-convex" in practice.) In the past 10 years, people have become better at getting past that. Theoretically, only one hidden layer is necessary for neural nets to be universal. Thus, for a long time, most research focused on single-layer networks because those were "good enough" to model any mathematical function. The problem is that convergence, for single-layer nets, can be very slow (especially given that you're often doing stochastic gradient descent when working large data sets). Single-layer nets are often very difficult to audit. There's a lot of "cross talk" where unrelated features are mapped to the same "space" in the network. So you have a "black box" that is hard to interpret. Deep neural nets (which is what ML researchers call "deep learning", until MBAs start abusing the latter term, as they have with "big data") have come back into style over the past few years, due to recent research into how to make them actually perform, and findings about superior convergence with the right conditions. In deep nets, there's often a problem of signals either amplifying or vanishing as they propagate throughout the net. The former leads to saturation (the neural net moves very slowly from a suboptimal but flat place on the error surface) and the latter is too linear and unlikely to pick up interesting features. Even now, making deep neural nets not sensitive to initial starting conditions is an unsolved problem, but there's been a lot of progress. I would hazard the guess that the convolutional technique is a lot more useful in deep neural networks than it is in single-hidden-layer neural nets.
- kmavm 12y agoHi, I am a research engineer in Yann LeCun's group at Facebook. I hate to seem to be picking a fight, but you're a prominent poster here, and this comment seems likely to garner a fair amount of attention. Unfortunately almost every sentence you've written betrays a subtle misunderstanding of the space, and the totality is quite misleading. > The naive approach (start at a "random" weight setting, use gradient descent) that works on a convex error surface fails catastrophically on deep, complex neural nets. (Shallow neural net training is a non-convex problem as well, but seems to be "less non-convex" in practice.) Starting at a random (no scare quotes needed) point in weight space and SGD'ing is exactly what Alex Krizhevsky, and all of the derivative convnets over the last two years, did and do. It works just fine; I sit at work doing it all day long. You need to have enough data to train on, big enough models, and enough flops to train the big models on the big data before your interns' grandchildren die. We have all of the above now. Aside: even single-layer neural networks do not have convex error surfaces; convexity, and funky error surface geometry, is not a relevant distinction between shallow and deep nets. There have been no magical optimization breakthroughs, it's still SGD with the same herbs and spices that were used in the 90's (momentum, e.g.). > The problem is that convergence, for single-layer nets, can be very slow (especially given that you're often doing stochastic gradient descent when working large data sets). "Stochastic" gradient descent just means doing lots of weight updates per epoch. Ceteris paribus, training on large, redundant data sets, stochastic converges faster than batch because it gets to consider more points in the weight space than batch per pass over the data. The problem with single-layer neural nets is not that they converge slowly; the problem is that the layer size needs to grow exponentially with the task size. Single-layer neural nets' universal approximation power is thus not of great practical consequence. The power of deep nets is the power of composition: f(g(h(x))) is a strictly more powerful model than f(x) holding the number of parameters constant. > Even now, making deep neural nets not sensitive to initial starting conditions is an unsolved problem, but there's been a lot of progress. You just initialize with a Gaussian ball around zero and explore whatever valley in the error surface you happen to be in. Works 100% dandy. > I would hazard the guess that the convolutional technique is a lot more useful in deep neural networks than it is in single-hidden-layer neural nets. It doesn't really make sense to talk about a "single layer convolutional net", because if you only have a single layer, and all you can do is convolve with it, then the output of your net will necessarily be a big pile of filtered versions of the input image. Unless your task is specifically to learn a target set of filters, it would make no sense to have a single layer convnet.
- Houshalter 12y agoIt can work the other way too. Ironically the academics who created ImageNet restrict who can download it and don't allow it to be used for commercial use.
- kastnerkyle 12y agoWell, they have no choice. Because technically the copyright of each image is still held by the people who took the images (or in some cases the people in the images). Are weights of a trained network based on ImageNet a derived work, a cesspool of millions of copyright claims? What about networks trained to appoximate another ImageNet network? There is a lot of legal gray area here unfortunately, so the signing of the license to work with the images seems like a CYA move to me.
- conjectures 12y ago'Are weights of a trained network based on ImageNet a derived work, a cesspool of million of copyright claims?' I'd argue not. The success of deep learning in vision seems to be in acquiring allocentric representations of objects. Copyright protects the expression and not the concept. Parameter weights describe the concept, not particular instantiations of it.
- kastnerkyle 12y agoThis is what I am hoping too - but derived works are a sticky subject. Imagine a scenario where someone gets of one of these "data vaults" from a large company, then trains a network and throws away the training data. You are still holding the "essence" of their datastore, even without the actual data. I guess we won't really know until someone goes to court over it.
- Houshalter 12y agoYou can't copyright "essence". There are weird legalities about derivative works, but this is an entirely different thing.
- transpy 12y agoCould you provide a link to the results of NLP for machine translation? I'm very interested in MT.
- lars 12y agoI've been hoping for a scikit-learn like wrapper around theano for a long time. Thank you for this, I think it could be extremely useful!