3 ms·
I think it's worth clarifying a few points here - this is a fairly naive view of deep learning being put forward. In particular - 1) what does 'automatic tren
by throwawaydl 12y ago
I think it's worth clarifying a few points here - this is a fairly naive view of deep learning being put forward.
In particular -
1) what does 'automatic trend discovery' mean? We've been able to do change-point detection, linear regression, etc. for hundreds of years. If you're talking about automatically learning a feature representation, then there are other algorithms that can do this, in a much simpler way. If you're arguing that it produces better representations, then make that argument.
2) This is almost _completely_ false, and indicates a substantial lack of experience of ML beyond deep learning. Other machine learning algorithms (SVMs, LR, even some decision tree algorithms) are _much_ easier to train in parallel - this is (partially) because your objective function has certain nice properties that allow you to combine partial solutions together that are produced in parallel (convexity, separability). When you're using gradient-based methods on an incredibly ugly non-convex function from a multi-layer neural network, you're in a completely different world.
Granted, there have been techniques coming out for training in multiple address spaces, but these are _hacks_ to get around the ugly structure of the problem, not the principled approaches that exist for other algorithms.
I don't have any perspective on your deeplearning4j library, but I'm skeptical of its utility given the existence of existing well-tested deep learning libraries written/contributed to by renowned experts in the field (e.g. cuda-convnet, caffe, torch). This stuff is a) very tricky to get right, b) very tricky to debug, and c) very performance sensitive. Just a quick pass through shows zero references to CUDA/GPGPU programming, so I'd suspect performance is going to be significantly worse than the aforementioned libraries.
- agibsonccc 12y ago1. In this case, we are talking about the pretrain part of neural nets. Automatic trend discovery comes down to doing feature extraction for the user. That being said, if I was inexperienced I wouldn't be teaching this stuff[1]. Am I the best machine learning practitioner out there? No. A lot of us aren't. I am all about making other people's jobs practical though. Yes, I am talking about learning better representations. See hinton's deep autoencoder work as a prime example of this comparing PCA to RBM based methods for topic detection[2]. 2. Google and people way smarter than I am seem to be doing just fine with this[3]. That being said, I didn't say that random forest (with whole companies built on this parallelism[3]) or any of the algorithms WEREN'T friendly. I would say one of the main appeals for deep learning is the scale of data with which it can benefit from. Feel free to be skeptical all you want, if the researchers want to take the time to write a full stack distributed framework, I welcome others in to the game. The problem with the packages out there right now, (being matlab, python) are training times, and integrating in to an actual ecosystem. I'm addressing this this year at 2 different talks[5][6]. Replying to your last point, I use blas underneath for all of the matrix calculations, I will be adding GPUs later this year, and yes you're right,this stuff is hard to make. I also wouldn't be publicizing it if I wasn't already using it in production applications. Frankly right now though, I use cpu matrices right now, because I can fire this up on AWS (without the limit of GPU RAM), and it's practical for hadoop deployments. Honestly whether we like it or not, GPUs take a lot to get right. NVIDIA[7] and AMD[8] are going to make my job pretty easy though. To end, if I was afraid of every little obstacle, why do anything in the first place? While you're hiding behind a throw away account, I'm actually trying to put this in the hands of people who don't have the time to learn every little thing about neural networks. At the end of the day, I follow the papers very closely and enjoy what I do. I also work on all sorts of different techniques for different problems combining different machine learning algorithms for different tasks (just like anyone else would). This framework is my way of getting this out to everyone else. If you have a deep learning framework, I'd love to see it, maybe I could learn a thing or 2. [1]: http://zipfianacademy.com/ http://zipfianacademy.com/ [2]: http://www.cs.toronto.edu/~fritz/absps/esann-deep-final.pdf http://www.cs.toronto.edu/~fritz/absps/esann-deep-final.pdf [3]: https://bigml.com/ https://bigml.com/ [4]: http://static.googleusercontent.com/media/research.google.com/en/us/archive/large_deep_networks_nips2012.pdf http://static.googleusercontent.com/media/research.google.co... [5]: http://hadoopsummit.org/san-jose/schedule/ http://hadoopsummit.org/san-jose/schedule/ [6]: http://www.oscon.com/oscon2014/public/schedule/detail/33709 http://www.oscon.com/oscon2014/public/schedule/detail/33709 [7]: http://www.jcuda.org/jcuda/jcublas/JCublas.html http://www.jcuda.org/jcuda/jcublas/JCublas.html [8]: http://developer.amd.com/tools-and-sdks/heterogeneous-computing/aparapi/ http://developer.amd.com/tools-and-sdks/heterogeneous-comput...