4 ms·
In my opinion the difference that Hinton students made was using GPUs to run deep learning models. First in 2010, when publishing Theano at NIPS, making deep l
by dchichkov 4y ago
In my opinion the difference that Hinton students made was using GPUs to run deep learning models. First in 2010, when publishing Theano at NIPS, making deep learning on a single computer node with a large GPUs very approachable ( using a Python framework, precursor to TensorFlow). Then in 2012, with Alex Krizhevsky masterfully using two GPUs to build the landmark model AlexNet, which outperformed traditional computer vision models and was easy to use in practice.
- ekleraki 4y agoFunny ... According to [1], the papers [42,43] showed that plain backprop was capable of solving complex tasks without pretraining on a GPU. From [1]: > Our team then showed that good old backpropagation [A1] on GPUs (with training pattern distortions [42,43] but without any unsupervised pre-training) I can't verify since I have a barely stable connection and little battery, but somebody else should be able to. [1] https://people.idsia.ch/~juergen/firstdeeplearner.html https://people.idsia.ch/~juergen/firstdeeplearner.html [42] H. Baird. Document image defect models. IAPR Workshop, Syntactic & Structural Pattern Recognition, p 38-46, 1990 [43] P. Y. Simard, D. Steinkraus, J.C. Platt. Best Practices for Convolutional Neural Networks Applied to Visual Document Analysis. ICDAR 2003, p 958-962, 2003.
- catgary 4y agoIsn’t backprop just reverse-mode AD, which has been around since the 60’s? I get the impression that a lot of physicists and scientific computing researchers have been rolling their eyes at various ML “breakthroughs” over the years.
- ekleraki 4y agoMy point was not backprop, but time of publication. However, "reverse-mode" AD is not as simple as just caching results. > BP’s continuous form was derived in the early 1960s (Kelley, 1960; Bryson, 1961; Bryson and Ho, 1969). Dreyfus (1962) published the elegant derivation of BP based on the chain rule only. > explicit, efficient error backpropagation (BP) in arbitrary, discrete, possibly sparsely connected, NN-like networks But that wasn't close enough to current BP > BP’s modern efficient version for discrete sparse networks (including FORTRAN code) was published by Linnainmaa (1970). Here the complexity of computing the derivatives of the output error with respect to each weight is proportional to the number of weights. That’s the method still used today. See [1] for more details. I do not know the exact implementation details of Linnainmaa's BP, but there are all sorts of improvements to plain reverse mode auto-grad through stuff like Jacobian-Vector products that save an enormous amount of memory. Nocedal's "Numerical Optimization" covers a great deal on how to implement autograd, and I can easily recommend it to anyone interested. [1] https://www.reddit.com/r/MachineLearning/comments/e5vzun/d_jurgen_schmidhuber_on_seppo_linnainmaa_inventor/ https://www.reddit.com/r/MachineLearning/comments/e5vzun/d_j...
- ShamelessC 4y agoSorry what's your contention? or do you have one? Backprop is backprop on CPU or GPU - it just runs (way) faster on GPU. So while you can do some stuff with NN's on CPU, you can do more stuff with NN's on GPU. Further, the sources you reference are from before advances made in software and hardware for massively parallel GPGPU architectures (CUDA), so it's kind of a moot point anyways. Those researchers may have had access to such a system, but you definitely couldn't just `pip install pytorch` on your gaming GPU in the 90's...
- deepnotderp 4y agoDan Ciresan’s DanNet ran on GPUs before Hinton’s students
- P-NP 4y agoSchmidhuber's DanNet pages: https://people.idsia.ch/~juergen/DanNet-triggers-deep-CNN-revolution-2011.html https://people.idsia.ch/~juergen/DanNet-triggers-deep-CNN-re... https://people.idsia.ch/~juergen/2010-breakthrough-supervised-deep-learning.html https://people.idsia.ch/~juergen/2010-breakthrough-supervise...
- ekleraki 4y ago> Sorry what's your contention? Time of publication. I merely pointed that Schmidhuber et al's results were published long before Hinton et al's work (I included year of publication for a reason), ergo the comment that was in support of Hinton was refuted and my comment gives further validity to Schmidhuber's claims of misattribution and plagiarism, while pointing at yet another paper tiger in the claims that the three pioneered the field. If anything, Schmidhuber may have deserved a Turing award on his own, and what we see is the result of politics. This is pure speculation, but to me the whole ordeal hints at manipulation from the companies where 2 of the 3 worked at, which would essentially increase the pool of candidates the respective companies have only by association with their employee.
- ShamelessC 4y agoAh, that makes sense. Thanks for clarifying.
- breuleux 4y agoSmall correction, Theano was made in Bengio's lab.