10 ms·
“Deep Learning has outlived its usefulness as a buzz-phrase”
- deleted 9y ago[deleted]
- rememberlenny 9y ago[Text from post] OK, Deep Learning has outlived its usefulness as a buzz-phrase. Deep Learning est mort. Vive Differentiable Programming! Yeah, Differentiable Programming is little more than a rebranding of the modern collection Deep Learning techniques, the same way Deep Learning was a rebranding of the modern incarnations of neural nets with more than two layers. But the important point is that people are now building a new kind of software by assembling networks of parameterized functional blocks and by training them from examples using some form of gradient-based optimization. An increasingly large number of people are defining the network procedurally in a data-dependant way (with loops and conditionals), allowing them to change dynamically as a function of the input data fed to them. It's really very much like a regular progam, except it's parameterized, automatically differentiated, and trainable/optimizable. Dynamic networks have become increasingly popular (particularly for NLP), thanks to deep learning frameworks that can handle them such as PyTorch and Chainer (note: our old deep learning framework Lush could handle a particular kind of dynamic nets called Graph Transformer Networks, back in 1994. It was needed for text recognition). People are now actively working on compilers for imperative differentiable programming languages. This is a very exciting avenue for the development of learning-based AI. Important note: this won't be sufficient to take us to "true" AI. Other concepts will be needed for that, such as what I used to call predictive learning and now decided to call Imputative Learning. More on this later....
- bra-ket 9y agoit's really a pity that after 75 years of AI research the best thing we've got is still based on gradient descent, a brute force trial and error.
- jclos 9y agoIts simplicity is its power. More complex methods (e.g. second order methods) tend to get attracted to saddle points and produce bad results. Some metaheuristics like evolution strategies are also used in some specific cases (reinforcement learning). Minibatch gradient descent + reasonable minibatch size + some form of momentum is the best we have.
- log_base_login 9y agoAs much of a pity that, 70 years later, we are still using transistor based computers originally derived from three wires stuck in a piece of rock[1] by some very innovative fellows at Bell Labs[2]? [1]http://images.computerhistory.org/revonline/images/500004836-03-01.jpg?w=600 http://images.computerhistory.org/revonline/images/500004836... [2]http://www.computerhistory.org/revolution/digital-logic/12/273 http://www.computerhistory.org/revolution/digital-logic/12/2...
- zeth__ 9y agoOur transistors have nothing to do with those transistors.
- pacavaca 9y agoAssuming that AI tries to mimic the way humans learn and evolve, those methods haven't changed for hundreds of thousands of years and brute-force trial and error is just one of them. It's kind of fundamental...
- comstock 9y agoIt doesn’t really try to mimic the way humans learn and evolve...
- maxtollenar 9y agoMostly because the loss function space is not well understood, we need to do some kind of descent
- deleted 9y ago[deleted]
- platz 9y ago> It's really very much like a regular progam, except it's parameterized, automatically differentiated, and trainable/optimizable. > People are now actively working on compilers for imperative differentiable programming languages. Do you have an example of either of these things?
- zengid 9y agoFrom the thread, LeCun mentions: "Look at papers by Jeff Siskind and Barak Barak A. Pearlmutter particularly their work on VLAD, Stalingrad and VLD."
- joshuamorton 9y agoI believe https://github.com/google/tangent https://github.com/google/tangent counts, though I'm not 100% sure.
- chombier 9y agoErik Meijer gave a talk about this at KotlinConf last year https://www.youtube.com/watch?v=NKeHrApPWlo https://www.youtube.com/watch?v=NKeHrApPWlo
- saycheese 9y agoPast HN coverage of Differentiable Programming: https://news.ycombinator.com/item?id=10828386 https://news.ycombinator.com/item?id=10828386
- ehsankia 9y ago1. Differentiable Programming is horrible branding. It's hard to say, not catchy, and not as easily decipherable. 2. Isn't the evolution of Deep Networks more advance setups such as GANs, RNNs, and so on?
- bennofs 9y agoI like it, at least from the little that I read about it. The name describes the core of what it is: differentiate programs, in order to figure out how changes the the program affect the output and using that for optimization purposes. Do we really need to invent obscure, new names for everything just so that it sounds catchy?
- tomdre 9y agoCouldn't agree more. A technique should be judged by its usefulness. Not by its catchiness.
- pasquinelli 9y agoto be fair though, the subject is a judgement of a technique's name, not a judgement of the technique itself.
- wastewaste 9y agonot to forget, naming is one of the harder problems considered in programming.
- kinkrtyavimoodh 9y ago> Differentiable Programming is horrible branding. It's hard to say, not catchy, and not as easily decipherable Tell that to the people who deliberately popularized the term Dynamic Programming for something that was neither dynamic nor programming. ____ (From Wiki) Bellman explains the reasoning behind the term dynamic programming in his autobiography, Eye of the Hurricane: An Autobiography (1984, page 159). He explains: "I spent the Fall quarter (of 1950) at RAND. My first task was to find a name for multistage decision processes. An interesting question is, Where did the name, dynamic programming, come from? The 1950s were not good years for mathematical research. We had a very interesting gentleman in Washington named Wilson. He was Secretary of Defense, and he actually had a pathological fear and hatred of the word research. I’m not using the term lightly; I’m using it precisely. His face would suffuse, he would turn red, and he would get violent if people used the term research in his presence. You can imagine how he felt, then, about the term mathematical. The RAND Corporation was employed by the Air Force, and the Air Force had Wilson as its boss, essentially. Hence, I felt I had to do something to shield Wilson and the Air Force from the fact that I was really doing mathematics inside the RAND Corporation. What title, what name, could I choose? In the first place I was interested in planning, in decision making, in thinking. But planning, is not a good word for various reasons. I decided therefore to use the word “programming”. I wanted to get across the idea that this was dynamic, this was multistage, this was time-varying. I thought, let's kill two birds with one stone. Let's take a word that has an absolutely precise meaning, namely dynamic, in the classical physical sense. It also has a very interesting property as an adjective, and that it's impossible to use the word dynamic in a pejorative sense. Try thinking of some combination that will possibly give it a pejorative meaning. It's impossible. Thus, I thought dynamic programming was a good name. It was something not even a Congressman could object to. So I used it as an umbrella for my activities."
- danbmil99 9y ago"See more of Yann LeCun on Facebook" popup, no access to the page. No, I don't want to create a Facebook account to read a blog post. Perhaps links to walled-garden pages where you need an account and need to be logged in should be prohibited or at least discouraged.
- dizzystar 9y agoAgreed. I think many posters here assume that everyone has an account at FB, WSJ, NYT, and so on. The web search trick doesn't work for all of these situations either.
- nkozyra 9y agoI had an option to click "not now" to dismiss.
- username223 9y agoYou need to up your crap-blocking game. This bookmarklet can help: javascript:(function() { (function () { var i, elements = document.querySelectorAll('body *'); for (i = 0; i < elements.length; i++) { if (getComputedStyle(elements[i]).position === 'fixed') { elements[i].parentNode.removeChild(elements[i]);}}})()})()
- kinkrtyavimoodh 9y agoI opened it in an Incognito window and it worked fine. Public posts on FB can be seen without logging in.
- tedivm 9y agoHow long did you wait on the page? At first when it loads you can read it just fine, but then an obnoxious pop up appears. There is a little "not now" link on the bottom that will get you back to the post, but I can see how someone would miss that (it's a great example of a bad UI).
- wadkar 9y ago
- oh-kumudo 9y agoJust a rebranding, though necessary one. Deep Learning is not really all about 'deep' anymore, many successful models don't really need a lot of layers.
- username223 9y agoSo it's still just neural nets? Cool -- we've seen that before.
- woodson 9y agoI’m in favor of getting rid of „neural“. That would help dispense with all the unhelpful discussions about how different DNNs work compared to the human brain.
- jclos 9y agoI agree. I've taken to calling them computational graphs in my lab, because that's what they are. There is nothing neural about a LSTM.
- rspeer 9y agoIs a single SGD layer a neural net? Is an image filter or an audio filter a neural net? Is matrix multiplication a neural net? This would strain the intended definition even farther than it's already been strained. But all of those are differentiable programming, and rightly so because they're all pieces that you use and compose together to make interesting learning mechanisms, including the ones that we vaguely refer to as "deep learning" now. I like the terminology. It's not about what the original long-abandoned motivation for the design was ("neural"). It's not about how gratuitously complex you can make it ("deep"). "Differentiable" is about how it works and how we design it.
- nightski 9y agoI'm not really buying your argument here. Neural networks are just a collection of artificial neurons. There is no requirement for multiple layers or depth of any kind.
- bitL 9y agoDarn! Just when I invested a lot of money and mastered Deep Learning!
- elchief 9y agoRemember back in 2017 when Deep Learning wasn't legacy? Those were good times
- tomdre 9y agoOff topic. Yann LeCun really looks like Michael Moore who looks like Peter Griffin.
- umanwizard 9y agoNo he doesn't, but who cares?
- dang 9y agoPlease don't post unsubstantive comments here.
- billconan 9y ago"working on compilers for imperative differentiable programming languages" what would be an example of such language?
- currymj 9y agoworking with dynamic neural net libraries that have autodiff (PyTorch, regular Torch, and I think also MXNet now) feels a lot like this. you just write normal functions, but after you execute them you can ask for gradients too.
- sedachv 9y agoCommon Lisp without any changes: https://people.eecs.berkeley.edu/~fateman/papers/ADIL.pdf https://people.eecs.berkeley.edu/~fateman/papers/ADIL.pdf Fortran with some changes required: http://www.ens.utulsa.edu/~diaz/cs8243/adifor.html http://www.ens.utulsa.edu/~diaz/cs8243/adifor.html C with some changes required: http://www.ens.utulsa.edu/~diaz/cs8243/adiff.html http://www.ens.utulsa.edu/~diaz/cs8243/adiff.html
- seanmcdirmid 9y agoBrainScript was one, though it seems deprecated ATM: https://docs.microsoft.com/en-us/cognitive-toolkit/brainscript-basic-concepts https://docs.microsoft.com/en-us/cognitive-toolkit/brainscri... Compilers might not be involved, interpreters can work just as well if most of the heavy lifting is in the solving.
- BucketSort 9y agoI believe this paper by Marcus ( https://arxiv.org/ftp/arxiv/papers/1801/1801.00631.pdf https://arxiv.org/ftp/arxiv/papers/1801/1801.00631.pdf ) earlier this week inspired this. Edit: I don't mean Marcus inspired the term differentiable programming; he inspired LeCun to emphasize the wider scope of deep learning after Marcus attacked it. In fact, LeCun liked a post on twitter rebutting Marcus' paper that also talks about differentiable programming: https://twitter.com/tdietterich/status/948811917593780225 https://twitter.com/tdietterich/status/948811917593780225
- seanmcdirmid 9y agoOr the other way around... He has been throwing around the term for awhile now.
- BucketSort 9y agoIt was most certainly the other way around. Marcus does not focus on differentiable programming in his paper. See my edit.
- sytelus 9y agoI don't think so. LeCun seems to oppose Marcus's views... Related: https://twitter.com/ylecun/status/921409820825178114?lang=en https://twitter.com/ylecun/status/921409820825178114?lang=en I think LeCun doesn't want a repeat of AI winter because of exponentially rising hype and expectations out of Deep Learning. There have been few examples like Selena which he seems to think that people are trying to ride the deep learning wave to generate false buzz (and cash!) for themselves.
- BucketSort 9y agoHe does oppose Marcus' views, but he also knows neural nets are only one approach to differentiable programming. The term is confusing though. It should read like "linear programming" does, but people are not interpreting it that way.
- Ormus 9y ago
- currymj 9y agoi remember people jokingly referring to stuff like word2vec (one layer, millions of dimensions) as "wide learning". this is definitely better branding than that.
- fjsolwmv 9y ago"Google uses Bayes nets like Microsoft uses 'if' statements" -+ Joel Spolsky, 15 years ago
- BucketSort 9y agoYou may have seen DeepMind's results last year where it trained 3D models to move through space in different ways, entitled "Emergence of Locomotion Behaviours in Rich Environments" ( https://arxiv.org/pdf/1707.02286.pdf https://arxiv.org/pdf/1707.02286.pdf , https://www.youtube.com/watch?v=hx_bgoTF7bs&feature=youtu.be https://www.youtube.com/watch?v=hx_bgoTF7bs&feature=youtu.be). If you have a look in the paper, the method they use "Proximal Policy Optimization" is a great example of differentiable programming that does not include a neural network. I actually realized this last month when I was preparing a talk on deep learning, because I thought it used deep neural nets in its application, but found that it didn't.
- guillefix 9y agoScanning through the paper, I see this "We structure our policy into two subnetworks, one of which receives only proprioceptive information, and the other which receives only exteroceptive information. As explained in the previous paragraph with proprioceptive information we refer to information that is independent of any task and local to the body while exteroceptive information includes a representation of the terrain ahead. We compared this architecture to a simple fully connected neural network and found that it greatly increased learning speed." It seems to me they do use neural nets. Proximal Policy Optimization is just a more novel way of optimizing them.
- everdev 9y agoIf you're going to do a "rebrand" atleast use a better name. A six syllable word doesn't exactly roll off the tongue.
- acostin 9y agoFrom DNN to DPP. Differential Programming Patterns?
- adamnemecek 9y agodifferentiable.js.io
- seanmcdirmid 9y agoI'm seriously pondering what this means for PL research. There has been some work in probalistic programming languages, and a significant part of the community would like to avoid imperative features. However, it seems like this is a chance for some real invigoration in PL research agendas.
- cs702 9y agoI wish we could come up with a catchier name, but I LOVE the idea of calling this programming, because that is precisely what we do when we compose deep neural nets. For example, here's how you compose a neural net consisting of two "dense" layers (linear transformations), using Keras's functional API, and then apply these two layers to some tensor x to obtain a tensor y: f = Dense(n) g = Dense(n) y = f(g(x)) This looks, smells, and tastes like programming (in this case with a strong functional flavor), doesn't it? Imagine how interesting things will get once we have nice facilities for composing large, complex applications made up of lots of components and subcomponents that are differentiable, both independently and end-to-end. Andrej Karpathy has a great post about this: https://medium.com/@karpathy/software-2-0-a64152b37c35 https://medium.com/@karpathy/software-2-0-a64152b37c35
- bmc7505 9y agoI hate the name, but LOVE the idea of calling this programming... What would you call it instead?
- seanmcdirmid 9y agoLeCun specifically calls out imperative programming, not just typical data flow methods.
- cs702 9y agoYou're right. I softened the reference to functional programming.
- contextfree 9y agoDoes this mean programming language nerds get to play too, maybe after boning up on our calculus and topology?
- seanmcdirmid 9y agoThey already are; e.g. Jeff Dean among many others. The question is will the PL academic community play as well. Conal Elliott has done a lot of work in this area about 10 years ago. His work is beautiful but maybe before it’s time.
- sitkack 9y agoCan you expand on this comment, I don't understand.
- flor1s 9y agoInteresting comment regarding Conal Elliott. I've always thought there is some similarity between probabilistic programming (specify a probabilistic model as a graph), functional reactive programming (specify some reactivity as a graph) and deep learning (specify some linear algebra / calculus / optimization operations as a graph). Too bad the word "graphical programming" would be interpreted as "visual programming" (or programming using plots and charts!) and not programming using an explicit graph structure.
- outlace 9y agoBut differentiable programming would exclude deep neural nets trained by evolutionary methods/genetic algorithms since those are gradient free. With the term deep learning I think the focus is correctly on the “deep” (compositional) nature of these models and not necessarily the training algorithm, of which there are many.