5 ms·
I find differentiable programming languages really fascinating. Think about this: a differentiable programming language is still a programming language. If the
by muds 4y ago
I find differentiable programming languages really fascinating. Think about this: a differentiable programming language is still a programming language. If the language is designed to facilitate a smooth optimization landscape, it's actually possible to "learn" programs with gradient descent. This opens the door to a lot of cool possibilities:
- programming languages which use neural networks as primitive functions (think `result = sum([mlp(input) for input in list])`. NN's are (understandably) notoriously bad at learning simple operators [1]. Differentiable programming over a language defined by aggregation functions (map/fold/sum/mean/etc.) allows us to bypass learning some simple functions.
- Flipping this around, we can use neural networks that use differentiable programs to regularize the outputs. Assume we have a NN that learns the speed of a car from a video. We know that a car's speed cannot exceed (say) 200mph. Make a differentiable program to express this and use it to regularize the output of the network.
- Reusing the image->NN->speed example again, use the differentiable program to identify speeds/conditions where using a neural network policy is unsafe and switch to a (less-performant) handmade policy instead.
Some more thoughts about this: https://atharvas.prose.sh/differentiable_dsls https://atharvas.prose.sh/differentiable_dsls
[1] https://dselsam.github.io/posts/2018-09-16-neural-networks-occams-razor.html https://dselsam.github.io/posts/2018-09-16-neural-networks-o...
- mjburgess 4y agoTwo points, (1) everything which makes programs useful is impure device access and state change, discretely sequenced over time (2) grad. desc. et al. do not learn discrete constraints (hence why NNs are bad at learning operators: they cant. x+x is defined fa. x; not fa x. in the training set).
- muds 4y ago> everything which makes programs useful is impure device access and state change, discretely sequenced over time I haven't heard about this before actually. I'd love to hear more about this! "Impure," here, is PL terminology for functions that affect global state/arguments when you run them. right? So, brainstorming a bit, what this means is that making a diff. programming language that treats a NN module as a pure function won't actually be beneficial? I'm not sure if I'm drawing the correct conclusion but this is a really interesting point. Don't have an answer for this (yet!). > grad. desc. et al. do not learn discrete constraints Great Point! To push back a little on this. You're right that any discrete constraint will always mess up the smoothness of the function (eg: less-than-g is not smooth at x=g). However, we can engineer our way around this by relaxing a discrete constraint to its closest smooth approximation! So, we can implement the less-than-g function as a sigmoid that is shifted by +/-g. This introduces a parameter to control the slope of the sigmoid. In practice, I haven't had much difficulty learning programs even with a really steep slope for the sigmoid.
- mjburgess 4y ago(1) Yes, the modern ML/AI lot seem to ambiguously use a purely mathematical meaning to "computer" -- which is useless. As useless as any pure mathematics. If we only had this a "computer" would be a theoretical curiosity, like a 200-dim sphere. The real-world computers we care about run algorithms whose semantics is given by the properties of the devices real computers use. This double meaning to "computer" has caused a lot of superstition in the ML/AI space. Real computers are engineering devices which shuffle electrical signals around to useful devices. There is no reason to think that "pure algorithms" have any use at all, as with, eg., a 200-dim sphere. They're only useful if they can be given a semantics which exploits useful properties of devices. (cf. with physics, where a 200-dim sphere could be useful if it models some actual system). (2) This isn't enough. Consider learning the rules of chess; or likewise, the inference rules of mathematics. f(x) = 2x^2, f'(x) = 4x, etc. Search spaces constructed for a grad. desc. search are very infinite; and the solutions we need are infinitely precise. Discrete approaches to search(ing for solutions) are necessary.
- data_maan 4y ago> As useless as any pure mathematics Are you hearing yourself talk? Do you know why you have (to just name one example out of many) thousands of pictures on your phone, and not just a few? Because of pure mathematics. Because of compression: Even JPEG2000 from back in the day uses intricate and beautiful compression algorithms based on wavelets.
- usgroup 4y agoI think this is partially true; there is some support for logical statements and control flow in differentiable programs -- at least in Jax. Further, think Deep Mind have a recent paper on a DL sequence learning methodology able to learn control strategies for lots of games simultaneously. I think this is a good example of learning discrete constraints with an NN.
- Sirenos 4y agoTwo counter-points (appreciate some counter-counter-points :): 1) The discrete sequencing is an epiphenomenon. The underlying processes are continuous changes in voltage and current flows. (I'm not sure if Planck scale considerations can throw a wrench in this though. Would love to be educated here.) 2) Our brains do not have ostensibly discrete neural processors. I don't think gradient descent is comparable to how the brain learns, but I think there is some reason to think that it is possible to learn symbolic processes in spite of having a processor that isn't especially built for it.
- mjburgess 4y agoYou're making a genetic fallacy here: that since the origin/ground of something has property C, it's product must have it too. This isnt so. I agree that reality is fundamentally continuous. However cognition isnt; and many things arent. A frequency is discrete. A length is continuous. These properties aren't eliminable for one another. Here, whilst i'd agree that all physical process going on (everywhere) have essential continuous properties; they also have essential discrete ones. The issue is that grad. desc. alone does not give you the right kind of discrete ones.