4 ms·
Wasn't that the paradigm behind Tensorflow?
by dynamite-ready 3y ago
Wasn't that the paradigm behind Tensorflow?
- ggr2342 3y agoAFAIK it was first brought forward by PyTorch. I may be wrong though.
- 6gvONxR4sf7o 3y agoIt's been a core part of automatic differentiation for decades.
- ggr2342 3y agoOh! Can you give some pointers where to read about it more?
- PartiallyTyped 3y agoThe original 2017 PyTorch paper is probably a good one.
- 6gvONxR4sf7o 3y agoThe wikipedia page for automatic differentiation and its references would be a good start. https://en.wikipedia.org/wiki/Automatic_differentiation https://en.wikipedia.org/wiki/Automatic_differentiation In particular, the section "Beyond forward and reverse accumulation" tells you how hard it is to deal with the graph in optimal ways, hence the popularity of simpler forwards and backwards traversals of the graph.
- cgearhart 3y agoIt certainly pre-dates PyTorch. Applying computational graphs to neural networks was the central purpose of the Theano framework in the early 2010s. TensorFlow heavily followed those ideas, and Theano shut down a year or two later because they were so close conceptually and a grad research lab can’t compete with the resources of Google.
- mhh__ 3y agoPyTorch's insight was to be very dynamic and simple. The underlying model is basically the same with the caveat that tensorflow uses/used something called a "tape".
- HarHarVeryFunny 3y agoBefore PyTorch there was Torch (which was Lua-based, hence the Py- prefix of the follow on PyTorch). With Torch, like the later TensorFlow 1.0, you first explicitly constructed the computational graph then executed it. PyTorch's later innovation was define-by-run where you just ran the code corresponding to your computational graph and the framework traced your code and built the graph behind the scenes... this was so much more convenient that PyTorch quickly became more popular than TensorFlow, and by the time TF 2.0 "eager mode" copied this approach it was too late (although TF shot itself in the foot in many ways, which also accelerated it's demise).
- dkislyuk 3y agoAs another commenter said, viewing a neural network as a computation graph is how all automatic differentiation engines work (particularly reverse-mode where one needs to traverse through all the previous computations to correctly apply the gradient), and there were several libraries predating Tensorflow following this idea. The initial contribution of Tensorflow and PyTorch was more about making the developer interface much cleaner and enabling training on a wider range of hardware by developing a bunch of useful kernels as part of the library.
- dekhn 3y agoI always thought of seeing most computations of functions as computational graphs- at least, when I used MathLink to connect Mathematica to Python, it basically gives you a protocol to break any Mathematica function into its recursively expanded definition. Konrad Hinsen suggested using python's built-in operator overloading, so if you said "1 + Symbol('x')" it would get converted to "Plus(1, Symbol('x')) and then sent over MathLink to Mathematica, which would evaluate Plus(1, x), and return an expression which I'd then convert back to a Python object representation. I don't think we talked about doing any sort of automated diff (in my day we figured out our own derivatives!) but after I made a simple eigendecomp of a matrix of floats, the mathematica folks contributed an example that did eigendecomp of a matrix with symbols (IE, some of the terms weren't 5.7 but "1-x"). Still kind of blows my mind today how much mathematica can do with computation graphs. IIUC this is the basis of LISP as well.
- dkislyuk 3y agoThe one distinction I would add with neural networks is that it's not just a recursive tree traversal that one would get when evaluating an arithmetic statement, but an actual graph: a computation node can have gradients from multiple sources (e.g. if a skip connection is added), so each node needs to keep accumulated state around that can be updated by arbitrary callers. Of course, optimized autograd / autodiff is more parallelized than node-based message passing, but it's a useful model to start with.
- 3y ago