3 ms·
He’s saying that PyTorch in essence is a DSL embedded in python which introspects itself to perform automatic differentiation. Libraries can’t normally peek at
by mlazos 3y ago
He’s saying that PyTorch in essence is a DSL embedded in python which introspects itself to perform automatic differentiation. Libraries can’t normally peek at what ops are performed on variables in a given language, only the compiler/runtime for that language can. In the end it’s semantics - is an overloaded operator a DSL or native language construct? I’d argue it’s a DSL because its definition is in the host language.
- credit_guy 3y agoYou can implement autograd as a library. Just take a look at this https://github.com/sradc/SmallPebble https://github.com/sradc/SmallPebble The first line of the description is: > SmallPebble is a minimal automatic differentiation and deep learning library written from scratch in Python, using NumPy/CuPy. The author of this library wrote a blog post, which is in my opinion the best introduction to reverse mode automatic differentiation that exists: https://sidsite.com/posts/autodiff/ https://sidsite.com/posts/autodiff/
- kragen 3y agothis is a nice blog post, but, two things: - it's very much taking the edsl approach - it's only implementing forward-mode
- credit_guy 3y agoIt is reverse mode, not forward mode. As for the edsl, you can claim it is, but why? There is no parser, no syntax, nothing that would make any user think they are learning a new language. Would you call numpy an edsl ?
- kragen 3y agothe text says it's reverse mode, but it's wrong. what the code in the post implements is forward mode embedded dsls don't have parsers; that's what makes them embedded. numpy is solidly in the center of the embedded dsl concept
- credit_guy 3y ago> the text says it's reverse mode, but it's wrong. what the code in the post implements is forward mode I see. It's actually worse than both forward and reverse mode. Let's say you have the simple program x1+x2+x3+x4+x5. Internally the program calculates x6=x1+x2, x7 = x6+x3, x8=x7+x4 and result = x9 = x8+x5. In forward mode you get all the partial derivatives for all intermediate variables w.r.t the inputs x1 through x5, so you get things like dx8/dx2, but you never calculate (or store in memory) dx8/dx6. In reverse mode you get the derivatives of the result w.r.t. all intermediate variables, so things like dx9/dx8, dx9/dx7, etc, but again you never need to calculate dx8/dx6. In this library you calculate recursively all dxj/dxi and store them, as long as xj has a direct or indirect dependency on xi. This is a strict superset of the derivatives calculated in both forward and reverse mode. Good catch.
- kragen 3y agohmm, really? i didn't realize that