3 ms·
Pardon my ignorance but I have a few questions... 1) This seems like it would only work for varying the values passed to variables in a function. Can you diff
by bcheung 7y ago
Pardon my ignorance but I have a few questions...
1) This seems like it would only work for varying the values passed to variables in a function. Can you differential a function other than just changing some values -- like using a genetic algorithm. Seems like this is just finding the right values to pass to a function not finding a different function itself.
2) Is this limited to only continuous / ordered values? How can you differentiate logic, branching, and what other functions to call? Seems like to do so would require mapping a function to some kind of n-dimensional euclidian space / manifold.
3) Why does this need to be an extension to a language? Can't ordinary interfaces in the OOP sense or monads in the FP sense be used to wrap functions and give them this functionality for free?
- ddragon 7y ago1) It's not restricted to the variables being passed, for example Zygote [1] has an example in the README with IO gradient(x -> fs[readline()](x), 1), and it's not using numerical differentiation (varying inputs to check the output), but finding a formula (approximate, not a closed form like a symbolic differentiation) of how changing the output affects the input. Genetic algorithms is a black box metaheuristic, so people would favor using gradient descent given that the code is differentiable, but how the model is optimized is technically open to any method. 2) Yes, it can differentiate all control flow. What happens is that each forward/backward pass will have a different graph based on what branches it would follow (effectively a subderivative). Each one of these graphs is independently differentiable and do not contain the control flow by themselves. 3) It doesn't need a language extension (Julia doesn't, but that's because you can access Julia IR using the language itself, not many language support this level of metaprogramming). The OOP strategy (like Pytorch) overload methods so every time you call the overloaded method it builds the graph. It depends on customized types and methods and does not support code that isn't using those custom types that can hold the graph (native/custom types), native control flow (only implicitly by changing what graph is being created), reduces the possible optimizations (you only have one subderivative, not the complete graph to optimize) and end up creating a DSL with it's own error messages and quirks that is less natural than just using the host language. I can't say how it would look using monads though. [1] https://github.com/FluxML/Zygote.jl https://github.com/FluxML/Zygote.jl