3 ms·
After reading "Backprop as Functor: A compositional perspective on supervised learning", it came to my attention that the spaCy backend (thinc) is built with hi
by f00_ 9y ago
After reading "Backprop as Functor: A compositional perspective on supervised learning", it came to my attention that the spaCy backend (thinc) is built with higher order functions instead of a computational graph (unlike tensorflow, chainer, or pytorch)
https://github.com/explosion/thinc#no-computational-graph--just-higher-order-functions https://github.com/explosion/thinc#no-computational-graph--j...
https://arxiv.org/abs/1711.10455 https://arxiv.org/abs/1711.10455
Could someone give me some more detail?
- syllogism 9y agoYou know, I'm still not 100% certain whether there's a substantive difference between the "computational graph" perspective and this "functor" approach. Actually the feeling is sort of eerily familiar, because I spent most of my PhD confused about whether these grammar formalisms I was working with where really just notational variants, or whether there were significant differences. About the grammar formalisms, I ended up deciding that in theory there wasn't, in practice there sort of was. About these neural networks, I think it's "just" implementation. Here's the linear layer implementation in Chainer: https://github.com/chainer/chainer/blob/master/chainer/functions/connection/linear.py https://github.com/chainer/chainer/blob/master/chainer/funct... We have the forward and backward pass organized as class methods here, and the intermediate state from the forward pass is saved into attributes in the instance. So on each call to the network, we make an instance of this LinearFunction class. In terms of what's being computed, there's really no difference between this and what happens when you call a layer in Thinc. It's just that the state gets captured in the outer scope of the closure. Maybe Thinc's way has a little less overhead, if there are fewer levels of indirection. Thinc uses the Chainer folks' GPU library --- so, unsurprisingly if you define the same network, the benchmarks are very similar. On the other hand...I do think the implementation matters! Here's a difference for you: if the library approaches it as "we're going to build a computational graph, and execute it", then the library is going to steal the control flow. If the library tells you "here are some functions, and some higher order functions to compose them", you have more access. PyTorch and Chainer doesn't steal the control flow to nearly the extent that Tensorflow does, but they still build up and tear down the state in their objects, and that makes it harder to intrude. (I'm the author of spaCy and Thinc)
- jph00 9y agoTo clarify the grandparent comment : pytorch also supports a functional approach rather than a computational graph. There's even a pytorch functional model zoo :)