2 ms·
It's a shame it comes across that way, since the article and colah's blog past are talking about completely different concepts. colah's blog post can be roughl
by throwaway2322 11y ago
It's a shame it comes across that way, since the article and colah's blog past are talking about completely different concepts.
colah's blog post can be roughly summarized as taking the (well-known) representation of NNs as DFGs of modules (since the very early 90s, if not earlier), and then showing how common combinators in functional programming languages correspond to subgraphs in common modules. Since both FPs and DNNs can be represented as DFGs, this isn't a particularly surprising (or novel) correspondence.
Several libraries used exactly these functional combinators for symbolically expressing NNs well before his blog post.
The trend of differentiable programming (as in, using gradient-based methods to train a network to learn an algorithm) is what I think the author is trying to highlight. This has been around since e.g. Das et. al, "Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory" (from 1992). There's a lot of history (and duplication) in this area, some keywords to search for being NTMs, RL-NTMs, Pointer Networks, Neural GPU, MemNNs, MemN2N, Stack RNNs, .... - the RAM (reasoning, attention, memory) workshop at NIPS 2015 had a bunch of recent work in this area.
The high-level idea is by augmenting a (trained) controller (e.g. an LSTM) with some external datastructure(s) (stacks, queues, memory, hierarchical memory) with the ability to read from and write to the external datastructure(s), and to propagating errors back to the controller (hence differentiable programming moniker), via hard (e.g. RL) or soft attention mechanisms.