6 ms·
Dynamic Memory Networks for Visual and Textual Question Answering
- inlineint 11y agoIt's fascinating. Could anybody answer what programming language is most likely to be used for development of this kind of systems? I haven't found information about it in the paper.
- olalonde 11y agoPython seems to be the go to language for anything AI/machine learning.
- ogrisel 11y agoPython and lua both have good frameworks for developing such models (tensorflow, theano possibly with lasagne or keras, torch, caffe...). The visual part need access to efficient GPU-based implementation of convolutions (typically using cuDNN from nvidia or neon from Nervana Systems). You could write bindings for other languages or even work directly in C++ but I think that an interactive programming language with a REPL like Python and lua is very helpful for faster iterative model building, interactive exploration and evaluation.
- inlineint 11y agoThank you for detailed answer. I didn't know about the frameworks you listed except tensorflow. Btw do you know anybody who use Julia for this?
- ogrisel 11y agoThere is https://github.com/pluskid/Mocha.jl https://github.com/pluskid/Mocha.jl but it seems more focused on vision tasks at least for now.
- nicklo 11y agoMetaMind papers are always pretty awesome. Couple of great highlights from this paper: The visual saliency maps in Figure 6 are astounding and make their performance on this task even more impressive as they give a lot of insight into what the model is doing and it seems to be focusing on the things in the image that a regular person would use to decide on the answer. Most striking was on the question "is this in the wild?", and the saliency was on the artificial, human structures in the background that indicated it was in a zoo. This type of reasoning is surprising as it requires a bit of a reversal in logic to come up with this way of answering the question. Super impressive. The proposed input fusion layer is pretty cool - allowing information from future sentences to be used to condition how to processes previous sentences. This type of information combining was previously not really explored, and it makes sense that it improves the performance on the bAbI-10k task so much as back-tracking is an important tool for human reading comprehension. Also its clever that they encode each sentence as a thought vector before compositing so they can be processed both forwards and backwards with shared parameters- doing so on just one-hot words or even ordered word embeddings would require two vastly different parameters since grammar is wildly different when a sentence is reversed. Lastly, on a side note, if 2014 was the year of CNN's, 2015 the year of RNN's, it looks like 2016 is the year of Neural Attention Mechanisms. Excited to see what new layers are explored in 2016 that will dominate 2017.
- Smerity 11y agoThanks for the great comment! For a direct link to the visual saliency maps from the paper (Figure 6) which @nicklo mentions: http://i.imgur.com/DRfaNxB.png http://i.imgur.com/DRfaNxB.png Being able to see where a neural attention mechanism is fascinating and allows for far more than just introspection. Indeed, one of the earliest and coolest examples are neural attention mechanisms learning to align words between languages without any help[1]! http://i.imgur.com/J5zFZzN.png http://i.imgur.com/J5zFZzN.png New papers on neural attention seem to be coming out every few weeks. I devoted much of my weekend to a naive implementation of hierarchical attentive memory[2] that promises O(log n) lookup (important for improving the speed of attention based algorithms on large input) and can be "taught" to sort in O(n log n) - really exciting stuff! For those interested, I highly recommend reading NVIDIA's "Introduction to Neural Machine Translation with GPUs"[3]. The three part overview is a great introduction. I'd be really happy to see 2016 as the year of neural attention mechanisms :) [1]: http://arxiv.org/abs/1409.0473 http://arxiv.org/abs/1409.0473 [2]: http://arxiv.org/abs/1602.03218 http://arxiv.org/abs/1602.03218 [3]: https://devblogs.nvidia.com/parallelforall/introduction-neural-machine-translation-with-gpus/ https://devblogs.nvidia.com/parallelforall/introduction-neur... (full disclosure: one of the authors of the paper)
- ogrisel 11y agoOut of curiosity, is this work being submitted for a conference or journal? Also have you tried to run DMN on the Text QA / Reading comprehension datasets from DeepMind? Teaching Machines to Read and Comprehend https://github.com/deepmind/rc-data https://github.com/deepmind/rc-data