4 ms·
MetaMind papers are always pretty awesome. Couple of great highlights from this paper: The visual saliency maps in Figure 6 are astounding and make their perfo
by nicklo 11y ago
MetaMind papers are always pretty awesome. Couple of great highlights from this paper:
The visual saliency maps in Figure 6 are astounding and make their performance on this task even more impressive as they give a lot of insight into what the model is doing and it seems to be focusing on the things in the image that a regular person would use to decide on the answer. Most striking was on the question "is this in the wild?", and the saliency was on the artificial, human structures in the background that indicated it was in a zoo. This type of reasoning is surprising as it requires a bit of a reversal in logic to come up with this way of answering the question. Super impressive.
The proposed input fusion layer is pretty cool - allowing information from future sentences to be used to condition how to processes previous sentences. This type of information combining was previously not really explored, and it makes sense that it improves the performance on the bAbI-10k task so much as back-tracking is an important tool for human reading comprehension. Also its clever that they encode each sentence as a thought vector before compositing so they can be processed both forwards and backwards with shared parameters- doing so on just one-hot words or even ordered word embeddings would require two vastly different parameters since grammar is wildly different when a sentence is reversed.
Lastly, on a side note, if 2014 was the year of CNN's, 2015 the year of RNN's, it looks like 2016 is the year of Neural Attention Mechanisms. Excited to see what new layers are explored in 2016 that will dominate 2017.
- Smerity 11y agoThanks for the great comment! For a direct link to the visual saliency maps from the paper (Figure 6) which @nicklo mentions: http://i.imgur.com/DRfaNxB.png http://i.imgur.com/DRfaNxB.png Being able to see where a neural attention mechanism is fascinating and allows for far more than just introspection. Indeed, one of the earliest and coolest examples are neural attention mechanisms learning to align words between languages without any help[1]! http://i.imgur.com/J5zFZzN.png http://i.imgur.com/J5zFZzN.png New papers on neural attention seem to be coming out every few weeks. I devoted much of my weekend to a naive implementation of hierarchical attentive memory[2] that promises O(log n) lookup (important for improving the speed of attention based algorithms on large input) and can be "taught" to sort in O(n log n) - really exciting stuff! For those interested, I highly recommend reading NVIDIA's "Introduction to Neural Machine Translation with GPUs"[3]. The three part overview is a great introduction. I'd be really happy to see 2016 as the year of neural attention mechanisms :) [1]: http://arxiv.org/abs/1409.0473 http://arxiv.org/abs/1409.0473 [2]: http://arxiv.org/abs/1602.03218 http://arxiv.org/abs/1602.03218 [3]: https://devblogs.nvidia.com/parallelforall/introduction-neural-machine-translation-with-gpus/ https://devblogs.nvidia.com/parallelforall/introduction-neur... (full disclosure: one of the authors of the paper)
- visarga 11y agoI see this system can take in a few statements and then apply reasoning in order to answer questions. Can we scale up this method to answer questions from Wikipedia pages, scientific papers and articles? Maybe it could be used in conjunction with a chat bot that was trained with RL like AlphaGo (which used millions of human games as input data). People would say something or ask a question, then the net provide a bunch of answers and then people would rate the answers that are most human like, giving it a good/bad signal to use in training the RL part of the system. That would make the bot natural and human-like in conversation. Couple that with the reasoning part for answering questions and we get a reasoned/intelligent chat bot. I'm wondering how far we are from being able to have normal conversations with bots - conversations that don't quickly get derailed or weird.
- neurangotan 11y agoHi thanks for a great paper. I have read it as well as the nvidia articles a lot of times but I am failing to grasp some important details. If you could shed some light that would be great. The attention mechanism is a feed forward single layer neural network. The problem is that the input sequence is variable length. How are the outputs of the bi-rnn fed to the attention mechanism. What happens if I have a 80 word sentence as input and what happens if I have a 10 word sentence as input ?
- evc123 11y agousually zero padding is used; a max_input_length is set somewhere in the code and a number of zeros equal to (max_input_ length - actual_number_of_words_in_input) is appended to the array of input word_ids so that all input sentences have the same length.
- kastnerkyle 11y agoThis depends on implementation (unrolled RNNs vs true recurrence). Each minibatch needs to be the same length, but that is it. And that is even implementation dependent - if your core RNN had a special symbol for "EOS" it could always handle it in another way. Normally you pad each minibatch to the same length (length of the longest sequence in that minibatch), then carry around an additional mask to zero out any "unnecessary" results from padding for the shorter sequences. The BiRNN (using all hidden states) + attention mechanism is the thing that allows variable length context to be fed to the generative decode RNN. A regular RNN (using all hidden states) + attention, or even just the last hidden state of an RNN can all be used to map variable length to fixed length sequences in order to condition the output generator. You will note that padding to the length of the longest sequence in a minibatch wastes computation - people often sort and shuffle the input so that sequences of approximately the same length are used in each minibatch, to maximize computation. If you padded to the overall longest sequence (rather than per minibatch), you would pay a massive overhead computationally.