6 ms·
Deep learning for visual question answering: demo with Keras code
- harperlee 11y agoThere is a very weird behaviour in this image that (to me, at least) speaks a lot about the (low) consistency of this method: if upon the last image a sport-neutral question is asked, then the answer is: 40.52 % tennis 28.45 % soccer But then "Are they playing soccer?" is asked, and the answer jumps to the following: 93.15 % yes What would it happen with tennis then? How can this make sense? My only rationalization would be along the lines of, "Well, it guessed it was a sport at least", but no person would answer these question like that...
- canistr 11y agoHumans without experience with either sport could potentially run into the same difficulty distinguishing them. For instance, how would one tell the difference between Aussie-rule football and rugby based on a picture?
- harperlee 11y agoBut my point is, if the network believes with the most confidence that this is tennis, shouldn't it answer it's not soccer in the next one? Or at least with a hugely less confident answer? The fact that it answers like that makes me think it is not a very stable source of question answering.
- pateheo 11y agoIf the question has some inittial guess like "soccer" then this guess should be added up, as a Bayessian point of view.
- tachyonbeam 11y agoIt's not based on a bayesian system, it's a neural network that does question answering. There are no logical constraints in there. Subtly different wordings of the question could produce different answers, even if the meaning is the same.
- harperlee 11y agoI know, my point is that there doesn't seem to be a part of the system in which the two sports are weighted against, and one consistently outweights the other one. So the model is so sensible that even with exactly the same image, and the same target (identifying this as soccer), you get wild variations by changing the question - and the question not even providing a clear bias.
- iamaaditya 11y agoHi Harperlee I ran the question you asked -- "Are they playing tennis?" and following is the result 99.93 % yes 00.05 % no 000.0 % right 000.0 % left 000.0 % black and white
- harperlee 11y agoThanks for adding more information! This is very interesting. So what's limited in the response then is being able to discriminate among the separate semas of the answer. The network is sure of a lot of common things among the two sports but can't communicate that perhaps tennis is more likely than soccer...
- nickcano 11y agoIdeally there should be a way for the model to say "soccer is a type of sport, let's apply this question to every other sport and only say yes if soccer is the best answer", right? I'm not familiar with Word2Vec, but, presumably, there may be some way to get that information and work it into some logical layer on top of the model. One thing I see all the time with machine learning is trying to solve everything with pure learning, which is quite similar to human intuition in that it can be hard to rationalize. But the thing about humans is that we have these logical constructs that filter our intuition, so it makes sense to create generic abstractions of these constructs and apply them as a sort of filter to the output of a trained model. Of course this has probably been done before, though most of these project I see don't seem to attempt it.
- iamaaditya 11y agoEven though this is "Question Answering", it is trained as a classification model. Thus the model will try to come up with one of the top "1000" answers it has seen during the training. This certainly limits the possibility of answers and sometimes returns very weird answer. It is not for lack of trying that all the top papers in visual question answering end up doing this as a classification task. Results are really poor when it is used as RNN generation, and also extending it more than top 1000 answers does not yield any better results. 87% of the questions in training + validation is within 1000 unique answers. Latest models have started using more complex form of memory and more tightly integrating the question vectors. One of the top model called DPPNet trains a separate matrix from the question vector (chain of GRUs) to find correspondence on the image filter weights. Their idea is that some question have more relevant areas in the image features. Yet another model DMN+, by Metamind uses dynamic memory network which they build to do language question answers but the extension to images work pretty good. Surprisingly the models that use visual attention are not the best and I think it is mostly because this kind of model requires even more data and longer training. Just taking 10 different crop of the question image and doing voting of answer beats attention models (based on numbers reported by these papers). Right now I am working on converting "End to end network" -http://arxiv.org/abs/1503.08895 http://arxiv.org/abs/1503.08895 to this task. I tried working on Neural Turing machine but I could not make it work for this kind of task, but it was mostly because of lack of indepth understanding of NTM. Any feedback from you guys are welcome. P.S Thanks fchollet for writing Keras and for this post. Can't wait to try Keras 1.0
- igul222 11y ago> It is not for lack of trying that all the top papers in visual question answering end up doing this as a classification task. Results are really poor when it is used as RNN generation I'd be curious to know if you have a reference for this. Given that the answers are one word, a word-level RNN language model output should basically be the same thing as a straight 1000-way softmax.
- iamaaditya 11y ago1. Model Q+I [1] Q+I+C [1] ATT 1000 ATT Full ACC. 0.2678 0.2939 0.4838 0.4651 Where ATT Full represents using all the words in the vocabulary, as you can see it performs worse than "Most frequent 1000 answers". Source: Chen, K., Wang, J., Chen, L. C., Gao, H., Xu, W., & Nevatia, R. (2015). ABC- CNN: An Attention Based Convolutional Neural Network for Visual Question Answering. arXiv preprint arXiv:1511.05960. 2. (a) Several early papers about VQA directly adapt the image captioning models to solve the VQA problem [10][11] by generating the answer using a recurrent LSTM network conditioned on the CNN output. But these models’ performance is still limited [10][11] (b) our own implementation of this model is less accurate on [2] than other baseline models Above two quotes are from - Xu, Huijuan, and Kate Saenko. "Ask, Attend and Answer: Exploring Question-Guided Spatial Attention for Visual Question Answering." arXiv preprint arXiv:1511.05234(2015). However, I think my words were sloppy, as I could not find more concrete proof in the literature, but I will revisit them with detail to recollect where I read about RNN generating answers not overachieveing softmax classification over Top K distribution of answers. Also, I would like to note that, I am not using only "one word answers" as the possible set of answers. It contains few two words, and very few three and four word answers. Here is the distribution Key == Length of words | Value == Count of answers with those many words Counter({1: 855, 2: 112, 3: 32, 4: 1})