4 ms·
AFAIK, converting an image to a text summary isn't really a thing by itself. The related work would be "visual reasoning" which is the ability to ask things abo
by runnerup 3y ago
AFAIK, converting an image to a text summary isn't really a thing by itself. The related work would be "visual reasoning" which is the ability to ask things about the image in natural language and get responses back also in natural language.
I believe the current SOTA test for NLVR is VQAv2[0] or GQA[1].
0: https://visualqa.org/ https://visualqa.org/
1: https://arxiv.org/pdf/1902.09506.pdf https://arxiv.org/pdf/1902.09506.pdf