3 ms·
Looks like the Toronto group have been working on something very similar as well: http://lanl.arxiv.org/abs/1411.2539 http://lanl.arxiv.org/abs/1411.2539 Has a
by benanne 12y ago
Looks like the Toronto group have been working on something very similar as well: http://lanl.arxiv.org/abs/1411.2539 http://lanl.arxiv.org/abs/1411.2539
Has anybody been able to find the Google paper? The article says it's on arxiv, but I can't seem to find it there. All that seems to be published so far is this blog post: http://googleresearch.blogspot.be/2014/11/a-picture-is-worth-thousand-coherent.html http://googleresearch.blogspot.be/2014/11/a-picture-is-worth...
- dumitrue 12y agoJust went out: http://arxiv.org/abs/1411.4555 http://arxiv.org/abs/1411.4555
- iandanforth 12y agoI'm curious to hear your thoughts about learning object saliency from these datasets. Most human generated images have built-in biases toward framing things humans care about, and all of the captions will reflect the relative importance (to humans) of pictured objects. Captioning images, for humans, is a subset of a much more general skill set. Humans can scan a broad visual scene for salient components, focus on those while ignoring non-salient objects, and then organize their thoughts about what has been seen in such a way as to produce an extremely low dimensional description of the scene (a descriptive sentence.) Human's also have the advantage of immediate feedback to their generated descriptions from peers or parents. I haven't seen much work that has attempted to tackle datasets that aren't pre-framed by humans, or ones that try to scale reinforcement learning. I'd love to hear your thoughts or get suggested reading if any pops to mind.