4 ms·
I'm a bit surprised by the examples. The original captions contain extra context rather than descriptions of the images. Perhaps that doesn't matter, as the ab
by interloxia 2y ago
I'm a bit surprised by the examples. The original captions contain extra context rather than descriptions of the images.
Perhaps that doesn't matter, as the abstract suggests that's not the purpose.
On the other hand, training on captions they generate that are also incorrect is problematic. Perhaps they should have cherry-picked better examples without errors, such as details about the sugar and cream, reflections of the train in the water, or the location of the watermark and sign. I didn't look at the code, but perhaps each image feature can be written with a confidence score.
- petercooper 2y agoYes, the "Western Kingbird" is a good example of that. That's very useful context. I've not read the paper yet, but this could be trivially solved by just appending the extended description to the original one. You could ask the model if the original description seems accurate or not as well to weed out any duds (like the picture of a cake described as a 'twin room').
- skerit 2y ago> The original captions contain extra context rather than descriptions of the images Another good example on how accessibility can even help people without disabilities. You can see this on ALT text on Mastodon, and I find it really useful.