5 ms·
Seeing a lot of text-to-image out there recently. Does anyone know what the current state of the art is on image-to-text? Thinking something similar to Midjourn
by marvinkennis 3y ago
Seeing a lot of text-to-image out there recently. Does anyone know what the current state of the art is on image-to-text? Thinking something similar to Midjourney's /describe command that they added in v5
- mkaic 3y agoWhile it's not publicly available yet, I have strong suspicions that multimodal GPT-4 may actually be SOTA in image-to-text. The examples shown in the Sparks of AGI paper were extremely impressive imo, though of course those are cherry-picked so it's unclear how well the model will perform on non-cherry-picked images.
- jah242 3y agoThis is text + image -> text but pretty cool and still might be of interest to you: https://llava-vl.github.io https://llava-vl.github.io
- marvinkennis 3y agoJust entering "Describe this image" in the chat prompt got me exactly what I was looking for. Thanks!