3 ms·
I would be more interested in image-to-text models. Does someone know of any decent model? I saw the GPT4 demo, and they showed that they do image-to-text... bu
by danwee 3y ago
I would be more interested in image-to-text models. Does someone know of any decent model? I saw the GPT4 demo, and they showed that they do image-to-text... but then that was actually a fake (i.e., the model was interpreting the image filename).
- jkea 3y agoMidJourney has the describe function which is kind of like that. Not sure how decent it actually is
- theRealMe 3y agoCan you provide a source on gpt4 image model being fake? I haven’t heard that before, though I have wondered why I haven’t heard anything about the image part and don’t have access to image processing myself.
- runnerup 3y agoAFAIK, converting an image to a text summary isn't really a thing by itself. The related work would be "visual reasoning" which is the ability to ask things about the image in natural language and get responses back also in natural language. I believe the current SOTA test for NLVR is VQAv2[0] or GQA[1]. 0: https://visualqa.org/ https://visualqa.org/ 1: https://arxiv.org/pdf/1902.09506.pdf https://arxiv.org/pdf/1902.09506.pdf
- minimaxir 3y agoFor a fast-but-less-robust model, you can use a ViT encode/GPT-2 decoder model: https://huggingface.co/nlpconnect/vit-gpt2-image-captioning https://huggingface.co/nlpconnect/vit-gpt2-image-captioning For a more-robust-but-hard-to-run model, you can use BLIP2: https://huggingface.co/Salesforce/blip2-opt-2.7b https://huggingface.co/Salesforce/blip2-opt-2.7b
- grumbel 3y agoCLIP Interrogator[1], which is also build into AUTOMATIC1111, gives quite reasonable results, at least if all you need is a prompt, it can't handle complex interactions: Image: https://i.imgur.com/husplYZ.png https://i.imgur.com/husplYZ.png Output: "a white horse with a sign that says rexel's in space, pixelperfect, inspired by Paul Kelpe, official simpsons movie artwork, alternate album cover, in style of nanospace, by Apelles, pickles, pespective, pop surrealism, ingame, in a space cadet outfit, sifi" [1] https://huggingface.co/spaces/pharma/CLIP-Interrogator https://huggingface.co/spaces/pharma/CLIP-Interrogator
- jah242 3y agoThis is text + image -> text but pretty cool and still might be of interest to you: https://llava-vl.github.io https://llava-vl.github.io