6 ms·
given that the same model can both: 1. tell me about a cat (given a prompt such as "describe a cat to me") 2. recognize a cat in a photo, and describe the cat
by 2bitencryption 4y ago
given that the same model can both:
1. tell me about a cat (given a prompt such as "describe a cat to me")
2. recognize a cat in a photo, and describe the cat in the photo
does the model understand that a cat that it sees in an image is related to a cat that it can describe in natural language?
As in, are these two tasks (captioning an image and replying to a natural language prompt) so distinct that a "cat" in an image excites different neurons than a "cat" that I ask it about? Or is there overlap? Or we don't know :)
I wonder if you could mix the type of request. Like, provide a prompt that is both text and image. Such as "Here is a picture of a cat. Explain what breed of cat it is and why you think so." Possibly this is too advanced for the model but the idea makes me excited.
- thomashop 4y agoDefinitely possible. OpenAI's CLIP model already embeds images and text into the same embedding space. I don't know exactly how this particular model works but it is creating cross modal relationships otherwise it would not have the capacity to be good at so many tasks.
- ravi-delia 4y agoHow confident are we that it doesn't just have basically 600 smaller models and a classifier telling it which to use? Seems like it's a very small model (by comparison), which is certainly a mark in it's favor.
- sinenomine 4y agoYou can optimize pictures straight through it, and the pictures represent the combinatorial nature of the prompt pretty well. This contradicts the "flat array of classifiers" model.
- Der_Einzige 4y agoYou might find looking into the "lottery ticket hypothesis" fascinating.
- minimaxir 4y agoCLIP has a distinct Vision Transformer and distinct Text Transformer model that are then matmul'd to create the aligned embedding space. Gato apparently just uses a single model.
- sailingparrot 4y ago> OpenAI's CLIP model already embeds images and text into the same embedding space. Not really. Embeddings from images occupy a different region of the space than embeddings from text. A picture of cat and the text "cat" do not resolve to the same embedding. This is why DALL-E has a model learning to translate CLIP text embeddings into CLIP image embeddings before decoding the embedding.
- bungula 4y agoOpenAI actually found these "multimodal neurons" in a result they published a year ago: https://openai.com/blog/multimodal-neurons/ https://openai.com/blog/multimodal-neurons/ Similar to the so-called "Jennifer Aniston neurons" in humans that activate whenever we see, hear, or read a particular concept: https://en.wikipedia.org/wiki/Grandmother_cell https://en.wikipedia.org/wiki/Grandmother_cell
- visarga 4y agoCheck out "Flamingo" https://twitter.com/serkancabi/status/1519697912879538177/photo/1 https://twitter.com/serkancabi/status/1519697912879538177/ph...
- hgomersall 4y agoI think the critical question here is does it have a concept of cattyness? This to me is the crux of a AGI: can it generalise concepts across domains? Moreover, can it relate non-cat but cat-like objects to it's concept of cattyness? As in, this is like a cat because it has whiskers and pointy ears, but is not like a cat because all cats I know about are bigger than 10cm long. It also doesn't have much in the way of mouseyness: it's aspect ratio seems wrong.
- stnmtn 4y agoI don't disagree with you, and I think that what you're saying is critical; but it feels more and more like we are shifting the goalposts. 5 years ago; recognizing a cat and describing a cat in an image would be incredible impressive. Now, the demands we are making and the expectations we keep pushing feel like they are growing as if we are running away from accepting that this might actually be the start of AGI.
- underdeserver 4y agoOf course we are. This is what technological progress is.
- Veedrac 4y agoIf you've seen much DALL-E 2 output, it's pretty obvious they can learn such things. Example: https://old.reddit.com/r/dalle2/comments/u9awwt/pencil_sharpener_designed_by_lamborghini/ https://old.reddit.com/r/dalle2/comments/u9awwt/pencil_sharp....
- webmaven 4y ago> As in, are these two tasks (captioning an image and replying to a natural language prompt) so distinct that a "cat" in an image excites different neurons than a "cat" that I ask it about? Or is there overlap? Or we don't know :) We only have very limited (but suggestive) evidence that the human brain has abstract "cat" neurons involved in various sensory and cognitive modes. Last time I paid attention, there was reasonably strong evidence that an image of a cat and reading the word cat used some of the same neurons. Beyond that it was pretty vague, there seemed to be evidence of a network that only weakly associates concepts from different modes, which isn't consistent with most people's subjective experience. But then we have other evidence that what we think of as our subjective cognitive experience is at least partly a post-hoc illusion imposed for no apparent reason except that it creates an internal narrative consistency (which presumably has some utility, possibly in terms of having a mind others can form theories about more readily).