3 ms·
I feel like multimodal models that can read images should work differently than they do. My understanding is that multimodal models basically first generate an
by foota 14d ago
I feel like multimodal models that can read images should work differently than they do. My understanding is that multimodal models basically first generate an image embedding and then the model is trained to interpret that embedding, but in the same way that text is lossy, it seems like the embedding would be as well. Why don't multimodal models learn to interpret images themselves without an embedding? Or e.g., by passing some "prompt" to the embedding model?
- thfuran 14d agoWhat does interpreting images mean in practice if you exclude the possibility of feature extraction or any other sort of implicit embedding?
- foota 14d agoI'm not an ML expert, but I was thinking of a sort of "guided" embedding. E.g., give the image model some prompt for what it's trying to do? I don't understand why multimodal models generate an embedding that doesn't understand what the model is trying to "figure out". I think this is similar to how Gemma 4 12B is implemented, but even then I don't think the single layer image embedding is "aware" of the context.
- seanhunter 13d agoYou already do give the image model a prompt to tell it what to do. That’s not something the embedding can use independently of how the model is already using it. In general an embedding doesn’t have intent or awareness in the way you’re looking for. “Embedding” just means one mathematical structure stuffed inside another. So for example the real number line is embedded into the Cartesian plane as each axis- that’s an embedding. Now in this case specifically, the embeddings in any kind of transformer model encode the meaning of the thing they represent into vectors (which is what the model itself actually operates on). You can train the embedding to be more useful for a particular task at inference time, which already happens.
- foota 12d agoI think what I'm saying is that I don't understand why the embedding exists. I assume it's some kind of training and inference cost issue? But why can't the Gemma architecture linked above just learn to represent pixels in the LLM model's embedding space directly, rather than having the embedding from 48 x 48 pixel chunks? Or rather, give the embedding model some context to produce the embedding? (Which, as you note, wouldn't really be an embedding anymore, but seems like it would better understand fine detail)