3 ms·
Speaking as someone with only a tenuous grasp of how VLMs work, this naïvely feels like a place where the "embodiement" folks might have a point: Humans have th
by akavi 2y ago
Speaking as someone with only a tenuous grasp of how VLMs work, this naïvely feels like a place where the "embodiement" folks might have a point: Humans have the ability to "refine" their perception of an image iteratively, focusing in on areas of interest, while VLMs have to process the entire image at the same level of fidelity.
I'm curious if there'd be a way to emulate this (have the visual tokens be low fidelity at first, but allow the VLM to emit tokens that correspond to "focusing" on a region of the image with greater resolution). I'm not sure if/how it's possible to performantly train a model with "interactive" data like that, though
- slashdave 2y agoThese models have learned to focus on specific portions of an image (after all, this is the stated purpose of a transformer).
- efskap 2y agoIsn't this the attention mechanism, the reason we're using transformers for these things? Maybe not greater resolution per se, but focusing on a region with greater neural connectivity
- akavi 2y agoAh, good point! But the model is downstream of the "patch" tokenization, so the cut-down in resolution (compression) of the image has already occurred prior to the point where the model can direct greater "attention". I think the synthesis is that I'm proposing a per-pixel tokenization with a transformer block whose purpose is to output information at a compression level "equivalent" to that of the patch tokens (is this what an autoencoder is?), but where the attention vector is a function of the full state of the LLM (ie, inclusive of the text surrounding the image)). Naïvely, I'd think a layer like this that is agnostic to the LLM state needn't be any more computationally costly than the patching computation (both are big honks of linear algebra?), but idk how expensive the "full context attention" feedback is... (I apologize to anyone who actually understands transformers for my gratuitous (ab|mis)use of terminology)
- Brechreiz 2y ago>Humans have the ability to "refine" their perception of an image iteratively That's not related to embodied cognition.
- akavi 2y agoIs embodied cognition not at least in part about interactivity? I perform action (emit tokens) and receive feedback (non-self-generated tokens)
- kromem 2y agoLots and lots of eye tracking data paired with what was being looked at in order to emulate human attention processing might be one of the lower hanging fruits for improving it.
- caddemon 2y agoHumans are actually born with blurry vision as the eye takes time to develop, so human learning starts with low resolution images. There is a theory that this is not really a limitation but a benefit in developing our visual processing systems. People in poorer countries that get cataracts removed when they are a bit older and should at that point hardware-wise have perfect vision do still seem to have some lifelong deficits. It's not entirely known how much early learning in low resolution makes a difference in humans, and obviously that could also relate more to our specific neurobiology than a general truth about learning in connectionist systems. But I found it to be an interesting idea that maybe certain outcomes with ANNs could be influenced a lot by training paradigms s.t. not all shortcomings could be addressed with only updates to the core architecture.