4 ms·
The python zoom in seems performative. A vision model already has access to all the data, how does zooming in help it? Still very cool that it can!
by tippytippytango 1y ago
The python zoom in seems performative. A vision model already has access to all the data, how does zooming in help it? Still very cool that it can!
- energy123 1y agoYeah, once it gets converted into tokens how does "zooming in" somehow increase information content?
- nutrientharvest 1y agoIt's cropping the original image then tokenizing it again with less downsampling, not cropping its internal representation.
- simonw 1y agoYeah, I'm a little unconvinced by that. My best guess there is that the vision input has quite a restricted resolution and "zooming in" (really, cropping to an area) lets it get more information about the region of the photo because it's not as "fuzzy". Just a hunch though.
- Legend2440 1y agoVision models are typically bad at small details. If there’s too much stuff going on at once, they can’t focus on the entire image.