3 ms·
How do LLMs get information from images? Do they have to run essentially the opposite of an image generation model, taking an image and converting it into a des
by voidUpdate 3mo ago
How do LLMs get information from images? Do they have to run essentially the opposite of an image generation model, taking an image and converting it into a description? I'm just concerned that the description wouldn't be able to encapsulate the information needed to differentiate exactly what is wrong with a shoulder. The image -> text model would need to know what it should actually report back to the LLM about the image, so that it doesn't just say "this is an MRI of a shoulder" or similar. It would be like a layperson describing a bridge, and asking an engineer if the bridge is safe based on that description
- weird-eye-issue 3mo agoNo, it does not work like that, it actually can process the image itself there is not an intermediate image to text step
- voidUpdate 3mo agoHow does a Large Language Model process images then?
- jappgar 3mo agoIt can only deal in tokens, so you're essentially right that it creates a textual description before describing it back to you. This process is obviously incredibly lossy and details are easily missed
- Mr-Frog 3mo agoolder vision LLMs chopped up images into patches which were projected into the same embedding token space as words. Newer ones use an encoder to more efficiently project an image into token space. Then it runs through the same attention layers as the text component.