3 ms·
AFAIK the data does not need to be text.
by firecall 6mo ago
AFAIK the data does not need to be text.
- teaearlgraycold 6mo agoWell diffusers are trained unsupervised on raw pictures. I don't know how they train multi-modal LLMs on images, but yes obviously they are consuming other media than just text. I don't think, but would be happy to be corrected, that models glean much of their "knowledge" from non-textual training data.
- mikert89 6mo agoyou couldnt be more wrong
- teaearlgraycold 6mo agoPlease tell me more. When I ask an LLM a question, and get a text response, can that response incorporate non-textual information from visual training data?