4 ms·
I'm particularly interested how GPT-4 manages multi-modal processing. Do the images share the same domain as the text inputs, or is there some location in the m
by numberalltheway 4y ago
I'm particularly interested how GPT-4 manages multi-modal processing. Do the images share the same domain as the text inputs, or is there some location in the model inputs that is ~for images only~. The Technical Report states that "the model generates text outputs given inputs consisting of arbitrarily interlaced text and image"[1], but that doesn't really clear up how the images are being treated here.
[1] https://arxiv.org/pdf/2303.08774.pdf https://arxiv.org/pdf/2303.08774.pdf