2 ms·
And if it the model is supposed to be so attentive to context, why did it show a desert instead of "dessert"? After all, they just ate "launch".
by neckro23 2y ago
And if it the model is supposed to be so attentive to context, why did it show a desert instead of "dessert"? After all, they just ate "launch".
- yorwba 2y agoThe model can only attend to context that is part of the input. Most likely they created the image grid by independently feeding the model each prompt together with the reference image. (And the point is to show off that the model output remains consistent despite this independent generation process.)