3 ms·
I'm not sure I see the behavior in the Gemini 2.0 Flash model's image output as a strength. It seems to me it has multiple output modes, one indeed being masked
by blixt 1y ago
I'm not sure I see the behavior in the Gemini 2.0 Flash model's image output as a strength. It seems to me it has multiple output modes, one indeed being masked edits. But it also seems to have convolutional matrix edits (e.g. "make this image grayscale" looks practically like it's applying a Photoshop filter) and true latent space edits ("show me this scene 1 minute later" or "move the camera so it is above this scene, pointing down"). And it almost seems to me these are actually distinct modes, which seems like it's been a bit too hand engineered.
On the other hand, OpenAI's model, while it does seem to have some upscaling magic happening (which makes the outputs look a lot nicer than the ones from Gemini FWIW), also seems to perform all its edits entirely in latent space (hence it's easy to see things degrade at a conceptual level such as texture, rotation, position, etc.) But this is a sign that its latent space mode is solid enough to always use, while with Gemini 2.0 Flash I get the feeling when it is used, it's just not performing as well.