4 ms·
My prediction: Either these models will have to expose some editable intermediate step, or they have to become as smart as a human. From the users point of vie
by mucle6 3y ago
My prediction: Either these models will have to expose some editable intermediate step, or they have to become as smart as a human.
From the users point of view, a sentence is turned into an image. But there is an underlying structure to the pixels that we aren't allowed to tinker with.
Many choices are made like, placement, color pallete, level of detail, emotion evoked, or sub image descriptions (What should this bush look like, what kind of cloud should this be?)
These image generators are so useful because they fill in all the details you leave out, but you have no ability to tinker with the intermediate details they choose, and you can't just give them a paragraph of details either.
- BoorishBears 3y agoAll of those things are already enabled for these models depending on the UI you're using to accessing. Discord even added a special inpaint UI to their application for Midjourney
- michaelt 3y agoThere are models available that give you more control - in some senses, at least. For example, you can use Stable Diffusion with 'ControlNet' [1] where for example, you can input an 'openpose' to choose the pose of people in the scene. There's also a 'Regional Prompter' [2] which lets you use different prompts for different areas of the image, giving you some control over the composition. You can also use 'inpainting' to regenerate select parts of your image if, for example, you don't like the shape of the clouds. Of course this stuff isn't perfect - for example, you'll get hands with the wrong number of fingers sometimes, no matter what you specify. And you can't easily generate things like multi-frame cartoons without characters clothes changing between frames. [1] https://github.com/Mikubill/sd-webui-controlnet https://github.com/Mikubill/sd-webui-controlnet [2] https://github.com/hako-mikan/sd-webui-regional-prompter https://github.com/hako-mikan/sd-webui-regional-prompter
- ShrigmaMale 3y agoas mentioned controlnet and inpainting help. my guess is img2img will improve to fill this edit step.
- Sohcahtoa82 3y agoFrom what I understand about DALL-E 3 is that it integrates with ChatGPT so you can ask it to change the image with simple language. So if you generated, say, a picture that ended up including a bush, you could ask it to change it from a boxwood to a rose bush, and it could do it. As it is, MidJourney already has the the ability for you to select an area of a generated image and make it regenerate it. You can even change the prompt, though it won't always do what you think. I think a lot of the challenges you're describing won't apply to DALL-E 3's ChatGPT integration.