5 ms·
The way I see it, input being confined to a "text description" is the next immediate problem that needs to be solved. I don't think we can rely on textual input
by educaysean 4y ago
The way I see it, input being confined to a "text description" is the next immediate problem that needs to be solved. I don't think we can rely on textual inputs for much longer as the human language is too imprecise and/or verbose. It's hard to imagine what exactly the optimal interface would be, but I'm thinking we'll need ways to dictate attributes for each entity being represented, the backdrop, and the view composition all as separate individual components. Ideally all these components can also be reusable and provide reproducibility guarantees without having to share a global "seed" as well.
- pishpash 4y agoLike an AI-assisted photoshop, but that's not restricted by language, only interactivity. Down the line you'll need a direct mind meld because some ideas don't have words to describe them, but that's not the "next immediate" problem.
- ihuman 4y agoPeople quickly created photoshop plugins to integrate stable diffusion when SD came out https://twitter.com/wbuchw/status/1563162131024920576 https://twitter.com/wbuchw/status/1563162131024920576
- nicd 4y agoA "mood board" of inspiration images could be an interesting input method. With deep vision models, we can already separate different "levels" of concepts: a high level subject ("person riding a horse"), textures (the horse hair), medium (painting vs 3d rendering vs photograph), etc. It'd be interesting to have a "smart mood board" that goes from text prompt, to visualizing that hierarchy with different options. Then the user could interactively increase or decrease different parameters, ultimately iterating alongside the computer to realize their creative vision.
- ReactiveJelly 4y agoThat one person already used the other half of the auto-encoder to go from an image of "ugly Sonic" from the Sonic The Hedgehog movie, to a special "{sonic}" token, and then they could put that into prompts. It is not far off.