4 ms·
DALL-E within ChatGPT uses GPT-4 to rewrite what you ask for into a good text-to-image prompt. You could probably do something similar with Stable Diffusion wit
by fassssst 3y ago
DALL-E within ChatGPT uses GPT-4 to rewrite what you ask for into a good text-to-image prompt. You could probably do something similar with Stable Diffusion with just a little upfront effort tuning that system prompt.
- IanCal 3y agoSomewhat, but dalle3 is hugely better at understanding a description and relationships.
- dragonwriter 3y agoLLMs in general are, and that can be leveraged by using an LLM to set up layout for Stable Diffusion. https://github.com/TonyLianLong/LLM-groundedDiffusion https://github.com/TonyLianLong/LLM-groundedDiffusion
- dragonwriter 3y ago> You could probably do something similar with Stable Diffusion with just a little upfront effort tuning that system prompt. And, indeed, someone has: https://github.com/sayakpaul/caption-upsampling https://github.com/sayakpaul/caption-upsampling