3 ms·
I think this is mainly theoretical at this point. In my experience, current technology doesn’t seem to be utilizing the additional information that comes from n
by throwaway1851 4y ago
I think this is mainly theoretical at this point. In my experience, current technology doesn’t seem to be utilizing the additional information that comes from natural language all that well. For example: prompt Dalle2 for “a dinner plate on a stack of pancakes” and you will get ordinary images of pancakes on plates, not the other way around.
Edit: an experiment comparing tags/BOW vs natural language sequences in image generation tasks would be interesting to see.
- GaggiX 4y agoI think this does not work mainly because of the unusual situation you describe, such as "a horse riding a person"; most of the time Dalle 2 is really good at following the prompt.