4 ms·
Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike
by ehsankia 4y ago
Top 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts.
Here's an example: Stressful Shapes
Dall-E: https://i.imgur.com/JBkSh0y.png https://i.imgur.com/JBkSh0y.png
Midjourney: https://i.imgur.com/C02Zq3i.png https://i.imgur.com/C02Zq3i.png
On the other hand, here's a specific prompt: "nerdy yellow duck reading a magical book full of spells"
Dall-E: https://i.imgur.com/FMKZ8zc.png https://i.imgur.com/FMKZ8zc.png
Midjourney: https://i.imgur.com/lpsg6af.png https://i.imgur.com/lpsg6af.png
- musicale 4y agoI gather Midjourney was trained primarily using Journey album covers?
- simon_kun 4y agohttps://laion.ai/blog/laion-aesthetics/ https://laion.ai/blog/laion-aesthetics/
- dmitriid 4y agoI find Midjourney to be biased towards an artistic representation (for some definition of artistic) When Dall-e is happy to produce children's scribbles or poor imitations.
- kaetemi 4y agoTry 'poorly drawn ... by a 5 year old using crayons' in Midjourney.
- dmitriid 4y agoEven then Midjourney is more high-quality :) See https://imgur.com/gallery/U5zJMcU https://imgur.com/gallery/U5zJMcU Comparison of two prompts, "poorly futuristic landscape by a 5 year-old" and "poorly drawnn highly detailed futuristic landscape dotted by mahcinery and tall buildings by a 5 year-old" Also, https://imgur.com/gallery/jvEClos https://imgur.com/gallery/jvEClos Comparison of "poorly drawn red sports car in the street of a city by a 5 year-old" Edit: forgot about crayons :D
- deleted 4y ago[deleted]
- ehsankia 4y agoThat's honestly one of my favorite prompts. It's funny to think I use this state of the art AI to generate crayon drawings, but they look so great! https://i.imgur.com/jKcNkat.png https://i.imgur.com/jKcNkat.png
- jug 4y agoNow that I'm aware and biased, DALL-E's first image indeed looks very much like stock photo training. This would also make sense given how they can correlate the image with words completely for free due to pretty extensive metadata. What puzzles me is if the Getty Images logo can sometimes appear. If you only have a Getty account, you get rid of the logo and can legally use them royalty free?
- tough 4y agoNo, but you can input your getty image to StableDiffussion img2img and see what's out
- GuB-42 4y agoBut still "king of belgium giving a speech to an audience, but the audience members are cucumbers" is very specific. And I don't see the king of Belgium anywhere, two pictures have absolutely nothing to do with the prompt (no king, no speech, no audience, no cucumber), one has the speech and audience but no king or cucumber. Graphically, they are deep into the uncanny valley. Only the third image is kind of right, if you really stretch your imagination.
- m000 4y agoIIRC, DALL-E filters requests related to politicians/celebrities. A friend had tried to make some funny stuff involving the Greek PM a couple of months ago, and it plainly refused. Now, it seems to process the request, but it will not show anyone resembling the person you asked for.
- Terretta 4y agoI’ve been comparing Dall-E, MidJourney, and StableDiffusion. Goes to show how much training set and implementation choices matter. But in all cases, you have to think of the underlying labeled text-to-image sets as paint colors to mix, and prepare a palette accordingly. Still haven’t figured out how to get what I want, but to your point, one can get closer. - - - Not sure if this is why, but with OpenAI’s Dall-E, you can’t use public figures. You can use proxies, such as “60 year old banker with salt and pepper hair” and then fill in the rest, e.g. “handsome 60 year old banker with salt and pepper hair giving a speech while standing above 12 cucumbers”: https://i.imgur.com/qYKOWM1.jpg https://i.imgur.com/qYKOWM1.jpg Telling it oil painting can fudge who the person is, then pick one that’s close and generate variations: https://i.imgur.com/QRbV7aM.jpg https://i.imgur.com/QRbV7aM.jpg Or use a reasonable photo and then use edit and in-painting to try to improve the implausible subject. This takes a photo from the first prompt above, erases the lower half of image, and makes a new prompt for the lower half, while keeping just enough of the upper half to orient the collage, e.g. “[photo_edit] + banker giving a speech to cucumbers bin full of cucumbers”: https://i.imgur.com/OmOK1HF.jpg https://i.imgur.com/OmOK1HF.jpg - - - Over on MidJourney, where it’s happy to use public figures so long as you’re not violating terms of service about their use, first a couple prompt experiments with King Philippe of the Belgians. https://i.imgur.com/KkgIz2w.jpg https://i.imgur.com/KkgIz2w.jpg https://i.imgur.com/ekf9ypG.jpg https://i.imgur.com/ekf9ypG.jpg Then one upsized plausible painting from among those, where the actual command was “King Philippe of Belgium talking in a large group of cucumbers --q 2 --uplight” which is pretty basic. https://i.imgur.com/QWUaNFv.jpg https://i.imgur.com/QWUaNFv.jpg
- still_grokking 4y ago> On the other hand, here's a specific prompt: "nerdy yellow duck reading a magical book full of spells" > Dall-E: https://i.imgur.com/FMKZ8zc.png https://i.imgur.com/FMKZ8zc.png How well it learned all the common prejudices! "nerdy" == wears glasses I'm applauding. I'm looking already forward to AGI based on the current approaches… It will lead us finally into a better world, for sure. /s
- stubish 4y agoHow do you visually show 'nerdy' without resorting to the glasses stereotype? Your prompt is specifically requesting a prejudiced image.
- still_grokking 4y agoSure. And the AI serves the expected stereotype. Isn't that great? The world will become a better place with AI everywhere. We need especially more AI in law enforcement, and such… AI should make important decisions. Because it bears the same prejudices as humans. So it can replace humans just great. ;-)
- stubish 4y agoThe AI is making no decision. The person entering the prompt made the decision to include the term. It is performing the same function as a pencil. Now, if 'criminal' rendered as a black male 90% of the time rather than a crouched white male wearing a cheesy burglar mask and a sack over his shoulder, then I could see your point about perpetuating prejudice rather than stereotypes.