7 ms·
These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like In
by webwielder2 4y ago
These are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
- smlacy 4y agoPrompt Engineering can help a lot but yes, you're basically right: People are generating many, many images and sharing only the best ones with the fewest artifacts. For simple prompts with little additional guidance, all the diffusion image generators I've seen/used will produce output about like what the author linked most of the time. There are always a few gems, and honing in via prompt engineering helps immensely.
- yreg 4y agoI disagree, with proper promptcrafting you can expect far far better results than the one in the op, before any cherry picking. (see my comment sibling to yours) With Dalle-2 I get a satisfying result in >50% of attempts and I'm a beginner. With Midjourney the result almost always looks great, but often misses some part of what I wanted. I'd say Stable Diffusion is similar. The results are seldom crap, but it's difficult to bend it to produce unusual situations. And in SD it's difficult to keep the entire objects in the frame, but that's a different problem.
- gojomo 4y agoOf course people are more likely to share the best iamges – or in this case, the one most illustrative of their concern (about watermarks). Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.
- tnzk 4y agoYou're right. This shows the prompt and it doesn't have such style directives https://ibb.co/gz5RDkB https://ibb.co/gz5RDkB
- gojomo 4y agoAlso, a construction like 'but' that tries to override another expectation may be suboptimal. I gave the same concept a few tries, with more 'sweeteners'. First batch, for prompt "news photo of the King of Belgium giving a speech to an audience that is entirely cucumbers, award-winning, well-composed, detailed surroundings" – & it's a bit better: https://labs.openai.com/s/9YF5WxF1GoZVdzpLBAQYp2Zg https://labs.openai.com/s/9YF5WxF1GoZVdzpLBAQYp2Zg https://labs.openai.com/s/wBfHevs9hIZXvzkJ686mFmn3 https://labs.openai.com/s/wBfHevs9hIZXvzkJ686mFmn3 https://labs.openai.com/s/M0i029fZnYjQXHFobUpw7eun https://labs.openai.com/s/M0i029fZnYjQXHFobUpw7eun (best of batch imo) https://labs.openai.com/s/Hf4z0M9M3KBr6IaaKsjEt9Mx https://labs.openai.com/s/Hf4z0M9M3KBr6IaaKsjEt9Mx A few more tries didn't manage to create any photorealistic shots with actual cucumbers-in-seats – perhaps due to the absurd contrasts required – but shifting to a 'cartoon' style with the prompt "editorial cartoon of the King of Belgium giving a speech to many cheering cucumbers, professional illustrator" got a lot closer: https://labs.openai.com/s/4IonSKYkl0okhNvzJmEAH30K https://labs.openai.com/s/4IonSKYkl0okhNvzJmEAH30K (good) https://labs.openai.com/s/ZeadCzZ9WqeASYXPlOb13wDV https://labs.openai.com/s/ZeadCzZ9WqeASYXPlOb13wDV (good) https://labs.openai.com/s/nERf6bALKEBsQQBvPVsAH7o4 https://labs.openai.com/s/nERf6bALKEBsQQBvPVsAH7o4 (good) https://labs.openai.com/s/7dsZu3bwtfGZZxxJYtg9lTVf https://labs.openai.com/s/7dsZu3bwtfGZZxxJYtg9lTVf If I had more time & credits to burn, I suspect working off those could eventually hit something really apt... but it takes some work & tinkering.
- sh4rks 4y agoOP did say what prompt they used
- gojomo 4y agoI've often seen people show off their autogenerated images and report only approximate paraphrases of their actual prompts. There's one screenshot showing the prompt – but in $CURRENT_YEAR, I view all screenshots with at least a little suspicion, especially when there was a way to highlight the pseuod-watermarked image – OpenAI's native 'Share' – that would've provided stronger proof, direct from OpenAI, of exactly the prompt associated with an image. Hoaxes are everywhere! I've added DALL-E bottom-right color-squares to non-DALL-E images, & seen others do the same, as a subtle joke. So I generally believe the OP, but don't rule-out the possibility there's been tampering to make some point.
- ehsankia 4y agoTop 1% is a bit exaggerated, but there is definitely a lot of not good stuff. I find that Dall-E does especially poorly with underspecified prompts too, unlike something like Midjourney which can give visually pleasing photos for even the most abstract concepts. Dall-E tends to do better with concrete and specific prompts. Here's an example: Stressful Shapes Dall-E: https://i.imgur.com/JBkSh0y.png https://i.imgur.com/JBkSh0y.png Midjourney: https://i.imgur.com/C02Zq3i.png https://i.imgur.com/C02Zq3i.png On the other hand, here's a specific prompt: "nerdy yellow duck reading a magical book full of spells" Dall-E: https://i.imgur.com/FMKZ8zc.png https://i.imgur.com/FMKZ8zc.png Midjourney: https://i.imgur.com/lpsg6af.png https://i.imgur.com/lpsg6af.png
- musicale 4y agoI gather Midjourney was trained primarily using Journey album covers?
- simon_kun 4y agohttps://laion.ai/blog/laion-aesthetics/ https://laion.ai/blog/laion-aesthetics/
- dmitriid 4y agoI find Midjourney to be biased towards an artistic representation (for some definition of artistic) When Dall-e is happy to produce children's scribbles or poor imitations.
- kaetemi 4y agoTry 'poorly drawn ... by a 5 year old using crayons' in Midjourney.
- dmitriid 4y agoEven then Midjourney is more high-quality :) See https://imgur.com/gallery/U5zJMcU https://imgur.com/gallery/U5zJMcU Comparison of two prompts, "poorly futuristic landscape by a 5 year-old" and "poorly drawnn highly detailed futuristic landscape dotted by mahcinery and tall buildings by a 5 year-old" Also, https://imgur.com/gallery/jvEClos https://imgur.com/gallery/jvEClos Comparison of "poorly drawn red sports car in the street of a city by a 5 year-old" Edit: forgot about crayons :D
- pdntspa 4y agoI've been reading some folks saying that "prompt engineering" is a legit future vocation in a world where AI has taken over a lot of creative work And from my experience getting high-quality output from AIs takes a bit of finesse. Not quite unlike crafting a good Google query so... yes
- simon_kun 4y agodoubtful - I'm tuning GPT3 with good midjourney prompts as we speak
- lovemenot 4y agoIs command language the ultimate interface? I doubt it. Similar to how GUI supersedes CLI in most use cases, we should be able to indicate "warmer" / "colder" preference to generate new images from previous attempts.
- pdntspa 4y agoI'm not a huge fan of controlling things with text prompts but it does seem to be the best way to describe the image you're looking for
- aaron695 4y ago
- grumbel 4y agoUm, have you read the prompt? It looking weird is simply the result of "the audience members are cucumbers". The more crazy your prompt is, the worse the results will generally get. On top of that DALL-E2 has generally issues with anything dealing with multiple objects. A single person will render fine, groups of people will generally give artifacts. Attributes will also be spread across all objects in the scenes, not just the ones you specified in your prompt, so doing anything more complex will require manual uncropping und inpainting, not just a single prompt. Anyway, if you avoid the obvious weak spots and holes in the training set, DALL-E2 output is for most part pretty amazing out of the box. It's really more a top 50% than a top 1%. The biggest bias when it comes to published DALL-E2 images are the prompts. Most prompts you see online are not the actual prompts, but funny descriptions made by a human after the fact. The actual prompt are often much longer and sometimes completely different.
- soderfoo 4y agoI have found being as direct as possible and removing duplicate or superfluous words works best. Perhaps this rewrite may yield better results: "King of Belgium gives a speech to an audience of cucumbers"
- yreg 4y agoOp constructed a horrible prompt. First of all, using king Philippe I. is against the ToS, so let's go with a generic "king". Let's not confuse the AI with "buts", just say that he is giving the speech to cucumbers. Lastly, specify some style, because this would probably not work out as a photo. My single try is not bad at all and it could definitely be improved. https://labs.openai.com/s/3OUmUxKefJCeLhAk4hkeKX4V https://labs.openai.com/s/3OUmUxKefJCeLhAk4hkeKX4V
- yreg 4y agoI tried it with Stable Diffusion as well. You can use actual people and the model is even pretty decent at many of the famous ones. On the other hand, it is more difficult to get it to produce absurd results like these. my prompt: King Philippe I. of Belgium giving a speech surrounded by [[[[large green vertical cucumbers]]]], digital art in the style of Greg Rutkowski https://files.catbox.moe/1ej1a4.png https://files.catbox.moe/1ej1a4.png
- throwaway675309 4y agoIs the usage of the surrounding brackets some kind of keyword weights specific to stable diffusion?
- yreg 4y agoI think it's specific to SD. [Square brackets] increase the weight while (simple brackets) decrease the weight. In the cucumber case I used them to force the model to take into account the less believable part of the prompt, because otherwise stable diffusion often ignores such parts.
- JimDabell 4y agoThey are the worst I’ve seen as well. Yes, people tend to share the best of the best. However these results seem especially bad, like bottom 10% bad.
- tkgally 4y agoAs I mentioned here a couple of weeks ago [1], I tested DALL-E with prompts for paintings and drawings in three standard genres: still life, landscape, and portrait. The prompts for portraits yielded a lot of grotesquely unacceptable faces, but almost all of the DALL-E output for the still lifes and landscapes was perfectly fine. [1] https://news.ycombinator.com/item?id=32433821 https://news.ycombinator.com/item?id=32433821
- NoMoreBro 4y agoThere is always some cherry-picking, but prompt engineering is an art per se, you become better and better working on it. I just started this experiment https://www.instagram.com/unshushproject https://www.instagram.com/unshushproject or without Instagram https://unshush.com https://unshush.com and spent A LOT of hours and patience to become good at it. Now I'm very proud of my results and I'm working on doing better. It's a bit risky to invest too much time because every generator is different and they change the underlying model frequently (see the beta of MidJourney yesterday), but if you do it for passion or curiosity there is no problem. Now I'm experimenting with a local installation of Stable Diffusion (well, not really "local" because I have an old computer) and the prompt is only one of the things you can tweak. There are num_inference_steps, guidance_scale and other parameters.
- neurostimulant 4y agoYes, "prompt engineering" in actually a thing. People shares various tips and tricks on the internet to engineer their prompts for best results. Lots of trial and errors required. Example: https://news.ycombinator.com/item?id=32088718 https://news.ycombinator.com/item?id=32088718
- naillo 4y agoThey can definitely be that bad quite frequently. I've actually been a lot happier with stable diffusion outputs lately (doesn't hurt that they're free too).
- InvOfSmallC 4y agoI have access. You get for trials for each query. I have to say that usually there is only one that is good on those three. Sometimes you need to refine your query. I'm pretty impressed as a user.
- smileybarry 4y agoIt definitely requires some very detailed descriptions and sifting through to find a good one. One time I've regenerated a prompt as well because the existing 4 were just not that good. But I did get some great ones at a pretty good usable:unusable ratio.
- grungegun 4y agoFor diversity, Dalle 2 has a random chance of injecting "women" or "black" after a prompt. When this happens, at least for me, it generally destroyed the quality of the images. Probably "King" was identified as a gendered word. You can find some discussion of this on the subreddit r/dalle2. Sometimes the images are quite poor, but in this case, openAI is doing additional tampering. A twitter user figured out which words they were using by generating a lot of images with the starting prompt "A sign being held that says "